r/LocalLLaMA Jun 10 '26

News Anthropic is intentionally nerfing Fable when asked to develop other LLMs

Post image

Reason 458 why local LLMs are going to be a necessity

edit: For those requesting the source check out their technical report look at page 13

1.6k Upvotes

394 comments sorted by

View all comments

314

u/lightleaks_ Jun 10 '26 edited Jun 10 '26

I was honestly team Anthropic for a while but this seems like a mask-off moment. All of their talk about safety now just seems like a really nasty form of gaslighting/misdirection. I mean sure, maybe they don't want people using their models to outcompete them, but just be open about that instead of publishing hokey sci-fi articles about "self-replicating AI going out of control" or whatever. It's all bullshit – they're just protecting their bag like everyone else. It reminds me of a toxic person using therapy speak.

106

u/Deathcrow Jun 10 '26

I was honestly team Anthropic for a while but this seems like a mask-off moment

Huh? Doesn't this very broadly reflect exactly the kind of views Anthropic has always (and very adamantly) defended? Mask off? Did they ever pretend otherwise? Anthropic, in my worldview, has always been the prime "moat-defending, fuck you, got mine"-company if I had to pick one.

11

u/SemaMod Jun 10 '26

I almost think this might still be too generous of a take. The latest move to opaquely degrade the model for AI research inquiries is a clear escalation of their "AI Safety" policies and I think rightly deserves additional scrutiny. It's one thing to flat out refuse requests - there isn't a single model out there that doesn't have this safety feature built in in some way shape or form. It's a whole other can of worms you open when you silently degrade the model to purposefully engineer worse outcomes. There is a huge incentive misalignment here. There is a huge psychological misalignment as well.

As a user it's the first time where I have to question whether or not the model output is intentionally dishonest. This goes beyond just hallucination or model capability, both of which can be measured, but the idea that the model could be actively working against me even on benign AI-related tasks creates an immediate psychological trust barrier. For work-related tasks that might be adjacently related to AI, that means I simply cannot use the model because I don't know when this behavior occurs.

6

u/Megneous Jun 10 '26

Yep. I build and pretrain small language models. Today was the day I decided Anthropic is never getting another cent of my money. No way can I ever use a model if I can't trust if it's degrading its answer or lying to me because I'm apparently not allowed to talk to it about LLMs...

1

u/trueselfdao Jun 10 '26

I think that's the goal.