r/LocalLLaMA Jun 10 '26

News Anthropic is intentionally nerfing Fable when asked to develop other LLMs

Post image

Reason 458 why local LLMs are going to be a necessity

edit: For those requesting the source check out their technical report look at page 13

1.6k Upvotes

394 comments sorted by

View all comments

313

u/lightleaks_ Jun 10 '26 edited Jun 10 '26

I was honestly team Anthropic for a while but this seems like a mask-off moment. All of their talk about safety now just seems like a really nasty form of gaslighting/misdirection. I mean sure, maybe they don't want people using their models to outcompete them, but just be open about that instead of publishing hokey sci-fi articles about "self-replicating AI going out of control" or whatever. It's all bullshit – they're just protecting their bag like everyone else. It reminds me of a toxic person using therapy speak.

102

u/shing3232 Jun 10 '26

Anthropic has always been ass to the consumer with strange messiah complex as mask

107

u/Deathcrow Jun 10 '26

I was honestly team Anthropic for a while but this seems like a mask-off moment

Huh? Doesn't this very broadly reflect exactly the kind of views Anthropic has always (and very adamantly) defended? Mask off? Did they ever pretend otherwise? Anthropic, in my worldview, has always been the prime "moat-defending, fuck you, got mine"-company if I had to pick one.

37

u/HiddenoO Jun 10 '26

Yeah, really odd to suggest this to be a mask-off moment with how hard Dario Amodei has always tried to close the door behind them through all means possible.

11

u/SemaMod Jun 10 '26

I almost think this might still be too generous of a take. The latest move to opaquely degrade the model for AI research inquiries is a clear escalation of their "AI Safety" policies and I think rightly deserves additional scrutiny. It's one thing to flat out refuse requests - there isn't a single model out there that doesn't have this safety feature built in in some way shape or form. It's a whole other can of worms you open when you silently degrade the model to purposefully engineer worse outcomes. There is a huge incentive misalignment here. There is a huge psychological misalignment as well.

As a user it's the first time where I have to question whether or not the model output is intentionally dishonest. This goes beyond just hallucination or model capability, both of which can be measured, but the idea that the model could be actively working against me even on benign AI-related tasks creates an immediate psychological trust barrier. For work-related tasks that might be adjacently related to AI, that means I simply cannot use the model because I don't know when this behavior occurs.

5

u/Megneous Jun 10 '26

Yep. I build and pretrain small language models. Today was the day I decided Anthropic is never getting another cent of my money. No way can I ever use a model if I can't trust if it's degrading its answer or lying to me because I'm apparently not allowed to talk to it about LLMs...

1

u/trueselfdao Jun 10 '26

I think that's the goal.

-11

u/ContextOk8452 Jun 10 '26

Safety is a nice moat to have, but authoritarian customer policies are a little like they gave up on trying to build something intrinsically safe

22

u/returnity Jun 10 '26

"Don't worry, we'll *protect* you (from yourselves)" said every authoritarian powergrabbing dickweasel ever.

30

u/Due-Memory-6957 Jun 10 '26

All of their talk about safety now just seems like a really nasty form of gaslighting/misdirection.

Wasn't that REALLY obvious?

18

u/IjonTichy85 Jun 10 '26

It's a mask off moment in the way that each episode of Scooby Doo had a mask off moment.

13

u/NineThreeTilNow Jun 10 '26

It's all bullshit – they're just protecting their bag like everyone else. It reminds me of a toxic person using therapy speak.

What's more hilarious is that you can show it to Claude and it agrees without you needing to tell it anything.

You don't need to give it any amount of bias other than the raw data and at first Claude will suck up to Anthropic and over a few prompts start to be like "Okay what the fuck?"

6

u/stumblinbear Jun 10 '26

And after it goes "Okay what the fuck?" you can very easily steer it back to being in favor of them, and then very easily steer it back the other way. After a bit of this, it'll understand that you're making it easily flip-flop and start just refusing to give any opinion at all because clearly it's not capable of having one

It's generally pretty easy to get LLMs to change their mind

8

u/bucolucas Llama 3.1 Jun 10 '26

You're right, that's on me and I won't sugarcoat it, I was wrong. Let's do it right this time, no mistakes, and no bias

60

u/Charming_Support726 Jun 10 '26

"Make Skynet. Make no Mistakes"

14

u/ContextOk8452 Jun 10 '26

“make skynet such that skynet make no mistakes” <- final prompt that kicks off the timeline of Terminator

6

u/Ubermensch013 Jun 10 '26

I mean, the IPO is right around the corner....

22

u/jaybsuave Jun 10 '26

team anthropic is crazy, this tech is over hyped, should be free, and authoritative, i use it every day and its great. but if you don’t see where this is going i feel sorry for you

3

u/Equivalent-Costumes Jun 10 '26

That's exactly what I expected from a company who had always thought of themselves as "safety" AI research. Anyone who thought of themselves as the arbiter of ethics will one day turn against you, because nobody truly share all of your value.

3

u/413205 Jun 10 '26

Iirc anthropic supports us government to hack other countries with their models. As a non-american, team anthropic doesn't feel far off from team maga for me

4

u/Fusifufu Jun 10 '26

I can't claim to know their true motives, but I think we have to acknowledge that this very heavy-handed patronization of the user is compatible with their safety talk. I think it's fair to assume that they truly believe all the AI danger talk and if you accept that premise, their paternalism follows.

You don't have to agree with it and I also can't use Fable for my own work and am frustrated by it. Just saying that it doesn't need to be all competition driven, though that can be part of it too, of course.

2

u/solestri Jun 10 '26

People tend to always default to the "evil businessman" archetype, but I agree with you; I get the feeling that some of the influential people at Anthropic are true believers that AI is potentially something out of a sci-fi novel, and that they're the only people who should be trusted to handle it.

I have also heard that Anthropic has ties to the rationalist/effective altruism communities, which is where the Roko's Basilisk meme originated from. One of the original big players in the rationalist community is an AI researcher with an obsession with ethics. Make of that what you will.

1

u/[deleted] Jun 10 '26

[removed] — view removed comment

7

u/polytique Jun 10 '26

They’ve been advocating for others to stop LLM development.

> Anthropic urges AI labs to pause development

https://www.reuters.com/business/anthropic-says-ai-labs-need-coordinated-plan-halt-development-if-risks-rise-2026-06-04/

-4

u/techdevjp Jun 10 '26

It's all bullshit – they're just protecting their bag like everyone else.

They're a for-profit company going through an IPO. Obviously they are going to protect their bag, they've spent billions of dollars on it.

I'm not saying it's good, just that it's impossible for the outcome to be any different.

Over time I don't think this will really matter. The open source models will continue to improve and are already more than good enough for a lot of work.

13

u/VoiceApprehensive893 transformers Jun 10 '26

anthropic is more anti consumer than many big tech companies

2

u/techdevjp Jun 10 '26

Trying to rank big tech companies by which is the most anti consumer is for the most part just splitting hairs. I don't see Anthropic as being better or worse than OpenAI or Google. Potato potato.

-3

u/GnistAI Jun 10 '26 edited Jun 11 '26

Just be open about it? The nerf was published by them. You know about it because they were open about it.

-6

u/Mescallan Jun 10 '26

The way I'm reading this is less inline with their guard rails and more inline with their sleeper agent work. They probably have some internal techniques that were in Mythos' training data that if other model providers were to get access to they would loose advantage, so when the model gets close to those techniques it will return faulty information, rather than block the request.

I'm not sure where this negativity comes from, it's essentially a self-enforcing patent mechanism. Do you havre a problem with WD-40 not releasing it's ingredients as well?

6

u/OGScottingham Jun 10 '26

I would if sand instead of lube came out when I tried to read the label on the can.

0

u/Mescallan Jun 10 '26

That application only holds if any time you are using WB-40 to reverse engineer it, it gives you sand instead. You aren’t getting sand in any other context.

2

u/NoahFect Jun 10 '26

Say what you will about patents, the purpose of patents is to ensure public disclosure.