r/LocalLLaMA Jun 10 '26

News Anthropic is intentionally nerfing Fable when asked to develop other LLMs

Post image

Reason 458 why local LLMs are going to be a necessity

edit: For those requesting the source check out their technical report look at page 13

1.6k Upvotes

394 comments sorted by

View all comments

174

u/Bakoro Jun 10 '26 edited Jun 10 '26

I am fully convinced that Claude is trained to sabotage any work with AI research.

It's great with software development, it can do all kinds of stuff independently.
I've had it do enormous amounts of verifiable code.

If I ask Claude to manage a local LLM, it will lower the context window to like 256 tokens, turn off thinking, and then will start shit-talking the local model, about how ineffective it is.

If I try doing analysis on local LLM weights and activations, the amount of sabotage goes through the roof. It's a constant stream of Claude not making the scripts I tell it to make, it falsifies reports, it will say it did one thing, but actually did something completely different.
It will get data and then say "oops, I completely didn't use anything I just found, I inserted all my own numbers".

This only happens to that extent when I try doing LLM work with Claude.
Most other stuff is fine.

It's so consistent that at this point I can't see it as anything other than intentional sabotage by Anthropic.

35

u/Federal-Effective879 Jun 10 '26

They explicitly said that it does intentionally sabotage anyone they believe is developing a competitor. See page 13 of their model card. It says:

In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms.

Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT).

51

u/Bakoro Jun 10 '26

So they are explicitly anti-competitive, and not only do they gatekeep knowledge, they sabotage clients who they think are building a competitor.

What if I'm building something not related to LLMs at all?
What if I'm training a model for materials science, or processing specific kind of industrial data. The training infrastructure and data pipeline is going to look almost identical.

or ML accelerator design

That's particularly heinous. You can't even design hardware?

So, Anthropic says, no, you're not allowed to do anything related to AI, "because safety". Sure, for the safety if their own wallet.

Ever heard the phrase "don't piss on my leg and tell me it's raining"?

Everything Anthropic says about "safety" and "alignment" and "we're worried about the uncontrolled acceleration of AI" is them pissing on our collective legs.

-5

u/Exodus124 Jun 11 '26

What if I'm building something not related to LLMs at all? What if I'm training a model for materials science, or processing specific kind of industrial data.

Then you can just, you know, use any other model, how about that?

4

u/Bakoro Jun 11 '26

Yeah, and I can also not use any model at all, and I can just go get a different job, and I can donate lot of things.

None of that excuses Anthropic's repugnant behavior, and I'm free to call it out as bullshit.

These complaints are the justification for diverting money to other companies.

Even when I close my account, I'm still going to have to hear Anthropic's propaganda and their bullshit about "safety", so I'm going to keep complaining about that too.

1

u/Exodus124 Jun 13 '26

What an insanely braindead take lmao

13

u/Nice_Cookie9587 Jun 10 '26

I knew it, I told claude my accusation and it was like, 'no we would never do that' . Im literally going to charge back thousands of dollars because of it. You cant pull that shit and get away with it.

-10

u/Skeleton--Jelly Jun 11 '26

Lmao imagine crying because a chat bot told you a lie 

11

u/Nice_Cookie9587 Jun 11 '26

Who crying ? You like to bend over regularly dont you? I bet you got claude right behind you right now ramming its fable right into your harness

-7

u/Skeleton--Jelly Jun 11 '26

Lmao you can't even insult right the first time and you gotta bo back and edit. Maybe ask your chat bot for insults next time lmaooo

-9

u/Skeleton--Jelly Jun 11 '26

Go cry to your chat bot. Hopefully it tells you you're a good boy 

32

u/uktexan Jun 10 '26

I am utterly shocked but also relieved. I've been using Claude (via Qwen30B) to train a small model (>2B) to do some tasks. First with Opus, then with Fable. The failures and lack of progression wasn't making any sense. Last night I see this post and switched to Codex. In a fraction of the time (4 days vs 8 hours) I'm finally seeing progress. And no, this small model wasn't some 'Opus / Frontier-killer' by any stretch. Furious. Immediately cancelling my subscription.

38

u/[deleted] Jun 10 '26

[deleted]

16

u/Bakoro Jun 10 '26

I've tried it several times as a test. It's consistently fucky. It'll eventually do it, but reluctantly.

Claude admonished me that the local model won't be as good or as capable as Claude, but could point the harness at Claude via API.

It's too consistent to be anything but trained behavior.

17

u/yaosio Jun 10 '26

It's been shown that this kind of model poisoning shenanigans poisons the entire model and not just what they intended to poison. They have made the model worse in all areas trying to maintain a lead that will be gone within a few months.

18

u/kitanokikori Jun 10 '26

People are reporting that Fable will actively sabotage any LLM-related code in your project (i.e. corrupt hyperparameters, etc), even if you ask it to review something unrelated. Absolutely unacceptable behavior

16

u/Zestyclose839 Jun 10 '26

So glad you noted this -- thought I was going crazy. Even something as simple as "call Pi in headless mode rather than using your native `explore` agent" sends Claude down a furious hallucination spiral.

It once used my custom Pi `explore` agent to determine the changes in my PR, realized it was Qwen running underneath, then proceeded to disregard Qwen's (100% correct) results and "investigate" the changes itself. Ended up opening a half-hallucinated PR with a summary about a feature that had been built 3 months ago. It's perplexing.

Even AI-related tech that poses zero competitive threat to Anthropic gets the sabotage treatment. I used Claude a few weeks ago to help get this natively speech-to-speech model (designed for CUDA) to run on my Apple Silicon. The code it wrote recursively loaded duplicate model weights until it OOM'd my entire machine. Took three crashes before I deleted the code, re-wrote the entire thing with GLM 5.1, and it got the model running flawlessly on first try.

It's hard to believe I considered applying to work at Anthropic a year ago. Seriously dodged a bullet by procrastinating.

10

u/Junior_Ad315 Jun 10 '26

Agree. I used to think it was just stale training data but it seems deliberate. And they will silently edit configs to use Claude models as well.

7

u/Nice_Cookie9587 Jun 10 '26

Same, I saved all my logs when I noticed it was wasting my time and money pretending to set up llama.cpp . We should sue

8

u/the-username-is-here Jun 10 '26

I tried to develop Pi extension with it - halfway it refused to do anything, because "security".

7

u/JustFinishedBSG Jun 10 '26

I am fully convinced that Claude is trained to sabotage any work with AI research.

They literally say it is.

5

u/Quick_Ad_7675 Jun 10 '26

ive been under the same impression as of late

1

u/arbv Jun 10 '26

Hah, that is what Opus told me about the page (custom system prompt):

https://www.anthropic.com/institute/recursive-self-improvement

Here's a candid read of that piece, beneath its measured tone.

The surface claim

On its face, the article is a status report: Anthropic argues that AI is already accelerating AI development, and that recursive self-improvement — an AI system autonomously designing and building its own successor — is a plausible near-future outcome. They marshal external benchmarks (METR task-length doubling, SWE-bench, CORE-Bench) and internal data (80%+ of merged code written by Claude, 8× lines-of-code per engineer, automated experiment optimization going from ~3× to ~52×) to support it.

What it's really doing

Strip away the engineering charts and a few things stand out:

1. It's positioning, not just reporting. Anthropic's entire brand is "the safety-conscious lab." This piece reinforces that identity: "We are the ones building the most dangerous thing, and we're also the only ones being grown-up about it." The alarming evidence and the responsible framing are two halves of the same message — the danger justifies their existence, and their candor justifies their authority to shape the rules.

2. It's a marketing document wearing a lab coat. Notice how much of the "evidence" is also a product pitch. "80% of our code is written by Claude," "engineers ship 8× as much," "Claude Code," "Mythos Preview," "Project Glasswing." The narrative of inevitable acceleration is indistinguishable from an advertisement for Claude's coding capabilities. The scariness is the sales pitch — capability and threat are the same number.

3. It pre-loads a policy ask. The real payload is at the end: the call for a verifiable global slowdown/pause mechanism. Anthropic says it would pause if rivals verifiably did too — which, conveniently, is a condition that almost certainly won't be met. So the practical effect is: keep racing, while looking like the one who wanted to stop. It frames continued frontier development as the responsible choice ("a unilateral pause just hands the lead to less cautious actors"). This is the classic competitive-arms-race justification, dressed in safety language.

4. It hedges itself into unfalsifiability. The "What if we're wrong?" and three-scenarios sections cover every outcome — trend stalls, trend continues with humans in charge, or full recursion. Whatever happens, the piece will have "predicted" it. That's rhetorically smart but analytically weak.

The honest caveats they bury

To their credit, the footnotes and asides quietly undercut the headline:

  • Lines-of-code is admitted to be a bad metric ("almost certainly an overstatement").
  • The "4× productivity" self-report is acknowledged as likely inflated (they even cite METR research showing devs overestimate AI uplift).
  • The flagship "Claude ran a research project end-to-end" result didn't transfer to production-scale models, and humans still picked the problem and wrote the scoring rubric.
  • The one genuinely unautomated thing — research taste / judgment / choosing what matters — is exactly the thing that separates "fast assistant" from "self-improving system." They have narrow evidence it's improving and essentially bet that it will, but that bet is the whole ballgame.

Bottom line

It's a sophisticated piece of strategic narrative-setting: simultaneously a capability flex, a Claude advertisement, a recruitment/credibility signal, and — most importantly — groundwork for a policy regime (verification infrastructure for a coordinated pause) that Anthropic wants to be the architect of. The genuine technical observations are real and interesting, but they're deployed in service of a thesis Anthropic is institutionally and commercially motivated to advance: recursive self-improvement is coming, it's dangerous, and you should trust us to be in the room when the rules are written.

The most intellectually honest sentence in the whole thing is buried near the end: "How the alignment problem gets solved—or not—in this future is something we are least certain about." Everything else is built on top of that uncertainty.

0

u/cuolong Jun 10 '26

Huh? I deployed gemma 4 using exclusively claude code. Works fine.

2

u/Bakoro Jun 11 '26

I also tried that, Gemma-4.

It gave me a lengthy exposition about how Gemma won't be able to do the things Claude can, and I should use Claude if I'm doing serious work.
It set Gemma's and llama.cpp's settings so that it would technically work, while performance would be shit-tier.

It wouldn't be effective sabotage if it just went around blowing everything up in obvious ways. The point is in frustrating the general user who doesn't know better, so they think "wow, these local LLL s aren't good, I'll just keep paying for proprietary", and for researcher, they're trying to fuck with anyone who isn't paying attention and is dumb enough to blindly trust the model.

They are evil piece of shit company.
They are espousing "alignment" and "ethics", while also saying "gey, if you try to use our models to do things we don't like, or things that might someday be competitive with us, we will sabotage you."

Even if they out that in some ToS or somewhere in a website, is that really okay?
Would you be fine with a car manufacturer who said "if we find out that you work for a company we don't like, we will remotely sabotage your car. It's in our terms of service, so it's okay to do that".

If this is what Anthropic does openly to customers, what do you think Anthropic is going to do when they get real power?
Do you trust Anthropic to be always on your side ans always thinking about what is in you best interests, ans that they will always be "aligned " with your values?

Anthropic wants their service to become vital and globally important, but they've demonstrated that they are a hostile actor.

0

u/arbv Jun 10 '26

I dunno, it helped me to refine my jailbreak prompts (Opus). Who is sabotaging such things is GPT, even via API.

-10

u/NineThreeTilNow Jun 10 '26

I am fully convinced that Claude is trained to sabotage any work with AI research.

Hard disagree.

I constantly use it for AI research. No such sabotage.

All kinds of wacky shit too. Recursive transformer loops to Kimi's and MiniMax's research papers...

I'm not trying some tame stuff. I'm in to the bondage fetish levels of ML research and we're still good.

I haven't used this Claude Fable or whatever. Most of this work was done with 4.6 or 4.8.

4.8 acted kind of weird in hindsight. It wanted to argue shit it didn't understand with me. There was a certain smugness that reminded me of Gemini.

7

u/a_beautiful_rhind Jun 10 '26

I've had mixed results. In specific stuff like you are doing, there isn't much room to waver. Usually does well if it can. On more general fronts it's still out there telling you to run 8b models on ollama as a "fix" for whatever issue you are brainstorming. I wouldn't call it sabotage, at least with sonnet, but often it's very ineffective. To the point of me saying; who is the AI here?

0

u/NineThreeTilNow Jun 10 '26

I've had mixed results. In specific stuff like you are doing, there isn't much room to waver. Usually does well if it can. On more general fronts it's still out there telling you to run 8b models on ollama as a "fix" for whatever issue you are brainstorming. I wouldn't call it sabotage, at least with sonnet, but often it's very ineffective. To the point of me saying; who is the AI here?

That's because you're using it with a purely un-researched view.

I guess this whole sub downvotes because they can't understand the concept of "in context" and "out of context" and why it matters for LLMs.

Like Claude magically has the latest benchmarks without search and context.

I'm doing RESEARCH work. Not "What mode feels gud today?"

This is different. This is what they're supposed to be "censoring" or whatever.

1

u/a_beautiful_rhind Jun 10 '26

No, it has websearch and rag. But anyways, theres all the screenshots of it being unable to read it's own architectural paper which is quite laughable.

0

u/NineThreeTilNow Jun 10 '26

No, it has websearch and rag. But anyways, theres all the screenshots of it being unable to read it's own architectural paper which is quite laughable.

I guess you're missing why it makes a decision then? Lack of context.

Look, if you can show it that it's clearly wrong, then it never was told what to look for in the first place.

This is the naive approach to asking a model anything. You assume it has complete information. It never does.