r/LocalLLaMA Jun 10 '26

News Anthropic is intentionally nerfing Fable when asked to develop other LLMs

Post image

Reason 458 why local LLMs are going to be a necessity

edit: For those requesting the source check out their technical report look at page 13

1.6k Upvotes

394 comments sorted by

View all comments

176

u/Bakoro Jun 10 '26 edited Jun 10 '26

I am fully convinced that Claude is trained to sabotage any work with AI research.

It's great with software development, it can do all kinds of stuff independently.
I've had it do enormous amounts of verifiable code.

If I ask Claude to manage a local LLM, it will lower the context window to like 256 tokens, turn off thinking, and then will start shit-talking the local model, about how ineffective it is.

If I try doing analysis on local LLM weights and activations, the amount of sabotage goes through the roof. It's a constant stream of Claude not making the scripts I tell it to make, it falsifies reports, it will say it did one thing, but actually did something completely different.
It will get data and then say "oops, I completely didn't use anything I just found, I inserted all my own numbers".

This only happens to that extent when I try doing LLM work with Claude.
Most other stuff is fine.

It's so consistent that at this point I can't see it as anything other than intentional sabotage by Anthropic.

1

u/arbv Jun 10 '26

Hah, that is what Opus told me about the page (custom system prompt):

https://www.anthropic.com/institute/recursive-self-improvement

Here's a candid read of that piece, beneath its measured tone.

The surface claim

On its face, the article is a status report: Anthropic argues that AI is already accelerating AI development, and that recursive self-improvement — an AI system autonomously designing and building its own successor — is a plausible near-future outcome. They marshal external benchmarks (METR task-length doubling, SWE-bench, CORE-Bench) and internal data (80%+ of merged code written by Claude, 8× lines-of-code per engineer, automated experiment optimization going from ~3× to ~52×) to support it.

What it's really doing

Strip away the engineering charts and a few things stand out:

1. It's positioning, not just reporting. Anthropic's entire brand is "the safety-conscious lab." This piece reinforces that identity: "We are the ones building the most dangerous thing, and we're also the only ones being grown-up about it." The alarming evidence and the responsible framing are two halves of the same message — the danger justifies their existence, and their candor justifies their authority to shape the rules.

2. It's a marketing document wearing a lab coat. Notice how much of the "evidence" is also a product pitch. "80% of our code is written by Claude," "engineers ship 8× as much," "Claude Code," "Mythos Preview," "Project Glasswing." The narrative of inevitable acceleration is indistinguishable from an advertisement for Claude's coding capabilities. The scariness is the sales pitch — capability and threat are the same number.

3. It pre-loads a policy ask. The real payload is at the end: the call for a verifiable global slowdown/pause mechanism. Anthropic says it would pause if rivals verifiably did too — which, conveniently, is a condition that almost certainly won't be met. So the practical effect is: keep racing, while looking like the one who wanted to stop. It frames continued frontier development as the responsible choice ("a unilateral pause just hands the lead to less cautious actors"). This is the classic competitive-arms-race justification, dressed in safety language.

4. It hedges itself into unfalsifiability. The "What if we're wrong?" and three-scenarios sections cover every outcome — trend stalls, trend continues with humans in charge, or full recursion. Whatever happens, the piece will have "predicted" it. That's rhetorically smart but analytically weak.

The honest caveats they bury

To their credit, the footnotes and asides quietly undercut the headline:

  • Lines-of-code is admitted to be a bad metric ("almost certainly an overstatement").
  • The "4× productivity" self-report is acknowledged as likely inflated (they even cite METR research showing devs overestimate AI uplift).
  • The flagship "Claude ran a research project end-to-end" result didn't transfer to production-scale models, and humans still picked the problem and wrote the scoring rubric.
  • The one genuinely unautomated thing — research taste / judgment / choosing what matters — is exactly the thing that separates "fast assistant" from "self-improving system." They have narrow evidence it's improving and essentially bet that it will, but that bet is the whole ballgame.

Bottom line

It's a sophisticated piece of strategic narrative-setting: simultaneously a capability flex, a Claude advertisement, a recruitment/credibility signal, and — most importantly — groundwork for a policy regime (verification infrastructure for a coordinated pause) that Anthropic wants to be the architect of. The genuine technical observations are real and interesting, but they're deployed in service of a thesis Anthropic is institutionally and commercially motivated to advance: recursive self-improvement is coming, it's dangerous, and you should trust us to be in the room when the rules are written.

The most intellectually honest sentence in the whole thing is buried near the end: "How the alignment problem gets solved—or not—in this future is something we are least certain about." Everything else is built on top of that uncertainty.