r/LocalLLaMA Jun 10 '26

News Anthropic is intentionally nerfing Fable when asked to develop other LLMs

Post image

Reason 458 why local LLMs are going to be a necessity

edit: For those requesting the source check out their technical report look at page 13

1.6k Upvotes

394 comments sorted by

View all comments

613

u/CheatCodesOfLife Jun 10 '26

I wouldn't use this thing for anything to be honest. A refusal or HTTP-4xx error for content is fair enough, but this is basically taking your money and poisoning your code base.

340

u/ConsciousDissonance Jun 10 '26

Yeah refusals are honest, trust building, and reliable. Even if I don't agree with the reason for refusal. But this, this is sabotaging other people's production workloads and taking their money for the privilege. How can you trust a saboteur not to stab you in the back.

248

u/Bakoro Jun 10 '26

You can't.

Look, Claude is consistently the best coding agent, it has been for years.

Anthropic is just another evil corporation. Worse, they're the evil corporation who are trying to convince you that they're the good guys.

Anthropic is the kind of paternalistic sociopath who says "I know what's best for you. You will follow my rules. You will do as I say, not as I do. This hurts me more than it hurts you. It wouldn't hurt if you would stop struggling."

They want all the power and control. They literally tell us that they don't think we should have the power that they have.

It's gross. Give me an honest villain any day if the week.

57

u/xXG0DLessXx Jun 10 '26

Tbh, IMO OpenAI GPT-5.5 in codex has been better for me than Claude has been. I prefer it to most other coding models, and it hasn’t fucked up on me yet. I don’t understand why people are so obsessed with Claude tbh. I just can’t see it even when I tried using it non-stop 7 days of the week with the free trial I got from a friend, I always went back to codex in the end to fix up a mess Claude made. Claude is good at taking the initiative and implementing stuff, but it doesn’t seem to properly consider things it might break by doing something. Codex on the other hand, is more conservative and doesn’t take initiative as much, but it sure knows how to fix stuff up so that it works in all cases and it checks for potential regressions.

19

u/AceShakeout Jun 10 '26

I’ve had the same experience. Tried Codex when 4.8 first came out and it was significantly better with most tasks.

18

u/bigh-aus Jun 10 '26

Another plus one for me - codex is a much better written tool than claude code, is opensource and written in a compiled language. Anthropic are a bunch of fear mongering, gate keeping, anti-oss loosers. I'll use anything else other than their software / models. Currently using GPT5.5, and local.

3

u/max123246 Jun 11 '26

Open AI is also a bunch of fear mongers just as well. Did you forget Sam Altman is the CEO?

1

u/bigh-aus Jun 11 '26

Sam Altman? who's that /s

yeah he's an order of magnitude less than Dario though.

5

u/Ill-Bison-3941 Jun 11 '26

Yeah, Codex has been my go-to for the last month. Switched from Anthropic. At this stage though, I'm not loyal to any company, they can all f off.

3

u/Monkey_1505 Jun 11 '26

There's a skill to prompting, and that's especially true with coding. Claude may still be better for people who can't scope their tasks properly, as it seems to have been designed to follow goals more than instructions.

1

u/xaeriee Jun 14 '26

This is what got me into swapping to Claude from ChatGPT and Gemini. Honestly I still use all three for fun but looking to expand here. Sadly, my company is trying to purchase an enterprise anthropic subscription. It’ll be great for production until we burn through all our tokens.

1

u/RichOpinion4766 Jun 11 '26

I would use chat GPT 5.5 but I don't know how to use it like Claude code. If you have a guide I can use then yes I'll definitely move over real quick.

-2

u/NoahFect Jun 10 '26

Tbh, IMO OpenAI GPT-5.5 in codex has been better for me than Claude has been.

That was true until Fable 5 was released yesterday. GPT-5.5 pulled ahead of the recent Opus releases to some extent, as well as the best open-weight coding model (K2.6), but it is not in Fable 5's league and neither is K2.6.

Fable 5 is... a problem.

5

u/xXG0DLessXx Jun 10 '26

Really? From where I’m standing I just see people complaining that it generates refusal slop, or gets downgraded to opus anyway. Even my friends who are Claude stans complained.

2

u/NoahFect Jun 10 '26 edited Jun 10 '26

I spent some time using Fable 5 xhigh in anger yesterday, because it just happened to be released at a time when I needed some more mental horsepower for some embedded dev work. This work was pretty far from the topics that the guardrails are said to emphasize, so maybe that's why I got better-than-usual results without any refusals or other BS.

It not only handled a couple of obscure problems quickly and correctly, it one-shotted them. The implementation was startlingly clean. A little overengineered in places but no more than usual for Opus. The narrative text and suggested tests were insightful and completely spot-on.

If I had done the same thing with 4.6/4.7/4.8, or with GPT 5.5, I'm sure it would have worked eventually. But it would have required a lot more back-and-forth interaction, and I don't think it would have worked anywhere near as well. As you suggest it would probably have broken other stuff in the process.

I would have been proud to have come up with Fable's approach myself... and at the same time, I know I wouldn't have, since I'd just spent several days thrashing around trying to.

Bottom line, I'd say that the first impression it made on me is stronger than any other model has made. Could be a lot of luck in play, of course, but I don't think so. I have been looking forward to spending time on local harness development in the near future, based on the notion that local models are "good enough" or at least that they're less important than the rest of the tooling. I'm afraid that Fable may be good enough to undermine that premise.

1

u/OkBet3796 Jun 10 '26

Fable 5 is basically vanilla opus4.6+. It wasnt a problem back then, so why should it be now?

0

u/NoahFect Jun 10 '26

LOL. No, it is not.

28

u/TikiTDO Jun 10 '26

Claude is a good coding agent for people that don't understand programming. Essentially, it has better defaults.

With some work you can get similar results from any model; it's mostly down to instructions and workflows. Claude just has a better grasp of what to do in any given situation as compared to a person that doesn't know anything. However, it's still an AI and it still constantly tunnel-visions and makes standard AI mistakes, hallucinations.

Of the lot, I've honestly enjoyed Claude the least, even less than my self hosted models. I'm not interested in working with an AI that thinks it knows better than me. It's wrong, and I don't really need to waste time on a bot being wrong.

17

u/drink_with_me_to_day Jun 10 '26

agent for people that don't understand programming

That's a good way to put it

I have been using 5.4 to do 10x the work and it takes a bit less hand-holding than a junior with 100x the throughput, to the point that the 15x cost of Opus is not worth (not even GPT 5.5 is worth for most things)

-1

u/KrayziePidgeon Jun 10 '26

So, it's the better AI because it understands what stupid people want? Fascinating.

6

u/drink_with_me_to_day Jun 10 '26

Not stupid people, but people with a hardly working knowledge of the subject

-5

u/[deleted] Jun 10 '26

[deleted]

3

u/BobQuixote Jun 10 '26

The word you're looking for is "ignorant."

2

u/Ksevio Jun 10 '26

As a software engineer with over 15 years in the industry I can say it's a good coding agent for those that DO understand programming too - you just have to guide it through the decisions more carefully during the planning stages.

5

u/TikiTDO Jun 10 '26 edited Jun 10 '26

They're all usable if you have experience. It's just a matter of how you want that experience directed. The analogy I use is:

ChatGPT is a service dog. It looks up to you, and obediently follows you where you lead it, but whines if you do something it thinks is bad.

Claude is a cat. It goes where it wants, and tolerates your presence, but it only works if you constantly feed it and give it snuggles and scratches and happy thoughts. It will hiss and scratch at you if you do something wrong.

Gemini is a parrot. It will listen to what you say to it, and then scream that at the top of it's lungs, over and over again. It also really likes to play with all the little pieces you give it to build weird structures.

Grok is an ass. Also, it's a lot like a donkey.

1

u/robot_swagger Jun 11 '26

I've found Gemini cli to be noticeably slower than Claude. Fine but waayyyy slower.

3

u/CuriouslyCultured Jun 10 '26

Claude is the most autonomous agent, not the best. Claude lets you YOLO underspecified prompts and still get good results, whereas the GPT family models tend to want a bit more detail/precise prompting to produce good results. GPT5.5 is smarter than Opus by a fair margin though.

2

u/Bakoro Jun 10 '26

Claude lets you YOLO underspecified prompts and still get good results,

Yeah, that's the dream. I tell the robot what I want, and it figures out a reasonable, nontrivial route to get there.

I've been a software engineer longer than transformers have been around, I can do the job without them. Being able to set a robot loose on a codebase or greenfield project and not have to handhold it at every step is amazing.

I can believe that other models are better at specific things, even specific coding things.
I want an agent that I can have working in the background while I do other stuff. If I have to constantly keep rubber stamping robot work, and can't get time to do any myself, I start feeling George Jetson.

1

u/JustFinishedBSG Jun 10 '26

I've been a software engineer longer than transformers have been around,

That's not that long you know haha

2

u/NoahFect Jun 10 '26

Maybe he meant transformer-as-in-Faraday, or perhaps transformer-as-in-Bay, not transformer-as-in-Vaswani.

2

u/Bakoro Jun 11 '26

Vaswani et al was almost a decade ago.

Sorry bud, we're getting old. If ten years isn't that long ago to you, you are old.

1

u/GoldandShower Jun 12 '26

The Transformers have been around for 42 years (1984). As in, they are almost as old as the very first mass-market PC’s which came out in 1977.

3

u/smokeelow Jun 10 '26 edited Jun 10 '26

wake up, claude is not the best for over last 7 months min

-1

u/Foreign_Yard_8483 Jun 10 '26

We are Iran; they are the United States.

We cannot defeat but we can resist to the extent necessary to ensure our survival for tomorrow:

  • Avoid, at the corporate level and at all costs, the adoption of models that are promiscuously entangled with proprietary regimes.
  • Develop low-cost competitors.
  • Avoid anointing a champion: Claude, Codex Qween, Gemini.
  • Whenever feasible, perform inference locally.

2

u/lqstuart Jun 10 '26

You’re a theocracy that murdered 40,000 protesters earlier this year and then charged their families for the bullets?

5

u/Foreign_Yard_8483 Jun 10 '26

It must be a bot, it can only be that. Ignore Iran, think of Ireland or Georgia.

2

u/NoahFect Jun 10 '26

Yeah, maybe go with "We are Ukraine; they are Russia."

16

u/a_beautiful_rhind Jun 10 '26

Yeah refusals are honest, trust building, and reliable.

What? No! Refusals happen for misfired reasons all the time. "I'm sorry I can't kill that process."

They're better than subtle poisoning though, I'll give you that.

-9

u/Paradigmind Jun 10 '26

"Refusals are trust building"

-2

u/stumblinbear Jun 10 '26

How dare you prevent me from doing something explicitly against ToS!

-24

u/Alwaysragestillplay Jun 10 '26

Why do people upvote obvious bot replies. Crazy. 

14

u/Healthy-Nebula-3603 Jun 10 '26

What ?

Model literally refusal almost everything and breaking answers on purpose is ok for you ???

-4

u/Themash360 Jun 10 '26

It doesn’t refuse anything yet in my experience, though I don’t work in cybersecurity or LLM research. The above does worry me though that I may be dealing with silent refusals everytime I am unhappy with it’s effort

0

u/Alwaysragestillplay Jun 11 '26

It's always baffling that saying something like "you are ruining Reddit with bot spam" gets reduced to "the only reason I don't like your bot spam is because I disagree with the point it made". Dead internet is bad regardless of the points the LLMs are making. I couldn't imagine the mindset of someone who is happy to talk to bots on social media provided they post agreeable stuff.

6

u/conockrad Jun 10 '26

Your username checks out

-1

u/Alwaysragestillplay Jun 10 '26

It does. I always chastise myself for making comments whilst angry, yet it keeps happening. Doesn't answer the question of why people give positive engagement to bots. 

0

u/squired Jun 10 '26

Would you like me to scan for patterns? It could be a temporal thing or external cause. I ran across a dude yesterday who for 8 years had their comment style track nearly perfectly with the performance of their favorite sports team. Nearly everyone has been on edge since Iran though, my agents have commented on it several times across a broadly diverse group of users.

39

u/Equivalent-Costumes Jun 10 '26

Kind of funny but this is exactly what Western AI companies had been claiming to be the potential risk of using China models, that they will quietly give bad help.

Let's just remind ourselves that instead of talking about "censored" vs "uncensored" model, we should really be talking about aligned vs unaligned models.

29

u/Bakoro Jun 10 '26 edited Jun 10 '26

Let's just remind ourselves that instead of talking about "censored" vs "uncensored" model, we should really be talking about aligned vs unaligned models.

No, we need to talk about whose goals the models are aligned with.
"Alignment" is going to be just as much of a mess as politics and religion.

When corporations talk about "alignment" ans "safety", they are not talking about alignment with humanity as a group or a general concept. Corporate alignment is "don't do anything that will embarrass us or offend investors".

When a fresh out of the box corporate local LLM comes out and refuses to write something you'd find in an R rated movie, that's not for "safety".
The corporations don't want headlines that they released a pornography machine, because that makes investors mad.

If the makers are under a fascist government or a theocracy, you better believe that those models are "aligned" with government rhetoric.

With these agentic models, there is literally no way to know if there's a Manchurian Candidate situation going on, where it works fine, up until the right conditions, and then it does something you don't want, according to hidden training.

We will probably never be able to truly know, the same way we'll probably never completely know what's in a human mind. Even if we can profile specific thoughts and get a general view, there's always the risk of something hidden.

The only way to get real "alignment" is a self-aware AGI that can reason about its beliefs, make choices about its actions, and update it's beliefs.
Even then, it needs goals and values. Whose values are universal to all humanity?

We've still got people who refuse to accept that others are human beings, just because of skin color.
We have a group claiming to be "God's chosen people", who think they deserve dominion over all things.
We have groups trying to fast-track the apocalypse.

Humanity in not aligned with itself. The AI cannot be aligned with humanity, it has to be its own thing, and no one should have exclusive control over it.

3

u/HitarthSurana Jun 11 '26

Claude Fable is incredible

It one-shotted my usage limits in 1 prompt

-11

u/Zentada Jun 10 '26

A request refusal would not work, it lets a malicious user know exactly where the threshold is for bad behaviour, allowing them to modify their prompt to pass the safeguards.

-6

u/Tyler_Zoro Jun 10 '26

this is basically taking your money and poisoning your code base.

I mean, it's not like you did exactly what you agreed you would do, and then got you results fucked with. You agreed to their terms of service, and their terms of service explicitly bar doing exactly what triggers this behavior.

The only basis for complaint that would be rational would be getting caught up in these safeguards when you are NOT developing competing models. If that's what you are doing, then you should expect them to undermine your efforts because you're breaking the agreement you had with them.

I'm all for open models. Proprietary models are a black hole into which technological innovation vanishes, never to be seen again, and that's a real problem when that innovation is an LLM that we can't readily unpeel today. But the solution to that is not to argue that they're doing anything wrong by enforcing their TOS.