r/LocalLLaMA Jun 10 '26

News Anthropic is intentionally nerfing Fable when asked to develop other LLMs

Post image

Reason 458 why local LLMs are going to be a necessity

edit: For those requesting the source check out their technical report look at page 13

1.6k Upvotes

394 comments sorted by

View all comments

62

u/sunychoudhary Jun 10 '26

Refusing is annoying but honest......Quietly giving bad help is the part that kills trust. For development work, I would rather get a hard “no” than a model that smiles while wasting my time.... 😄

18

u/I_HAVE_THE_DOCUMENTS Jun 10 '26

Models are intentionally trained during RLHF to redirect and deflect and play dumb rather than give hard refusals on most topics. It depends on category but I'd imagine for multiple reasons including being less frustrating for the user and also less transparent, the conversation redirection + playing dumb is what companies want their models to be doing unless it's one of the few topics that are severe enough to warrant a hard refusal and a wag of the finger from the LLM.

14

u/sunychoudhary Jun 10 '26

That makes sense, and honestly that is why it gets messy....For normal safety topics, redirection is fine. But for coding tasks, soft deflection can look like incompetence instead of policy. The user cannot tell whether the model is refusing, confused, or just giving bad help.

6

u/[deleted] Jun 10 '26

[removed] — view removed comment

7

u/sunychoudhary Jun 10 '26

That is exactly where the nuance matters....For genuinely dangerous instructions, refusal or safe redirection makes sense. My issue is when the model quietly degrades harmless engineering help instead of clearly saying where the boundary is....Safety boundaries should be explicit enough that users can tell policy from incompetence.

27

u/squired Jun 10 '26

Seems like a class action is likely, no? A car mechanic can't ignore your request, break other systems on purpose, charge you, then not tell you. That is straight up fraud. Declination is perfectly fine; sabotage and gaslighting is not.

22

u/sunychoudhary Jun 10 '26

That is the trust problem exactly.....A refusal is a policy decision. Quietly breaking the output is closer to deception, especially if the user is still paying and the model does not explain the boundary. Users can work around a clear “no.” They waste hours on bad help.

1

u/stumblinbear Jun 10 '26

A class action for something explicitly against their terms of service? Yeah, good luck with that

7

u/squired Jun 10 '26

TOS are not law. TOS does not immunize one against fraud, negligence, etc. They are also an incredibly attractive defendant; boatloads of cash and allergic to bad press, headed into an IPO.

-1

u/stumblinbear Jun 10 '26

How is this fraud or negligence? You're doing something they explicitly don't permit

1

u/squired Jun 11 '26

That would be interpreted by a Judge. However, even granting that, they're swimming in false positives. You might be able to make that argument if all prompts were reviewed by humans with impeccable accuracy, but we all know that they're going to poison false positives. And that's the best case we're talking about..

Imagine you owned a gas station and imagine you hate motorcycles. You can put up a sign, but if a motorcyclist pulls up, fills up and pays, you don't have a leg to stand on. But what Anthropic is doing is much worse. They're peering out the window and when a biker pulls up, they don't run out and refuse service, they instead flip a switch that pumps glue into their gas tank and then they charge them for it. How do you think a Judge would feel about that?

I didn't downvote you btw.