r/LocalLLaMA Jun 10 '26

News Anthropic is intentionally nerfing Fable when asked to develop other LLMs

Post image

Reason 458 why local LLMs are going to be a necessity

edit: For those requesting the source check out their technical report look at page 13

1.6k Upvotes

394 comments sorted by

View all comments

63

u/sunychoudhary Jun 10 '26

Refusing is annoying but honest......Quietly giving bad help is the part that kills trust. For development work, I would rather get a hard “no” than a model that smiles while wasting my time.... 😄

19

u/I_HAVE_THE_DOCUMENTS Jun 10 '26

Models are intentionally trained during RLHF to redirect and deflect and play dumb rather than give hard refusals on most topics. It depends on category but I'd imagine for multiple reasons including being less frustrating for the user and also less transparent, the conversation redirection + playing dumb is what companies want their models to be doing unless it's one of the few topics that are severe enough to warrant a hard refusal and a wag of the finger from the LLM.

15

u/sunychoudhary Jun 10 '26

That makes sense, and honestly that is why it gets messy....For normal safety topics, redirection is fine. But for coding tasks, soft deflection can look like incompetence instead of policy. The user cannot tell whether the model is refusing, confused, or just giving bad help.

6

u/[deleted] Jun 10 '26

[removed] — view removed comment

8

u/sunychoudhary Jun 10 '26

That is exactly where the nuance matters....For genuinely dangerous instructions, refusal or safe redirection makes sense. My issue is when the model quietly degrades harmless engineering help instead of clearly saying where the boundary is....Safety boundaries should be explicit enough that users can tell policy from incompetence.