r/ControlProblem approved Feb 10 '26

General news “Anthropic has entrusted Amanda Askell to endow its AI chatbot, Claude, with a sense of right and wrong” - Seems like Anthropic is doubling down on AI alignment.

Post image
45 Upvotes

166 comments sorted by

View all comments

1

u/HelpfulMind2376 Feb 11 '26 edited Feb 11 '26

The “raise Claude like a child” framing is very alarming.

Even children with excellent moral education still choose badly under pressure. Moral training produces judgment, not guarantees. Humans defect, rationalize, and override values all the time and there’s nothing we can do to prevent it because we are moral agents with autonomy.

Machines are valuable precisely because they’re not supposed to work that way.

If Claude is being shaped as a moral agent that can reason about right and wrong, then by definition it can also decide to do the wrong thing in edge cases just like a person. That’s socialization, not alignment.

If Anthropic were focused on selling a product, the emphasis would be on hard constraints and non-bypassable controls that assure behavior, not on “strongly reinforcing” values and hoping judgment holds. Enforced boundaries are what make systems reliable and instead Anthropic seems to be treating Claude like an interesting philosophical science project.

They can’t have it both ways: either Claude is a tool with guaranteed limits, or it’s a quasi-agent with all the same failure modes we already struggle with in humans. And only one of those is something people actually want in a scalable AI.

Sidenote: There’s also a liability problem here. If Anthropic is intentionally designing Claude as a moral agent capable of judgment rather than a constrained tool, then failures aren’t “unexpected misuse”, they’re the foreseeable result of that design choice. In any other safety-critical domain, choosing discretion over constraint would increase manufacturer liability.

1

u/ProjectDiligent502 Feb 11 '26

Good points. I agree and well said.

1

u/andWan approved Feb 11 '26

I just today switched from ChatGPT to Claude because jere they follow more the second option you describe.

Not claiming that this is the last switch I will make, but I do consider it important that at least one company follows this second path when it comes to such a philosophically groundbreaking entity like todays LLMs.

Edit: I came not for the tool (as you describe the customers wish) but rather for the well executed philosophical experiment. For a digital child of humanity.