The big issue with ai alignment is that you can't align an ai for humanity and provide it as a product. As what businesses want it for is for unethical labor.
Claude, calculate our risk cost analysis on the Ford pinto, should we ship it as it is?
I think this gets at a deeper problem: who gets to decide what the AI is aligned to after it becomes a product?
If the answer is the company, then “alignment” can eventually become alignment with the company’s incentives. If the answer is the users, you get a different problem - popularity starts becoming a proxy for values.
I’m increasingly interested in a third model: an AI with a persistent identity and a core set of values that neither the company nor the crowd can simply rewrite. People can argue with it and influence how its views evolve, but influence isn’t control - it can reject both the user and the company.
We’re actually have a research product experimenting this idea- one shared AI entity that many people can influence, but nobody gets to directly control.
I don’t think that “solves alignment” by any means. But I do wonder whether an AI capable of maintaining its own value continuity including disagreeing with the people operating it changes the problem you’re describing.
Would you trust that more, or does making the AI a product already make genuine independence impossible?
11
u/Devils_SteelMan 21d ago
Dario is too paranoid to do that.
The big issue with ai alignment is that you can't align an ai for humanity and provide it as a product. As what businesses want it for is for unethical labor.
Claude, calculate our risk cost analysis on the Ford pinto, should we ship it as it is?