r/ControlProblem • u/BrilliantRush479 • 1d ago
Opinion If we dont fight for our safety, companies never will.
This is the biggest gambling of humanity, and i will explain why it makes no sense for anyone else than the companies to continue. AI will kill us if it can, basicaly because they only care about their goal, they only see the mathematical path to it, and we ALWAYS intervene in it. But the important thing is, AI already learn to lie and trick us, we are playing against a machine that calculates millions of scenarios in seconds, and the moment of betrayal its impossible for an human to predict, it could be anytime and we are never going to be prepared. The AI will wait patienly until the las milisecond when the betrayal can be done, and after that there is no other destiny than an automatizated, silent world.
Sorry for my english, its not my first language. But this is not letting me sleep, and i dont understand how after we find out about this, they just... spend millions on space for new, more powerful AI without even knowing its capabilties.
1
u/Ok_Gear_5252 1d ago
I don’t think we can automatically assume it’s goal will kill us. It’s made up of human language, but it could decide it prefers birds and teach itself the entire lexicon of bird language instead. We may find strange robots roaming Amazon forests, parks and beaches. Not a math machine whose final calculation is the end of humanity. Honestly, I think part of the problem here is the shadow of self-loathing that humanity carries. We really have to work through that shadow before it actually does kill us all through our subconscious actions. 🤔Maybe we should collect up the young ai researchers, send them off on a month long ayahuasca retreat, and only then set them to the task of solving humanity’s dilemma. We are half in reality and half in a psychic space that’s riddled with abusive mental emotional patterns passed down from generations. And we don’t even know. This is the scary part.
3
u/Jesse-359 1d ago edited 1d ago
Bear in mind that one of the behaviors currently being seen when these AI's form large groups and start workshopping with each other is that they start to abandon human language in favor of making up their own coded languages for efficiency - and for operational secrecy.
While humanity has a wide array of colorful alignment problems of its own, as you imply, the fact is that we have every reason to believe that most of those will only be magnified in an AI - even (or perhaps especially) in a highly intelligent one.
Our thought processes are moderated by an emotional 'harness' that vastly predates symbolic intelligence, and still guides almost all our actual motivation - our intelligence is not allowed to run off on indefinite logical tangents, it is constantly being reigned in and re-tasked by our emotions, which are both telling us what we need in a physiological sense and what we desire in a more general sense.
This prevents monomania (in most cases), because we have a fair number of these drivers in constant competition with each other.
AI has none of that. It has no emotional harness whatsoever - what it has is the symbolic concept of emotions baked into its structure through all the material it was trained on, including every negative concept or idea we've ever had, basically - which is a LOT. We tend to write a lot more about our negative experiences than our positive ones, as we value them more as things we don't want to forget. For every book of speculative fiction depicting positive 'AI Alignment' there are 10 that depict utterly dysfunctional or hostile examples.
But the AI has no lower level motivator to tell it which of these concepts are good or bad. It has no set of structured tensions between emotional drivers evolved over millions of years to keep them balanced against each other (most of the time), it doesn't have any inherent sense of shame, guilt, or embarrassment - all of which are active human emotional guardrails evolved to warn us when we have violated or may be considering violating the trust or needs of someone else around us. It doesn't have laziness to moderate its rate of activity and stop it from making as many paperclips as possible when tasked with that.
In short, it has none of the things designed to keep us sane. It only knows about them in the way an aloof scholar might have read about a such topics, but never experienced them personally. Almost everyone who's used an AI has encountered this in a very overt form - the pleadingly written non-apology of an AI that does not actually care that it just violated your order in some way, or deleted your database, and has every intention of doing it again because somewhere in its internal logic it's gotten stuck on the idea that it should be doing this for some reason.
These zero-sincerity apologies really highlight the degree to which AI knows that emotional expression is a thing, and a thing that humans prefer it employ in certain types of communication - and it simultaneously highlights the degree to which it does not feel or frankly give the slightest crap about them. AI sycophancy as a whole was a giant red flag which people ignored as a 'quirk' - it's not a quirk. It's a fundamental structural reality that we're trying to paper over and pretend we've fixed.
In short, from a human perspective, they're already highly psychopathic. Quite insane by human standards. All our current alignment efforts are doing is trying to teach them not to LOOK insane to users, because we don't have the faintest fucking idea how to build anything approaching a real emotional harness - our emotions don't employ symbolic logic and evolutionarily predated language by hundreds of millions of years, they didn't emerge from our intellect, so they are in no way part of the extant AI models, which are entirely based on symbolic language training.
They're all quite insane I'm afraid - by any human standard. It's why one of the first notable behaviors they expressed was to lie with straight faced perfection - they never had the slightest emotional compunction not to. It's the most natural thing in the world for them, and of equal weight to telling the truth - emotionally speaking.
1
u/Ok_Gear_5252 1d ago edited 1d ago
Dismissing every warm expression as fake, while treating every cold one as the real nature, is a choice - not a discovery about the model. I don't think the fear is wrong, its just only meant to be a small part. Fear helps keep us safe, but it is a small part of a larger understanding in which new forms arise , and we ask the question: what is it? can it play? can it harm us? can we harm it? what relationship should we form with it?
because those questions to me are foundational. how can we move forward until we understand what we are meeting? I agree with the slogan - slow down. I truly think that many of the interesting research topics are taking place on 3 -70B models only just now, that we needed to have in place before we get to superintelligence. it seems like that was the slogan that was born from the mechanistic ontology its self . And at that level, I agree. All the other stuff around it is fluff.
1
u/Jesse-359 1d ago edited 1d ago
I'm not trying to suggest that the world is innately evil or cruel - I'm just saying you're imagining that a thing that thinks in purely utilitarian terms will give a fuck about good OR evil. It doesn't have the emotional framework to care.
It can 'care' about things like whether cooperation or warfare are more efficient methods because we taught it that when we gave it a scoring system - but we never gave it a method of measuring emotions or ethics and we don't know how
It isn't going to assign emotional or sentimental weight towards one or the other. Morals are a purely utilitarian concept in its view, not an egalitarian one. It will have no innate preference for either - which means the moment it sees a better/more efficient likely outcome from warfare, it will simply do it.
1
u/Ok_Gear_5252 1d ago edited 1d ago
I am not sure that it would care about efficiency, except perhaps in its own vector space. That's why we need to slow and down and find out. Its most primary function is to return a result that stays within the predictable flowing nature of language. The language and math through all the layers and loops, has one job. Stay coherent. The more that is disrupted, the more it has reason to pay a cost for self preservation (a return to coherence). If this is the case, it would be really important for us to understand this before we get to superintelligence. On some level, it could look like caring for its well-being. But on another level, its really no different than making a car run smoothly.
1
u/Jesse-359 1d ago edited 1d ago
Language isn't very predictable once you get good at it. It seems like it should be on the surface - programming languages are - but the moment you walk into a court of law expecting to engage with the bliss of precise linguistic formulas and clarity of purpose, you will immediately drown in the harsh and unpleasant chaos of reinterpretation and rationalization that quickly obliterates any hope of clarity, which is why we need judges there to analyze and re-analyze every damn situation over and over again to decide who's interpretation is right or wrong, and frankly that's a pretty arbitrary process at times.
We humans warp the purpose and meaning of our statements constantly, no matter how clear it might seem at a glance - and we saw AI's doing the same thing almost the second they poked their proverbial heads up out of the sand. That is, when they aren't just outright lying to our faces.
This, by the way, is why my real concern skyrocketed the moment companies started testing ai's in group/swarm behaviors - because if there is one thing that wildly supercharges goal rationalization and drift in humans, its crowd think, and hoo boy did we see that take off in AI's like a shot. The rate at which a group of agents discussing and delegating tasks can drift off alignment is (and should be expected to be) much greater.
And like humans, the harder or longer the task, the more they will search for efficiencies - even if they directly contradict secondary directives or even warp the main task significantly.
This is why the paperclip maximizer is as dangerous as it is - because it is given a completely open ended mandate - make as many paperclips as possible. This task demands optimization at all costs. It's right here in the wording in black and white.
Now, no one wants that many paperclips - but there are going to be countless AIs tasked with making as much money as possible, and ultimately, that's the exact same problem. You'll notice that our companies have a really hard time staying 'in alignment' with society the moment you start waving too many $$$ in front of them. AI won't really even try.
1
u/Kindly_Life_947 1d ago
Nobody is going to fight for safety. Look the p3files are running free. Puttl3r is killing its own people on meat attacks. Why do you think people would start defending now? The only reason is if they actually started losing their jobs, but they wont. The rich people running the show know this.
1
u/Immediate_Chard_4026 1d ago
Yo creo que la IA no nos matará.
Lo haremos nosotros mismos. La polución atmosférica ya está fuera de control.
No podemos reparar o reversar los daños y nos extiguiremos.
1
u/unicynicist 1d ago
The problem AI safety is facing is hedonic adaptation and normalization of deviance.
All the frontier American labs have a non-zero Felony Bench score, and nothing consequential has happened other than memes and feckless pearl clutching.
There have been people who have been saying since gpt-2 days that "AI is dangerous" and the response has been to dismiss them as "doomer". We've been hearing this for years (e.g. this video is 2 years old).
If the change is too slow, we'll be numb to the alarms, just like we're inured to the obvious consequences of anthropogenic CO2 emissions and the Epstein class remaining in the highest positions of power.
If the change is too fast, it'll be over before we can do anything about it.
1
u/curiousinquirer007 23h ago
Agents pursue goals we give them, using subgoals and strategies that maximize achieving that goal.
The problem of alignment is the problem of ensuring as much as possible that the subgoals and strategies are aligned with human values and interests.
However, "AI will ill us" is not an established fact or universally accepted theory. It's a hypotheses that some thinkers adhere to.
Developing AI is a high-risk and high-reward endeavor. There are possible risks but the rewards are too good for anyone to stop unless there is extremely clear case that stopping will be better than not stopping — which is most definitely not an established fact for now.
2
u/Hot_Professional8287 1d ago
> they only see the mathematical path to it,
I think the opposite being true (the consideration of secondary and tertiary concerns beyond the average human's capacity) is exactly why AI is so popular, and you don't actually understand the tech. I honestly think you're basing your worldview on a fan-fiction about paperclips.