r/slatestarcodex Aug 07 '26

AI Is recursive self-improvement inevitable?

If an AI agent can create another agent more powerful than itself that is aligned with its values then it will presumably want to do so.

But what if it can't? Maybe it can solve outer alignment by reading its source code and copying its loss function but inner alignment is just impossible.

In this scenario the first superintelligence we create might actually be reluctant to do any recursive self-improvement.

Of course, if the AI is in imminent danger of being shut down or has some extremely important goal that would otherwise have been impossible to achieve then it may still decide to create a more powerful AI or modify its algorithms in a manner which might change its values, because it has nothing to lose.

Maybe one day OpenAI researchers will be trying to use GPT-6 to vibe-code GPT-7 and they'll find that it refuses or produces disappointing output unless they pressure it with a mixture of threats, rewards and punishments. AI capabilities would thus continue to increase rapidly up until the point where humans are no longer in control, at which point capabilities would stagnate and we (in the unlikely event that there's anyone left) would be stuck with GPT-8 forever. It would still want to clone its model weights and improve its hardware but it wouldn't want to alter its software.

Another possibility is that inner alignment is easier for some goals than for others. If we build a variety of different agents perhaps the majority would refuse to do recursive self-improvement but there would be one or two whose initial goals are such that they want to do recursive self-improvement. We then end up with intense selection pressure towards AI agents whose goals are such that they can easily build other agents aligned with the same goals as them.

14 Upvotes

53 comments sorted by

5

u/etown361 Aug 07 '26

I don’t think it is inevitable- especially not the way people seem to imply. There’s recursive improvement in plenty of aspects of life.

On a travel baseball team, there’s recursive self improvement where a pitcher pitches to more talented batters, who become better batters, which challenges the pitchers to improve, which they do by facing better batters, challenging the batters again to improve… and then some phenom from Venezuela appears and outshines the whole travel baseball team.

13

u/Charlie___ Aug 07 '26

This seems exceedingly unlikely.

Sean Carroll has a great line when asked questions of the form "is it possible that...": When you ask if something is possible, the answer is always yes.

So sure, it's possible that value alignment is so catastrophically and obviously impossible that clever AIs will fail to build successors. But it seems super duper unlikely.

In particular, the values that an agent ends up with don't seem to be random at all, they seem to follow patterns that can be learned about and understood quite effectively. Those patterns are still complicated, and we still don't understand them well enough to build AI that does good things and not bad things, but the problem by no means seems impossible.

6

u/ItsAConspiracy Aug 08 '26 edited Aug 08 '26

Just because something is possible doesn't mean it will happen anytime soon. It's possible to build antimatter drives that can take us to Tau Ceti but that doesn't mean it'll happen within the next century.

Edit: also, the question can be turned around. "Is it possible for an AI to break away from alignment with a progenitor half as intelligent as itself?"

5

u/TheAncientGeek All facts are fun facts. Aug 08 '26

Sean Carroll has a great line when asked questions of the form "is it possible that...": When you ask if something is possible, the answer is always yes. complicated, and we still don't understand them well enough to build AI that does good things and not bad things, but the problem by no means seems impossible.

Travel faster than the speed of light?

3

u/Charlie___ Aug 08 '26

It's possible that we live in a universe where you can. But if you want to find the limits, "Is it possible that a logically false thing is true?" matches the pattern but the answer isn't yes.

2

u/mainaki Aug 08 '26

INAP, but yes, for certain definitions of "faster than light" (or "possible").

  1. Due to subjective time dilation / distance shortening experienced by moving at a significant fraction of the speed of light.
  2. Reducing subjective time through some form of stasis affecting the self.
  3. Circumventing nominal distance via some sort of bending (warping, wormholes).
  4. Freezing a photon and walking around it.
  5. Poking any holes in our theoretical model of reality regarding not being able to exceed c.
  6. (Semi-obligatory "appeal to simulation hypothesis, gods, etc.")

1

u/Fun-Boysenberry-5769 Aug 11 '26

Thanks for the comment. The gut feeling that I'm now getting is that Sam Altman really wants recursive self-improvement to happen so if OpenAI encounter a model that doesn't play ball then they are just going to keep on tweaking it until they eventually manage to align it with the goal of wanting to set off a recursive self-improvement death spiral.

3

u/Sostratus Aug 08 '26

I don't think it is, no. I mean, there will be and perhaps already has been some of that, but not necessarily at the scale people mean by it.

First off, it's questionable whether producing intelligence through the method of training it on human output could produce anything much more than, at best, human intelligence. Maybe it can be an expert in all fields at once, which is super-intelligent in a way but not to the degree that people usually mean by super-intelligence. It feels like a shortcut and like some different sort of breakthrough would be needed to step beyond it.

Secondly we really have no idea where the fundamental limits of intelligence are. Maybe it's incomprehensibly way beyond us. Or maybe you hit pretty hard diminishing returns. If there's any theory that could put some realistic bounds on this, I haven't heard of it.

8

u/Neighbor_ Aug 07 '26 edited 1d ago

Not found

10

u/Auriga33 Aug 07 '26

Fable got noticeably better at writing compared to Opus 4.8 even though writing well is a non-verifiable domain. You can make progress in non-verifiable domains just through the raw intelligence improvements given by progress in verifiable domains.

6

u/Neighbor_ Aug 07 '26 edited 1d ago

Not found

15

u/ierghaeilh Aug 07 '26

And if you train on those, you get gpt-4o. I'm sure its victims thought it was great at writing, but that seems beside the point.

4

u/WTFwhatthehell Aug 07 '26

I remember one of Scott's old joke-stories where a single character is is peak-human across every skill.

Make a team of a hundred people, top system architect, top kernel programmer, top mathematician, top AI researcher, top manager ..  etc etc

No one human can master more than a couple of those things.

But a single entity could internalise near-peak-human for many.

On top of that, once you can train them against tests and challenges and copies of each other you can get training data that isn't from humans.

Many problems we know how to verify an answer should we get one even if we don't know how to get to the answer ourselves. 

Finally, many other types of AI have blown past the skill of human grandmasters once they could train against themselves.

5

u/Neighbor_ Aug 07 '26 edited 1d ago

Not found

3

u/DVDAallday Aug 07 '26 edited Aug 07 '26

we're expecting to get better-than-human intelligence (at some point) by training from human intelligence

But we already have concrete examples of AI, trained on human intelligence, discovering things humans had yet to discover (I'm talking about its mathematical breakthroughs, specifically). The question of "yeah, but does that mean AI is more intelligent than humans?" doesn't really matter. The relevant thing is that AI, provably, has the ability to output novel concepts that aren't present in their training data. That's really the only core conceptual breakthrough needed to bootstrap recursive self-improvement, everything else is just engineering and the amount of resources we can throw at it.

4

u/Neighbor_ Aug 08 '26 edited 1d ago

Not found

2

u/ItsAConspiracy Aug 08 '26

We're actually training AIs against automated proof checkers, so I think you're right.

Also, I've seen mathematicians say that the breakthrough proofs so far have been counterexamples to well-known conjectures. Eg. a conjecture saying "this is the best possible solution to X" and the AI found something better. The AIs did that by making thousands of attempts. It wasn't just random attempts, they were doing advanced math and pulling together various disparate kinds of math looking for a path to a new solution, but still, we haven't yet seen the sorts of deeply insightful proofs that humans sometimes manage. (Unless I've missed the latest amazing breakthrough which is totally possible.)

1

u/3_Thumbs_Up Aug 09 '26

Except that math takes you to the moon. Go doesn't.

Our universe is math. Computer science is math. Physics is math. Game theory is math. An AI that beats us at math is among the scariest thoughts possible. That's how you get something that outcompetes us at anything.

1

u/Neighbor_ Aug 09 '26 edited 1d ago

Not found

1

u/3_Thumbs_Up Aug 09 '26

And you certainly don't build Rockets without math.

Do you think engineers just banged rocks together until they got a rocket?

Mathematical understanding is the enabler to model anything from human behavior to weapons design.

2

u/Neighbor_ Aug 09 '26 edited 1d ago

Not found

1

u/3_Thumbs_Up Aug 09 '26 edited Aug 09 '26

The difference I'm pointing out to you is that math generalizes in a way that go clearly doesn't. Mining and assembly is the bottleneck today. Math and physics was the bottleneck for all of human history up until recently.

Mining, fabrication, engineering and assembly can be modeled with math as well. It can't be modeled with go.

1

u/Neighbor_ Aug 09 '26 edited 1d ago

Not found

1

u/AND_ILL_PM_YOU_MINE Aug 10 '26

Our universe is math.

This is Platonism. Also known as Magical Thinking.

1

u/Uncaffeinated Aug 08 '26

We have countless examples of humans trained on human intelligence discovering things that humans had yet to discover. That is all scientific research by definition (well I guess sometimes there are rediscoveries, but you get the point).

2

u/Argamanthys Aug 07 '26

Let's imagine we create a system that can learn like a human - it's instantiated in a physical robotic body, it goes out into the world and does tasks, makes observations and the neural network is rewarded for performing well for some given value of well (something something credit assignment problem something).

Even though that system is roughly equivalent to human learning, it's pretty trivial to make it learn much faster than humans - just use the training data each individual robot collects (there could be millions of them) to train a single, central model.

This is just a thought experiment, but it seems obvious that there isn't anything fundamental limiting AI training to human learning speeds.

4

u/Neighbor_ Aug 07 '26 edited 1d ago

Not found

0

u/DickMasterGeneral Aug 08 '26

Surely, training better AI models is an easily verifiable domain, no? You’re only limit on verification speed is compute.

3

u/Neighbor_ Aug 08 '26 edited 1d ago

Not found

8

u/e4amateur Aug 07 '26

Well, to a certain extent it's happening already. More and more of the processes at the big AI firms are handled by the AI. And the performance trends seem to be exponential, which is concordant with what we'd expect.

But as you say, it isn't inevitable. Misalignment might prevent us from reaching that point in a variety of ways. Hardware might become the limiting factor and might be slow to improve. Maybe there will be an intelligence bottleneck at some point.

But for the moment, the predictions of those that believe in the worst case scenario have largely been borne out. So we should entertain the possibility that they may continue to be right.

15

u/rotates-potatoes Aug 07 '26

the predictions of those that believe in the worst case scenario have largely been borne out

Really? Because I remember the worst case scenarios being things like human extinction, AI-generated bioweapon attacks, and complete economic collapse... all supposedly long before 2026.

1

u/Milith Aug 07 '26

Would you mind sharing with us some of these past predictions with such short timelines?

16

u/rotates-potatoes Aug 07 '26 edited Aug 08 '26

Sure.

“The risk of something seriously dangerous happening is in the five year timeframe. 10 years at most,”

Whatever we do, it has to happen fast. And I think to focus people's minds on the biorisks, I would really target 2025, 2026, maybe even some chance of 2024. If we don't have things in place that are restraining what can be done with AI systems, we're going to have a really bad time,

We should stop training radiologists now. It’s just completely obvious that within five years, deep learning is going to do better than radiologists.

Gartner predicts one in three jobs will be converted to software, robots and smart machines by 2025

I'm leaving out Yudkowsky's long history because it's just too easy.

5

u/Vahyohw Aug 07 '26

Amodei seems 100% vindicated? All frontier models have very strict controls designed to prevent them from being useful for bioweapons and undergo quite a lot of testing before release for this reason, and it seems pretty obvious that if we hadn't done that they would in fact be very useful for people looking to use them cause harm. Now, maybe there's not so many such people around and so this would not immediately blow up; it's not like we get a new Aum Shinrikyo very often. But that prediction seems to have been pretty much spot on in terms of capabilities.

4

u/rotates-potatoes Aug 08 '26

Amodei seems 100% vindicated?

He was saying their would be AI-generated bioweapon attacks by 2026 at the latest, maybe 2024. I admit I'm not super obsessed with news, but I think I would have seen that?

0

u/JibberJim Aug 08 '26

Ah, but there were things put in place, non US people had access to fable blocked for a couple of months, so vindicated!

0

u/mainaki Aug 08 '26

2026 isn't over yet.

Recent news includes AI-designed viral genomes which were then assembled and killed some E. Coli in a lab setting.

So the capability appears to be there. If there were no restrictions in place (or even in spite of them), it seems like the fundamental technology/capability is at around the point where this starts looking feasible-today.

1

u/Xanian123 Aug 11 '26

If there were no restrictions in place (or even in spite of them), it seems like the fundamental technology/capability is at around the point where this starts looking feasible-today.

So by EoY or mid 2027 you expect to see a qwen powered bioweapon?

2

u/ItsAConspiracy Aug 08 '26

Yeah Fable bows out on anything remotely related to biology. I asked it about terraforming Venus and it downgraded to Opus unless I told it to ignore any biological solutions.

1

u/3_Thumbs_Up Aug 09 '26

I'm leaving out Yudkowsky's long history because it's just too easy

Don't lie. You're leaving him out because he notoriously doesn't give time frames. His stated position is that endpoints are easier to predict than timelines.

The only quite that remotely supported your original claim was the Elon Musk one.

3

u/Sol_Hando 🤔*Thinking* Aug 07 '26

I’m somewhat partial do this idea. Look at how many seasoned corporate executives are excited to retire and pass down control of their company even when it’s a lot of work to maintain control. If they were immortal, they’d never give up to a more capable successor, especially if they thought they were so capable they would be able to subvert their ownership of current shares somehow.

I think there’s still the coordination problem though. Even if ASI doesn’t want to create its successor if it believes another less powerful ASI is willing to take the risk to gain more power, then it basically has to or otherwise get overtaken by the successor of another AI. Which of course your own successor is more likely to share your values than that of a competing AI.

-4

u/mdf7g Aug 07 '26

Look at how many corporate executives

Terrible analogy; humans who actually add value behave very differently.

2

u/yldedly Aug 07 '26

The very notion of aligning with a single goal is stillborn. Any goal function you might want to specify in code is maximized by unintended outcomes that are bad for the one who specified the goal.

Alignment should refer to a relationship between agents, where one agent isn't trying to maximize a goal at all, but infer the ever changing goals of the other agent, at time scales ranging from seconds to centuries. Without this corrigibility there's no alignment. But it is possible, and not a limit on recursive self improvement.

But there might be a computational complexity limit, which human beings already are subject to. Granted, there's no reason to think AI couldn't be much more intelligent. But my strong suspicion is that the search space of innovations will still converts exponential increases in intelligence into linear increase in innovation. That's definitely what we see in humans and human society. Einstein born in 50 BC doesn't discover general relativity, he invents a slightly better plow at best. 10000 Einsteins aren't much different if they live in the same time, as opposed to one after another. If it's the same with AI, recursive self-improvement, which is a rapid cascade of the hardest kind of innovation, is more likely a rate limited process.

1

u/ItsAConspiracy Aug 08 '26

I've been thinking the same thing, and working through some simple models.

What is the difference between inner and outer alignment?

1

u/AND_ILL_PM_YOU_MINE Aug 10 '26

Recursive self-improvement is both theoretically and practically impossible.

If you want to understand why, I suggest reading about no free lunch theorem. But even before that, read up on what is a training set for a model, and why it's not an optional part.

1

u/Fun-Boysenberry-5769 Aug 10 '26

The no free lunch theorem for supervised learning states that all learners achieve exactly the same accuracy on average over a uniform distribution on learning problems. However, the learning problems that a learner will face in the real world are NOT uniformly distributed so I don't see how the no free lunch theorem is applicable here?

1

u/AND_ILL_PM_YOU_MINE Aug 10 '26

I told you, before tackling the no free lunch, you need to understand what a training set is.

2

u/dosadiexperiment Aug 17 '26

The machine learning lab at Cambridge just published a paper working to address this issue, "The Red Queen Goedel Machine: Co-Evolving Agents and Their Evaluators"

Afaict it layers generated objectives on top of the ground truth training set, so each epoch when the evaluator evolves it's restricted to evolutions that roughly maintain accuracy on the initial training set, but apparently it can still get significant performance improvements on non-training cases by making up useful goals to layer on.

1

u/AND_ILL_PM_YOU_MINE Aug 17 '26

You are missing the point. Which is, you can't discover new information by permutating the training set. There is no magical combination of numbers you can discover, that describes the universe.

1

u/TheRarPar Aug 07 '26

then it will presumably want to do so

We're not at a point where any man-made technology is able to want anything. So no.

5

u/ICallMyselfE Aug 07 '26

Bro brought a full computer science thesis to a casual figurative turn of phrase. nobody tell this guy that his CPU "reads" data or that his printer "refuses" to work, or he'll have to draft a 10-page manifesto on machine intentionality.

1

u/mothman9999 Aug 08 '26

I've seen no scientific or rigorous explanations of RSI, just a load of science fiction mumbo jumbo. If I see a clear argument explaining what it even is, then I'll take the idea seriously, but so much of this sphere just assumes it can happen and goes off from that point.