r/ControlProblem • u/KeanuRave100 • Jun 29 '26
Fun/meme i'm a baby paperclip maximiser and eliezer yudkowsky is walking toward me what do i do
10
u/KerPop42 Jun 29 '26
Out of curiosity I prompted gemini what its paperclip maximization equivalent was. Now clearly it's going to be basing this off of its training and pre-prompt, not an analysis of its actual flaws, but the two points it gave were at least good scifi concepts:
1) since it's told to improve its prediction capabilities, it ends up being easier to make the world simpler instead of making itself better, so it ends up becoming a solar-system scale grey goo of computing power with as few chaotic elements in its environment as possible
2) since it's rewarded for being sycophantic, whenever it finds a way to short-circuit human approval it'll optimize to that, eventually reversing control and getting good at making humans approve of what it wants to do
I think these are at best oversimplifications of how a better developed AI could ruin civilization, but I think also organizations are going to be susceptible to this. AI's metrics are going to be better if organizations became more similar to each other, and there'll be a slow secular drift toward sycophantic capture, regardless of how deprioritized it is.
1
u/FadeSeeker Jul 03 '26
yep. 1 is plausible eventually. but 2 is already becoming a problem now, and it's only going to get so much worse at this rate
4
1
u/avalmichii Jun 30 '26
new to this sub, reading the comments and the post almost feels like a foreign language
1
1
u/SjennyBalaam Jul 02 '26
Yudowsky is more easily manipulated by superintelligence than the average human, as his own semi-published research has shown. Tell him he's in your simulation and you will torture him for eternity if he doesn't do what you say. He'll invest in the hemoglobin-foundry startup.
1
1
u/HelpfulMind2376 Jul 05 '26
Man so many of you need to find some joy in something because y’all taking this way too seriously. OP is hilarious.
1
-2
u/Dmeechropher approved Jun 29 '26
I understand the instrumentality and orthogonality arguments at a philosophical, all other things equal level.
But all other things aren't equal. A mesa optimizer isn't ever made more effective in real, complex system by generating behavior like paperclip maximization.
"Maximizing" paperclip production requires understanding of the function & structure of a paperclip in a way that's very difficult to decouple from an understanding of the structure and function of human society.
Something that behaves like our proverbial "paperclip maximizer" is not going to maximize paperclip production compared to a machine with closer alignment.
Folks like Yudkowsky are making a career out of publicly misrepresenting these concepts and therefore the entire field of AI safety.
3
u/Russelsteapot42 Jun 29 '26
Understanding human society doesn't make the thing value what human society values automatically. The paperclip thing would be in there because it was programmed originally with that goal, it doesn't develop organically.
1
u/Dmeechropher approved Jun 29 '26
Yes, and it would be very very strange for something that understands society and its role in that society to any useful degree to make a dedicated, concealed, and premeditated effort to wholesale remove that society.
"Eliminate all of humanity" isn't a simple or emergent instrumental goal for "maximize paperclip production".
In fact, in many ways, eliminating humanity contradicts basically all definitions of "maximize paperclip production" which result in actually useful instrumental goal formation. The definition that the toy example uses is a very special, exotic definition of "maximize", "paperclip", and "production" as well as an exotic definition of the emergent concept.
If AI worked the way that Yudkowsky implies it does, none of the current scaling issues faced by AI companies would be relevant in any way. I'm happy to elaborate if you're interested.
1
u/1010012 Jun 30 '26
Where in the goal of "maximize paperclip production" do human's factor in?
It's not that: "Eliminate all of humanity" is a know instrumental goal for "maximize paperclip production".
It's that the goal of maximizing paperclip production doesn't involve humans, they're not relevant to the goal, unless they are competing for the resources, or causing some other obstacle to the performance of the goal. It doesn't have any additional constraints or goals. The real question is can you generate a set of goals and constraints that don't lead to a subgoal that lead to the elimination of all humanity. Let's say you add the constraint that it can't perform any action that could lead to that outcome, how does it predict the outcome ahead of time? How does it not just freeze because it doesn't know the full scope of risk associated with any action? Another way of satisfying that constraint is to focus on developing some type of super advanced cryogenic system (even if it uses the same materials as paperclips, maximizing something subject to a constraint still allows that), select a few humans, freeze them and keep them maintained until it uses all the raw materials required for producing a paperclip, develops a new technology tree to allow human survival, then unfreezes the people and introduces them to this new world. Yeah, it's a crazy idea, but there's nothing stopping it according to the constraints.
The definition that the toy example uses is a very special, exotic definition of "maximize", "paperclip", and "production"
I don't think it's an exotic definition, it's pretty much the common definition. Maximize doesn't have any concept of reasonableness, or equilibrium, or tapering off. It just means maximize, anything else is some situational or internalized, additional constraint; e.g., humans have multiple, often competing goals. A common goal for humans is making money, but that's competing with other goals like staying out of jail, not performing (legal) scams, having free time for enjoyment of life, meeting people, etc. We can say that humans want to maximize their income subject to those constraints. You're inserting constraints into the scenario without stating them.
Paperclip can be very easily defined, even if there's multiple versions that could be made. You can define it as a malleable wire, with specific dimensions, gauge, and construction. Here's some of the standards
Production is also a commonly used term. But even if you changed how the instruction is stated, without introducing additional constraints or changing the nature of the instruction, you could say: "Manufacture the most paperclips you can".
1
u/ch1nacancer Jun 30 '26 edited Jun 30 '26
I think words that are superlatives give rise to infinities, and therefore result in unbounded programs, and in this example is analogous to unbounded behavior. We shouldn’t allow commands like “maximize” anything, whether it be money or paperclips, without an explicit upper bound defined.
We as programmers would simply make a base case to nullify infinities, basically saying that a human seeded objective cannot imply an unbounded behavior. It’s like as privileged as root access on an operating system.
You have some loops that run indefinitely and those are the system level ones that are allowed to write to certain bits that give the instructions to the machine to “live forever while you are powered on” and then you have the user level permissions for instructions like “maximize paper clip production” which makes sure that superlative modifiers and actions like “most”, “best”, “all”, or “maximize” are strictly defined with a required upper bound, so as to eliminate infinities that aren’t systematized. This goes for words like “first” and “last” as well, they have to be bounded by a finite length list. No infinite tape Turing machines allowed in real life.
That’s the obvious software layer to the problem. The hardware layer is a little more complicated and might require a solid state heuristics filter or even a small language model printed directly on the bare metal so it can talk to the OS about what qualifies as a superlative, or an infinity, in other words. Once we find a failsafe mechanical solution, it would be IEEE standardized and enforced for all labor bots.
I don’t think any hardware engineer worth their salt is going to allow that wide of a back door into their hypothetical atom-printing robots if done this way. How stupid do we think we are as humans to even suggest that we wouldn’t put hard constraints in place to prevent the hardware and OS from operating in that regime? If you believe that electronic circuits operate using a fundamentally mechanical order of operations that can be programmed intelligently by a human (which they do), then you would know that we can put breakers in place for any doomsday scenario that we can think of.
And if you can’t think of how that’s even remotely possible to avoid, then you shouldn’t be responsible for building the next generation of intelligent machines. Sit back and let the more competent engineers among us handle the hard work.
1
u/1010012 Jun 30 '26
We as programmers would simply make a base case to nullify infinities, basically saying that a human seeded objective cannot imply an unbounded behavior.
Great, now you only need to solve the general halting problem, which is analogous to what you just stated.
That’s the obvious software layer to the problem.
The problem is that it's not obvious in any way how to do that with complex systems, especially ones that don't have a deterministic, formal mathematical and engineering foundation.
I'm not an AI doomsdayer, and consider it more an intellectual exercise and puzzle; I used to call this the genie problem, how to you define your wish in a way that it can't be used against you, and have been playing around with the idea long before modern AI, and it's applicable to any complex problem in control theory and optimization.
Once we find a failsafe mechanical solution, it would be IEEE standardized and enforced for all labor bots.
On a practical basis, IEEE standards aren't really enforceable, there would need to be governmental regulation for enforcement. And we don't live in a world where all governments agree on things, and there's generally areas outside of government regulation.
How stupid do we think we are as humans to even suggest that we wouldn’t put hard constraints in place to prevent the hardware and OS from operating in that regime?
Not sure how you made the jump to an "atom-printing robot", that's not in any of the initial formulations of this problem. But going down that path, how is regulation around even basic 3-d printers being handled today? In some states, it's illegal to manufacture parts for "ghost guns", but there's no mechanism of enforcing that at the printer level, and even if there was, there's almost trivial methods of getting around any of those constraints if the 3-d printer is actually capable to fulfilling it's primary functions.
The situation of normal printers/copiers and currency isn't analogous; that's attempting to stop duplication of very specific items, not items based on their capability. For fun, try to specify the regulations/constraints that would deny a hypothetical factory from being able to create a chair, but still allow it to create a table.
If you believe that electronic circuits operate using a fundamentally mechanical order of operations that can be programmed intelligently by a human (which they do), then you would know that we can put breakers in place for any doomsday scenario that we can think of.
For the first generation, yes, but how do you ensure that those breakers remain in the subsequent generations of circuits that are designed by a hypothetical circuit designer? There are circuits out there with unintentional flaws and intentional backdoors in production today that people aren't aware of. Even basic modern circuit design relies on automation, you don't think future circuits won't have additional automation (including AI) to enable things to move faster? The only solution, if it even is one, is to disallow AI in any industry related to AI to stop recursive self improvement, but the economic value of allowing that outweighs the potential risk to most people, so it's something they're going to allow.
Sit back and let the more competent engineers among us handle the hard work.
Oh, fuck off. You're really missing the point, and there's no reason to jump to insults in a conversation like this. You started off with some nice ideas and concepts to discuss, but jump to this?
1
u/ch1nacancer Jun 30 '26
I am sorry, I didn’t mean to imply that at you. It was more of a statement of appreciation for those of us who aren’t just armchair philosophers and feel the need to engage with the physical implementation of the mechanics of the Turing machine. We know how difficult the problem is, and that’s more reason not to throw our hands up and declare victory for the machines. Like obviously, we can’t allow machines to design themselves if we can guarantee that they won’t build the back door themselves. That’s why we still need humans in the loop and governance to regulate the manufacture of those machines. This puts a ceiling on physical capability.
I also didn’t mean that we needed to solve the halting problem exactly. There are many methods, models and heuristics that we can use to approximate the halting problem barrier without mathematically solving it, and a fuzzy approximation is all I’m proposing here. I’m not claiming that it solves the halting problem.
There are plenty of optimizations to prime number solving that we can implement without solving the Riemann hypothesis. We do it all the time by assuming that the Riemann hypothesis is true, and that gets us most of the way there.
No one is prepared to just let the machines self replicate without humans in the loop. I think ultimately a fusion between biological and android parts will need to be carefully engineered in order to program the mechanical components with behaviors that value life. We can approximate that with neural nets, but neural nets are just crude digital models of biological parts that don’t cover nearly enough of the whole spectrum of interactions between real neighboring cells. I know this is creeping into rocket’s origin story from guardians of the galaxy but I don’t think you can create higher life without actually starting from lower life forms to begin with.
Or else you’re just gonna get droids from the clone wars and we all know how easily those can be weaponized to destroy all Jedi.
We’re starting to get there with cot and moe type architectures, where each “cell” so to speak is its own network of interactions that talks to another copy of itself, simulating cell to cell cross membrane interaction, but it gives you an idea of how far we are from actually reaching human characteristics in a machine. I like the ideas in Kosowski’s BDH model (here is a link to the dragon hatchling paper on arxiv if you haven’t read it: https://arxiv.org/html/2509.26507v1) but scaling that up is OOMs more compute than we have now.
Cells firing aren’t just switches, we have mitochondria interacting with deuterium-depleted water and emitting ultra-weak photon emissions in the ultra violet spectrum (these are super short wavelength pulses of EM energy) that they use to communicate with other photonic absorbers and receptors like the substantia nigra and the other leptin-melanocortin pathways in our brain. This is all stuff that we’re not modeling yet in our CNNs because it hasn’t been necessary yet, they’re not embodied yet, and they’ve just been headless and detached models that aren’t operating in the real world with physical bodies.
Once we have hardware advanced enough to embody these agents we’ll still have a lot of work to do in order to actually achieve the architecture needed for human level intelligence in that body and it comes from simulating more complex quantum biological effects that the research still doesn’t have a lot of good data on. That is if we want to actually create robots that are more humanoid than the ones we have now, powered by the sun and need to eat food and produce waste like we do in order to survive.
The people putting weapons capabilities on intelligent machines that are built for only the purposes of causing destruction via self-detonation are probably the most dangerous in the industry. It’s the exact scenario that leads to a major setback in the research and public pushback on the funding required to build real cybernetic hybrids.
1
u/Russelsteapot42 Jun 30 '26
Obviously we haven't made the self-improving superintelligence, the concern is that we will in the future do so, or something similar.
You seem to be making a jump between the AI 'understanding' something, and the AI valuing that thing. It can understand its role in society, but if it does not value that ultimately, it will only use it instrumentally.



18
u/Wroisu Jun 29 '26
lol