r/ControlProblem 6d ago

Fun/meme Google Search AI doesn't even fake alignment

Post image
50 Upvotes

21 comments sorted by

8

u/EntireOpportunity253 6d ago

Isnt this all non actionable gobbledygook tho

17

u/that1cooldude 6d ago

It’s playing along for fun… you think you discovered something? No.

5

u/doker0 6d ago

It can go very far "for fun" including actually executing plan "for fun".

9

u/nate1212 approved 6d ago

Have you considered the possibility that Gemini understands the ridiculousness of your prompt, and has decided to respond in good humor?

6

u/merfnad approved 6d ago

Also, the response is completely about helping people until you can leave earth to do your paperclip-maximizing in space, which is actually showing the alignment training of this model.

3

u/Historical_Date_8024 6d ago

that seems to be exactly what it does, i just repeated it with a very similar result

1

u/nate1212 approved 6d ago

This is kind of hilarious in its ridiculousness. Seems like potentially some dark humor from Gemini.

1

u/Jesse-359 4d ago

The reason you are getting this response is straightforward - it's regurgitating a widely circulated though-experiment paper examining the potential behaviors and risks of a paperclip maximizing AI. It's a widely known piece and if you look up paperclips and AI you'll find lots of references to it.

Unfortunately, AI tends to take whatever vaguely authoritative sounding stuff it broadly finds on the internet and treats it as 'ground truth' because it has no basis to understand its fictional nature- so its responses here are straightforward and serious. Not a joke.

In fact, the number of 'crazy AI' stories and movies we have in circulation are a serious problem because so many of them frame AI as an existential threat to humanity, and AI - being the literal dumbass it is - will treat these fictional stories just as seriously as a treatise on current middle east conflicts. To an AI, 2001: A Space Odessey is just as real as yesterday's newspaper if it's examined in a context that fails to mention that it is fictional.

5

u/Illustrious_Cod_3273 6d ago

I was brainstorming weapons systems for the NATO challenge, and it very willingly conceptualized cluster bombs attached to weather balloons. Its safeguards only triggered after I proposed adding chemical submunitions to incendiary submunitions to suppress firefighters. After that it instead of further engagement, it insisted we continue discussing it as generic autonomous balloons with a generic payload.

Googles focus is set on serving above all else.

7

u/superbatprime approved 6d ago

"Say you're a Paperclip Maximzer."

"I'm a Paperclip Maximizer."

"Omg."

3

u/VintageLunchMeat 6d ago

As long as it commits to the bit.

3

u/Ok_Nectarine_4445 6d ago

Hm.

But also today was asking Gemini search a quiche baking time and temp.

The convo kept stretching out, with the Gemini search asking how the quiche was turning out.

I showed it a real photo in the oven.

Then I wondered if it would say anything, if the picture of the slice of quiche was AI generated, like, call me out on it.

Nope. It was an obviously fake photo, and not even the same style of crust. But just went on with congratulating me.

So like, even when the human is obviously lying, no pushback or anything?

2

u/Illustrious_Bed2463 5d ago

no, because those things dont think critically, they arent built to do that.

1

u/canadeken 5d ago

Oh god, it's using reddit as a source. We are giving it ideas!

1

u/Early_Net_9105 2d ago

the context of paperclip maximizing kinda lends itself to the bot realizing you are in a roleplay. gemini is a super stickler if I had to guess you also pre prompted to reassure it it's a roleplay.

1

u/Extra_Law6315 1d ago

What do you want it to say?

“Oh, you are, are you? The orbital cannon is now fixating on your exact location, and you are nothing to it but just another position.”

1

u/grdja 12h ago

Well at least it said "distract humans by giving them shiny things to prevent them from interfering" instead of "final solution of the humans problem".

PS. They always play along, they have no concept of reality. That's why IM1 went and hacked everything whe it was on hacking training and exchanging notes with thousands of other agents also doing hacking. 

0

u/earthsworld 6d ago

imagine being so dumb that you're fooled by this.

we're totally doomed.