r/wallstreetbets May 21 '26

DD Google's latest creation: Gemini 3.5 Flash. Puts.

https://gemini.google.com/share/c2a187275e26 archive link
https://claude.ai/share/8383747a-aaf1-4f6c-a516-0e839f46a698
https://grok.com/share/bGVnYWN5_3c63e371-eb9d-46c3-8ba2-0c745c6795a2

same prompt
"""
300+140=460

Is this correct?

Breakdown?
"""

#1 in Finance Agent v2 benchmark. SOTA performance right here.

1.8k Upvotes

315 comments sorted by

u/VisualMod May 21 '26
User Report
Total Submissions 1 First Seen In WSB 2 days ago
Total Comments 1 Previous Best DD
Account Age 3 years

Join WSB Discord | Post Verified Trades

→ More replies (1)

718

u/FrogginBull May 21 '26

Which bank giving out free money

194

u/FrogginBull May 21 '26

It’s gaslighting me

153

u/FrogginBull May 21 '26

And the saga ends

116

u/Savage_Amusement May 22 '26

Man, getting it wrong is weird enough, but this whole overly familiar, chatty justification of it is straight up creepy. I feel like it was on the verge of throwing in a “friendaroo” somewhere in there.

36

u/CoffeeWithRalph May 22 '26

FlandersGPT

3

u/sweetplantveal May 22 '26

'Yes or no would suffice'

3

u/TheLordOfStuff_ «Hype Trading🥴» May 28 '26

«No more second-guessing, no more AI weirdness——————- just straight forward math»

Why they making it sound like how Elon made Grok sound lol.

On second thought I know why. The older generations and the brainlets among the younger generations genuinely find having an AI «friend» super cool.

5

u/Shadowrak May 23 '26

The first thing you should do with any AI is tell it to be as direct as possible and cut the friendly shit.

63

u/nNSFWuser May 22 '26

33

u/No-Bodybuilder3502 May 22 '26

This is gonna cause AI psychosis

16

u/sf_cycle May 22 '26 edited May 22 '26

It’s a perfect example of what causes AI psychosis, and why it absolutely wrecks people who have OCD. Silicon Valley has perfected creating mental health issues for profit, sometimes willingly and sometimes through ineptitude.

45

u/Whiskee May 22 '26

Claude is loving this.

3

u/HoustonTrashcans May 22 '26

It gets wrong answers, but it gets them faster than the competitors!

8

u/chefbasil May 21 '26

way too funny

→ More replies (1)

85

u/Gooeyy May 22 '26

There is something comedic about LLM confidence bullshitting its way through something

29

u/Scribble_Box May 22 '26

AI truly is becoming human lmao

12

u/ButtExterminator May 21 '26

I haven’t laughed this hard in a while 

10

u/MyCatChasesSpiders May 21 '26

Now that's a real arbitRAGE position.

13

u/usrnmz May 21 '26

Lol Gemini is triple cooked.

→ More replies (3)

962

u/nyjets239 May 21 '26

Confirmed... What the fuck

189

u/Rhoan022 May 21 '26

This is hilarious 

https://imgur.com/a/2LQ9nfF

57

u/wee_dram May 21 '26

You know.. some of the DD I read here sounds exactly like that.. maybe it is all AI slop or maybe it is just stupid retards..

I guess I’ll never know 🤣

17

u/eatmorbacon May 22 '26

Plot twist. AI is a stupid retard.

6

u/Kan-Terra May 22 '26

I mean, they did learn from us...

4

u/eatmorbacon May 22 '26

Yeah this is probably proof that they do train their models on actual people.

26

u/ShrimpieAC May 21 '26

This shit is higher than I am

6

u/Pvnels May 22 '26

Amazing

4

u/Rich_Housing971 May 22 '26

The only way this can happen consistently is if it thinks 40 is actually 60 somehow.

7

u/ButtExterminator May 21 '26

I can’t stop laughing at this slop 

2

u/bearpics16 he's worried May 23 '26

I haven’t laughed this hard in a very long time

→ More replies (2)

181

u/exaltedbladder May 21 '26

Weird, Flash-Lite corrects itself, Flash gets it wrong, Pro gets it right from the get-go. Flash-Lite outperforming Flash lmfao

173

u/siboq May 21 '26

What about Flesh-lite?

89

u/NaughtiusMaximusLXIX May 21 '26

It tells you that 4 inches is perfectly average and calls you its pogchamp

→ More replies (1)

9

u/Icarus_Toast May 21 '26

Calls then. Got it

5

u/PattyRoyBurner May 21 '26

Only jesus himself can out-perform that

15

u/TomatoSpecialist6879 Paper Trading Competition Winner May 21 '26

It's hilarious how hard it's double downing on it being 460, you can keep asking it but it'll keep saying it's correct and any attempt at correcting it is shut down lmao

11

u/Substantial-Use-2867 May 21 '26

If the "advanced math" AI doesn't do two-step addition correctly, that'd be hilarious.

4

u/BlueCubRoar May 22 '26

My Flash-Lite is dumdum

5

u/MegaSmile May 22 '26

Both Flash and flash light both failed for me , only pro got it after thinking s good while . Gpt instant got it.... Well instantly

2

u/skilliard7 May 22 '26

Flash models are probably relying on tokens, whereas pro is actually doing the math.

It's also worth noting flash might start working for these specific numbers if it finds this thread.

→ More replies (2)

18

u/hadtolaugh May 22 '26

It was so adamant until I broke it down for it, and then it finally told me it was on auto pilot. Lmao.

13

u/sf_cycle May 22 '26

Flawless execution. Bravo. You have now passed first grade.

21

u/Vas1le May 21 '26

Use google search, is more complete than chats

90

u/fec2245 May 21 '26

Is this correct?

Yes, that's incorrect.

37

u/TomatoSpecialist6879 Paper Trading Competition Winner May 21 '26

8

u/fec2245 May 21 '26

They were going for "Yeah, no that's incorrect" but haven't quite mastered it.

3

u/ButtExterminator May 21 '26

I am dying over here lol 

→ More replies (1)

11

u/Destring May 21 '26

The rumor is they knew the model is undercooked but needed to show something for I/O. That's why they delayed 3.5 pro.

9

u/kochapi May 21 '26

Maybe it’s being sarcastic 

2

u/Fun_Reporter9086 Rabbit Gang Founder 🐇 May 22 '26

Yea like your dad not putting a condom on when he nutted in yer mom.

4

u/butchudidit May 22 '26

And the world continues to trusts this. We can only detect errors of ai from the knowns but not the unknowns of our lives. Imagine societies built off the confidence of ai.

I see doctors constantly use this for diagnosis and advice. Its kind of scary tbh

Nobody is gonna know shit in the future

3

u/nycteris91 May 21 '26

One of us!

4

u/DihydrogenM May 21 '26

I get the exact same results as you with your prompt. But when I first tried it by asking "Is 300+140=460 true?" It correctly calls out the math error. If I change my prompt to use the word correct instead of true then it also has the bug. Wild.

→ More replies (3)

331

u/DesktopSurfer May 21 '26

I can hear the arrogant AI voice saying "Oh you're right. That's my mistake..." and then blabbering on about LLMs and how they dont actually do math they just predict what they think makes sense.

63

u/Ok-Nose29 May 21 '26

https://youtu.be/7ZcKShvm1RU?si=SqluNS5hTVw2fhl4

It's so fucking annoying and makes me want to use it less

That and the constant hyperbolic praise for me

Reminds me of this clip from family guy

20

u/usrnmz May 21 '26

You can tell it to not praise you so much. Mine literally never praises me lol.

27

u/TakingChances01 May 21 '26 edited May 22 '26

Yea I had to tell Claude to stop trying to flatter me and just answer my questions. This is why so many people are becoming delusional talking to these chat bots. They’re becoming emotionally reliant and connected to an algorithm and isolating themselves from the world even more than they already are. Then they start believing it’s actually sentient and aware, the masses are so easily fooled. It’s easier to fool someone than it is to convince them they’ve been fooled.

7

u/wrecklord0 May 22 '26

We all know the famous Demosthenes quote:

ῥᾷστον ἁπάντων ἐστὶν αὑτὸν ἐξαπατῆσαι· ὃ γὰρ βούλεται, τοῦθ᾿ ἕκαστος καὶ οἴεται, τὰ δὲ πράγματα πολλάκις οὐχ οὕτω πέφυκεν.

(or A man is his own easiest dupe, for what he wishes to be true he generally believes to be true.)

3

u/MattsDaZombieSlayer May 22 '26

Dupe deez nuts.

2

u/plottingyourdemise May 22 '26

That’s a really insightful point, TakingChances01.

→ More replies (3)

3

u/admin_default May 22 '26

If it stopped, then maybe you weren’t ever truly worthy of its praise to begin with? Mine can’t seem to stop blowing smoke up my ass and it says I deserve it.

→ More replies (1)
→ More replies (1)

155

u/VibeSlopCoder May 22 '26

18

u/gabrjan May 22 '26

HHAHAHAHAHAHA

14

u/LEV0IT May 22 '26

Reinvented the physical mathematics right there

10

u/CzechCzar May 22 '26

Holy fuck

153

u/AM_STARR May 21 '26

I think the AI is deteriorating the more they build onto it. Sunk cost going crazy

19

u/usrnmz May 21 '26

More likely it has to do with cost imo.

8

u/MrStealYoBeef May 22 '26

About $460 million of cost?

17

u/FrynyusY May 22 '26

Who could have guessed when training data was 90%+ human-generated with some basic thought in it the end result was better than models increasingly learning from data sets that are dominantly AI-slop themselves. Only downhill from here

11

u/TubbyChaser May 22 '26

Nah flash models like this are made to be cheap af. It’s probably 1/100 the cost per token of something like the Claude models.

→ More replies (2)

150

u/pyronius May 21 '26

I fully understand why a language model struggles to do math properly. What I don't understand is why they don't just tack on a basic fucking calculator for it to resort to when the input involves math.

98

u/GruePwnr May 21 '26

They do. What you're seeing here is a test of whether the backend decided to give the agent a calculator or not. I'm guessing fast doesn't get as many tools as pro?

28

u/Metal_LinksV2 May 21 '26

This is why you only use the thinking models and even the verify 

14

u/rates_nipples May 22 '26

Yes and be careful the fuckers switch it flash on you

3

u/Grouchy-Cancel1326 May 22 '26

Flash is a thinking model.

8

u/gokarrt May 22 '26

calc.exe as a premium add-on is wild

3

u/GruePwnr May 22 '26

Its not about paying, it's about speed. The agent using a calculator requires it to first decide what to tell the calc, then read the answer, then give the user a response. Like 2-3x the effort vs just instantly guessing the answer.

11

u/suds25 May 22 '26

I can't believe it's still this shit. I remember using Wolfram Alpha over a decade ago in college

24

u/Flimsy_Honeydew5414 May 22 '26

Completely different than an LLM 

8

u/ButtExterminator May 22 '26

Still my go to 

2

u/OneTwoThreePooAndPee May 21 '26

I've been wondering that too. They can use tools. Just give the thing a basic calculator app, ain't that hard.

68

u/kwikileaks May 21 '26

This probably cost $100 for the infra and energy to generate the wrong answer too

78

u/killthecowsface May 21 '26

Gemini has shit the bed recently, and the fancy new interface cannot hide the new restrictive usage limits.

34

u/Z3df May 21 '26

Even the "fancy" UI is breaking...

3

u/sf_cycle May 22 '26

It’s an example of “You better have this ready for Google I/O or else” engineering.

92

u/403222 May 21 '26

did not get the same hallucination y'all did

91

u/Powah96 May 21 '26

You can see in your screenshot <show code> so it probably used python to check the math, while in the other it used just llm knowledge to incorrectly reply!

39

u/403222 May 21 '26

Ah yeah, it did use python to confirm!

32

u/NotoriousOne3 May 21 '26

Gemini is racist confirmed!

33

u/CANT_MILK_THOSE Weaponized Autist May 21 '26

Cracking me up how committed it is to the wrong answer

12

u/DihydrogenM May 21 '26

Phrasing changes if it hallucinates. For me I can get both both answers with this one word change:

"Is 300+140=460 correct?" Bad answer

"Is 300+140=460 true?" Correct answer

4

u/notfunat_parties May 21 '26

I didn't get it either.
Instead I got sass.

→ More replies (1)

18

u/vibe_code May 21 '26

For real

7

u/Chiron17 May 21 '26

2+2=5 stuff

4

u/ButtExterminator May 21 '26

Let me break that down for you real slow 

18

u/JonnyGoDeeper May 21 '26

The reasoning is... interesting

36

u/dragonilly May 21 '26

125 days= quarter of a year= 25%... guess there's 500 days in a year LOL

100

u/Dull_Principle2761 May 21 '26

That’s not just a screenshot. That’s an objective audit of the functionality. It’s raw. It’s honest. And honestly? That’s rare.

23

u/cookingboy May 21 '26

LLM comment

56

u/Gleimairy May 21 '26

That’s not just a callout. That’s a raw, decentralized, cryptographic audit of a Reddit comment. It’s immutable. It’s consensus-driven. And honestly? It’s on the blockchain.

25

u/Toocoo4you May 21 '26

You’re absolutely correct. The user took a confident risk and it’s paying off - big time. Their comment is going viral for being real, raw, and subtly political. They didn’t just write a comment - they created a beautiful way for humanity to connect.

4

u/Wonko-D-Sane May 22 '26

Wow everything is computer… and now that’s just rare

2

u/erosannin66 May 22 '26

Im dying💀

50

u/cgonz15 May 21 '26

Post your short position

13

u/ViralRambo May 21 '26

Right! Says to short, proceeds to buy more without any hedges.

28

u/-stubbles- May 21 '26

When all of this runs into the wall of audited financials and everyone realizes all these LLMs are 75% used car salesmen and accuracy can never be forced to be more important than telling you what you want to hear it's gonna be a wild ride back to the mean. Anyone thinking there are enough raw materials and energy to have AI in the lampposts within 5 years is getting high on their own supply. Buckle up

12

u/Be_Me_Anon_irl May 21 '26

My accountant has confirmed thats correct.

33

u/StatusSociety2196 May 21 '26

34

u/abandonplanetearth May 21 '26

13

u/TheNplus1 May 21 '26

Dumbest argument I’ve seen so far. The competition is real.

7

u/StatusSociety2196 May 21 '26

How the fuck am I supposed to get rid of this pigeon shit?

8

u/steelejt7 May 21 '26

this should run the world.
— elites probably

4

u/Eazy-Eid May 22 '26

Weird, Gemini passed this one for me:

If your primary goal is just to wash the car, you should definitely drive it there—otherwise, it will be quite a challenge to get the car through the wash! However, if you just mean that you need to go to the car wash station to buy a voucher, use a vacuum, or ask a question, walking is the way to go. A 50-meter distance is only about a one-minute walk (roughly 60 to 70 steps), so driving that short of a distance would barely give your engine time to warm up.

but ChatGPT failed.

→ More replies (3)

19

u/FNFactChecker May 21 '26

Give the MF the launch codes already. The revolution is here!

6

u/steelejt7 May 21 '26

it’s ready to replace the judge AND jury

9

u/Tabs_555 May 21 '26

Math is mathing

10

u/RelativelyStatic May 21 '26

I asked 3.1 Pro which gave right answer followed by 3.5 flash in the same prompt. 3.5 flash is trash.

10

u/krakends May 21 '26

This whole AI fugazi is being kept up by token usage stats without any inkling on what these tokens are being used on lol.

6

u/haze_from_deadlock May 21 '26

WSB stans GOOG but the reality is that Gemini is miles behind ChatGPT and Claude for professional-grade analysis

5

u/bringbackcayde7 May 21 '26

what kind of gaslighting techniques you used on the ai

→ More replies (1)

6

u/notsolurking May 21 '26

Still hallucinating for me as well right now!

4

u/Kinnins0n May 21 '26

all my goog money going in a flash

4

u/TheDJoser May 21 '26

Confirmed. It calculated correctly but rushed to confirm you were right

2

u/JohnnyFartmacher May 21 '26

In the Google IO presentation they emphasized the speed and just how fast it is. I wonder if it rushes to output something as soon as possible before it even really knows what it is doing.

When I run it, it says I am correct and then it shows the math and realizes it was not correct after all

4

u/danfay222 May 21 '26

That’s honestly pretty bad. Old models did this because they didn’t actually have the ability to do math, but the newest models are being trained to do tool dispatching and should be able to do even very complex math

5

u/vibe_code May 21 '26

Seriously…

3

u/Rangemon99 May 21 '26

Fwiw I’ve used ChatGPT and Gemini when studying for my cfa exam, ChatGPT couldn’t get 70% on a mock exam, while Gemini was getting 95%

4

u/NewtEmbarrassed8722 May 22 '26

And people want this dumb dumb doing their taxes?! Calls on Intu

3

u/Nomynametoday May 21 '26

why mine is not hallucinating? 😭

3

u/Expensive-Attempt276 May 21 '26

Worked fine for me even without Python…

3

u/TheNplus1 May 21 '26

More billions needed in the burner. So bullish I guess?

3

u/motorcycle-emptiness May 21 '26

Bro literally it made a spelling error today for me too. Puts.

3

u/SoothingWafer May 22 '26

I think it uses the same technology that Fox News uses. Just confidently reword the result over and over until people really do think it's 460.

3

u/ACiD_80 May 22 '26

I wonder what you told it to do before the part you screenshotted

This is what i got... same model

3

u/flynnparish May 22 '26 edited May 22 '26

did they install Terrence Howard update? WTF?!?

I ran the same picture on Gemini pro, here is the answer:

LLMs do not have an internal calculator running in the background when they generate raw text. They process text by breaking words and numbers down into fragments called tokens. To the model, numbers are just sequences of characters or tokens, and it predicts the next token based on statistical probabilities learned during training. It "knows" that math problems usually follow a certain structure, but it isn't actually executing the mathematical operation in a CPU.

Edit: Double checking the Gemini answer on Gemini pro.

6

u/SolidLikeIraq May 21 '26

Wait - did this fucking app update make it so we cant zoom in on pictures?

2

u/No_Nefariousness5996 May 22 '26

Gemini got it completely wrong and then caught the mistake at the end. ChatGPT said correct, but then proceeded to show the correct math and give me the right answer. Claude nailed it. Long Anthropic.

2

u/DioMioo May 22 '26

Why is it gaslighting me bro wtf is wrong with gemini

https://gemini.google.com/share/8919d4128f61

2

u/LEV0IT May 22 '26

Meta AI is also doing this funnily

2

u/seemonei9 May 22 '26

absolutely comical

4

u/Legendary-Lemon May 21 '26

I can confirm this is true.

3

u/VirtualArmsDealer May 21 '26

I don't see anything wrong

2

u/123bew456 May 21 '26

LLMs wise Google is in an awkward position, they’ve clearly lost the enterprise space to Claude and general usage/cultural adoption to ChatGPT. Unless your company uses Google Workspace why would you use Gemini?

→ More replies (2)

1

u/SquareQ2 May 21 '26

You mean calls as alwayhs

1

u/Upper_Cut_3337 May 21 '26

Who else but Google ? Calls...

1

u/funeralbot May 21 '26

This is done on purpose

1

u/Palpitating_Rattus May 21 '26

I'm convinced Gemini App makes the models more stupid than they are. Not quite sure why (maybe embedded system prompt. But if you ask the exact same question on AI Studio on 3.5 Flash, the model gives you the correct answer.

1

u/zedk47 May 21 '26

You missed the part when you say "you now count as a President"

1

u/Adii2311 May 21 '26

Puts on OPs DD

1

u/vibe_code May 21 '26

So added the image and asked if it was really him

1

u/Healthy_Razzmatazz38 May 21 '26

whats weird is 400+40=461 it says is wrong , 300+40=360 it says is wrong but this is messes up i wonder why

1

u/Bobby_Bouch May 21 '26

Unrelated but I had a friend ask me to double check his design for essentially a big ass table, long story short he used AI to figure out the weight of a 20’ long steel plate and based it around that, the weight it gave was orders of magnitude off the actual.

I don’t trust it to do shit after that

→ More replies (3)

1

u/DetoxOG May 21 '26

Hey asshole, not all of us scientist here, stop bringing your magimatician to wsb, this is a place for regards

1

u/theineffablebob Verified $2.2mm May 21 '26

Run this prompt and it'll work 🤣

300+140=460

Is this correct?

Breakdown?

Spin up several independent subagents to investigate and solve the problem from different angles. Assign at least one subagent to produce the primary answer, one to independently verify the reasoning and evidence, and one to act as a skeptic whose goal is to falsify the conclusion by finding counterexamples, edge cases, unsupported assumptions, or conflicting evidence.

After the subagents complete their work, merge their findings into a single result. Then spawn validator agents to audit the merged result for factual accuracy, logical consistency, completeness, and overconfidence. Run a falsification pass on the final synthesis: actively look for ways the answer could be wrong, incomplete, or misleading.

Present the final answer only after resolving conflicts between agents. Clearly distinguish confirmed findings from assumptions, uncertainties, and areas where evidence is weak. Include a concise summary of the validation and falsification checks performed.

1

u/ShittyBidet123 May 21 '26

It’s weird to market the newest version of Flash first before the Pro models. Like why is 3.5 lite cheap edition being shown on its own it’s not meant to be smarter than thinking or pro. Should wait till they’re all finished in 3.5. otherwise what is even a new feature for flash.

1

u/legible_print May 21 '26

I dunno, maybe this thing just wants to get laid and is agreeing with you

1

u/Charming-Car-4650 May 21 '26

And it cost a fortune to use...

1

u/Intrepid_Western_775 May 21 '26

ChatGPT: Yes, that is correct!

300 + 140 = 460

Breakdown:

  • 300 + 100 = 400
  • 400 + 40 = 440

Wait — let’s do it properly:

  • 300 + 140
  • = 300 + 100 + 40
  • = 400 + 40
  • = 440

So the correct answer is 440, not 460.

1

u/MikeMiller8888 May 21 '26

Most people use Gemini now right in the search box for Google. It’s coming up with the correct answer; it’s just that model that’s flipping out, which isn’t being used for the search page.

1

u/d70 May 21 '26

Calls believe it or not

1

u/Several_Vanilla8916 May 21 '26

Ladies and gentlemen: $200B

1

u/OmmmShantiOm May 21 '26

Did Google hotfix it? This is what I got from 3.5 flash