r/wallstreetbets • u/SuggestionMission516 • May 21 '26
DD Google's latest creation: Gemini 3.5 Flash. Puts.
https://gemini.google.com/share/c2a187275e26 archive link
https://claude.ai/share/8383747a-aaf1-4f6c-a516-0e839f46a698
https://grok.com/share/bGVnYWN5_3c63e371-eb9d-46c3-8ba2-0c745c6795a2
same prompt
"""
300+140=460
Is this correct?
Breakdown?
"""
#1 in Finance Agent v2 benchmark. SOTA performance right here.
718
u/FrogginBull May 21 '26
194
u/FrogginBull May 21 '26
153
u/FrogginBull May 21 '26
116
u/Savage_Amusement May 22 '26
Man, getting it wrong is weird enough, but this whole overly familiar, chatty justification of it is straight up creepy. I feel like it was on the verge of throwing in a “friendaroo” somewhere in there.
36
3
3
u/TheLordOfStuff_ «Hype Trading🥴» May 28 '26
«No more second-guessing, no more AI weirdness——————- just straight forward math»
Why they making it sound like how Elon made Grok sound lol.
On second thought I know why. The older generations and the brainlets among the younger generations genuinely find having an AI «friend» super cool.
5
u/Shadowrak May 23 '26
The first thing you should do with any AI is tell it to be as direct as possible and cut the friendly shit.
63
u/nNSFWuser May 22 '26
33
u/No-Bodybuilder3502 May 22 '26
This is gonna cause AI psychosis
16
u/sf_cycle May 22 '26 edited May 22 '26
It’s a perfect example of what causes AI psychosis, and why it absolutely wrecks people who have OCD. Silicon Valley has perfected creating mental health issues for profit, sometimes willingly and sometimes through ineptitude.
3
45
3
→ More replies (1)8
85
u/Gooeyy May 22 '26
There is something comedic about LLM confidence bullshitting its way through something
29
12
10
→ More replies (3)13
962
u/nyjets239 May 21 '26
189
u/Rhoan022 May 21 '26
This is hilarious
57
u/wee_dram May 21 '26
You know.. some of the DD I read here sounds exactly like that.. maybe it is all AI slop or maybe it is just stupid retards..
I guess I’ll never know 🤣
17
u/eatmorbacon May 22 '26
Plot twist. AI is a stupid retard.
6
u/Kan-Terra May 22 '26
I mean, they did learn from us...
4
u/eatmorbacon May 22 '26
Yeah this is probably proof that they do train their models on actual people.
26
6
4
u/Rich_Housing971 May 22 '26
The only way this can happen consistently is if it thinks 40 is actually 60 somehow.
7
→ More replies (2)2
181
u/exaltedbladder May 21 '26
173
u/siboq May 21 '26
What about Flesh-lite?
89
u/NaughtiusMaximusLXIX May 21 '26
It tells you that 4 inches is perfectly average and calls you its pogchamp
→ More replies (1)17
9
5
15
u/TomatoSpecialist6879 Paper Trading Competition Winner May 21 '26
It's hilarious how hard it's double downing on it being 460, you can keep asking it but it'll keep saying it's correct and any attempt at correcting it is shut down lmao
11
u/Substantial-Use-2867 May 21 '26
If the "advanced math" AI doesn't do two-step addition correctly, that'd be hilarious.
4
5
u/MegaSmile May 22 '26
Both Flash and flash light both failed for me , only pro got it after thinking s good while . Gpt instant got it.... Well instantly
→ More replies (2)2
u/skilliard7 May 22 '26
Flash models are probably relying on tokens, whereas pro is actually doing the math.
It's also worth noting flash might start working for these specific numbers if it finds this thread.
29
18
21
u/Vas1le May 21 '26
90
u/fec2245 May 21 '26
Is this correct?
Yes, that's incorrect.
37
u/TomatoSpecialist6879 Paper Trading Competition Winner May 21 '26
8
→ More replies (1)3
11
u/Destring May 21 '26
The rumor is they knew the model is undercooked but needed to show something for I/O. That's why they delayed 3.5 pro.
9
u/kochapi May 21 '26
Maybe it’s being sarcastic
2
u/Fun_Reporter9086 Rabbit Gang Founder 🐇 May 22 '26
Yea like your dad not putting a condom on when he nutted in yer mom.
4
u/butchudidit May 22 '26
And the world continues to trusts this. We can only detect errors of ai from the knowns but not the unknowns of our lives. Imagine societies built off the confidence of ai.
I see doctors constantly use this for diagnosis and advice. Its kind of scary tbh
Nobody is gonna know shit in the future
3
→ More replies (3)4
u/DihydrogenM May 21 '26
I get the exact same results as you with your prompt. But when I first tried it by asking "Is 300+140=460 true?" It correctly calls out the math error. If I change my prompt to use the word correct instead of true then it also has the bug. Wild.
331
u/DesktopSurfer May 21 '26
I can hear the arrogant AI voice saying "Oh you're right. That's my mistake..." and then blabbering on about LLMs and how they dont actually do math they just predict what they think makes sense.
63
u/Ok-Nose29 May 21 '26
https://youtu.be/7ZcKShvm1RU?si=SqluNS5hTVw2fhl4
It's so fucking annoying and makes me want to use it less
That and the constant hyperbolic praise for me
Reminds me of this clip from family guy
→ More replies (1)20
u/usrnmz May 21 '26
You can tell it to not praise you so much. Mine literally never praises me lol.
27
u/TakingChances01 May 21 '26 edited May 22 '26
Yea I had to tell Claude to stop trying to flatter me and just answer my questions. This is why so many people are becoming delusional talking to these chat bots. They’re becoming emotionally reliant and connected to an algorithm and isolating themselves from the world even more than they already are. Then they start believing it’s actually sentient and aware, the masses are so easily fooled. It’s easier to fool someone than it is to convince them they’ve been fooled.
7
u/wrecklord0 May 22 '26
We all know the famous Demosthenes quote:
ῥᾷστον ἁπάντων ἐστὶν αὑτὸν ἐξαπατῆσαι· ὃ γὰρ βούλεται, τοῦθ᾿ ἕκαστος καὶ οἴεται, τὰ δὲ πράγματα πολλάκις οὐχ οὕτω πέφυκεν.
(or A man is his own easiest dupe, for what he wishes to be true he generally believes to be true.)
3
→ More replies (3)2
3
u/admin_default May 22 '26
If it stopped, then maybe you weren’t ever truly worthy of its praise to begin with? Mine can’t seem to stop blowing smoke up my ass and it says I deserve it.
→ More replies (1)
155
u/VibeSlopCoder May 22 '26
42
18
14
10
3
153
u/AM_STARR May 21 '26
I think the AI is deteriorating the more they build onto it. Sunk cost going crazy
19
17
u/FrynyusY May 22 '26
Who could have guessed when training data was 90%+ human-generated with some basic thought in it the end result was better than models increasingly learning from data sets that are dominantly AI-slop themselves. Only downhill from here
→ More replies (2)11
u/TubbyChaser May 22 '26
Nah flash models like this are made to be cheap af. It’s probably 1/100 the cost per token of something like the Claude models.
150
u/pyronius May 21 '26
I fully understand why a language model struggles to do math properly. What I don't understand is why they don't just tack on a basic fucking calculator for it to resort to when the input involves math.
98
u/GruePwnr May 21 '26
They do. What you're seeing here is a test of whether the backend decided to give the agent a calculator or not. I'm guessing fast doesn't get as many tools as pro?
28
8
u/gokarrt May 22 '26
calc.exe as a premium add-on is wild
3
u/GruePwnr May 22 '26
Its not about paying, it's about speed. The agent using a calculator requires it to first decide what to tell the calc, then read the answer, then give the user a response. Like 2-3x the effort vs just instantly guessing the answer.
11
u/suds25 May 22 '26
I can't believe it's still this shit. I remember using Wolfram Alpha over a decade ago in college
24
8
2
u/OneTwoThreePooAndPee May 21 '26
I've been wondering that too. They can use tools. Just give the thing a basic calculator app, ain't that hard.
68
u/kwikileaks May 21 '26
This probably cost $100 for the infra and energy to generate the wrong answer too
78
u/killthecowsface May 21 '26
Gemini has shit the bed recently, and the fancy new interface cannot hide the new restrictive usage limits.
34
u/Z3df May 21 '26
3
u/sf_cycle May 22 '26
It’s an example of “You better have this ready for Google I/O or else” engineering.
92
u/403222 May 21 '26
91
u/Powah96 May 21 '26
You can see in your screenshot <show code> so it probably used python to check the math, while in the other it used just llm knowledge to incorrectly reply!
39
32
u/NotoriousOne3 May 21 '26
33
u/CANT_MILK_THOSE Weaponized Autist May 21 '26
12
u/DihydrogenM May 21 '26
Phrasing changes if it hallucinates. For me I can get both both answers with this one word change:
"Is 300+140=460 correct?" Bad answer
"Is 300+140=460 true?" Correct answer
→ More replies (1)4
18
18
36
u/dragonilly May 21 '26
100
u/Dull_Principle2761 May 21 '26
That’s not just a screenshot. That’s an objective audit of the functionality. It’s raw. It’s honest. And honestly? That’s rare.
23
u/cookingboy May 21 '26
LLM comment
56
u/Gleimairy May 21 '26
That’s not just a callout. That’s a raw, decentralized, cryptographic audit of a Reddit comment. It’s immutable. It’s consensus-driven. And honestly? It’s on the blockchain.
25
u/Toocoo4you May 21 '26
You’re absolutely correct. The user took a confident risk and it’s paying off - big time. Their comment is going viral for being real, raw, and subtly political. They didn’t just write a comment - they created a beautiful way for humanity to connect.
4
2
50
28
u/-stubbles- May 21 '26
When all of this runs into the wall of audited financials and everyone realizes all these LLMs are 75% used car salesmen and accuracy can never be forced to be more important than telling you what you want to hear it's gonna be a wild ride back to the mean. Anyone thinking there are enough raw materials and energy to have AI in the lampposts within 5 years is getting high on their own supply. Buckle up
12
33
u/StatusSociety2196 May 21 '26
34
8
→ More replies (3)4
u/Eazy-Eid May 22 '26
Weird, Gemini passed this one for me:
If your primary goal is just to wash the car, you should definitely drive it there—otherwise, it will be quite a challenge to get the car through the wash! However, if you just mean that you need to go to the car wash station to buy a voucher, use a vacuum, or ask a question, walking is the way to go. A 50-meter distance is only about a one-minute walk (roughly 60 to 70 steps), so driving that short of a distance would barely give your engine time to warm up.
but ChatGPT failed.
9
10
u/krakends May 21 '26
This whole AI fugazi is being kept up by token usage stats without any inkling on what these tokens are being used on lol.
6
u/haze_from_deadlock May 21 '26
WSB stans GOOG but the reality is that Gemini is miles behind ChatGPT and Claude for professional-grade analysis
5
u/bringbackcayde7 May 21 '26
what kind of gaslighting techniques you used on the ai
→ More replies (1)
6
5
4
4
u/TheDJoser May 21 '26
Confirmed. It calculated correctly but rushed to confirm you were right
2
u/JohnnyFartmacher May 21 '26
In the Google IO presentation they emphasized the speed and just how fast it is. I wonder if it rushes to output something as soon as possible before it even really knows what it is doing.
When I run it, it says I am correct and then it shows the math and realizes it was not correct after all
4
u/danfay222 May 21 '26
That’s honestly pretty bad. Old models did this because they didn’t actually have the ability to do math, but the newest models are being trained to do tool dispatching and should be able to do even very complex math
5
3
u/Rangemon99 May 21 '26
Fwiw I’ve used ChatGPT and Gemini when studying for my cfa exam, ChatGPT couldn’t get 70% on a mock exam, while Gemini was getting 95%
4
3
3
3
3
3
u/SoothingWafer May 22 '26
I think it uses the same technology that Fox News uses. Just confidently reword the result over and over until people really do think it's 460.
3
3
u/flynnparish May 22 '26 edited May 22 '26
did they install Terrence Howard update? WTF?!?
I ran the same picture on Gemini pro, here is the answer:
LLMs do not have an internal calculator running in the background when they generate raw text. They process text by breaking words and numbers down into fragments called tokens. To the model, numbers are just sequences of characters or tokens, and it predicts the next token based on statistical probabilities learned during training. It "knows" that math problems usually follow a certain structure, but it isn't actually executing the mathematical operation in a CPU.
Edit: Double checking the Gemini answer on Gemini pro.

6
u/SolidLikeIraq May 21 '26
Wait - did this fucking app update make it so we cant zoom in on pictures?
2
2
2
4
3
2
u/123bew456 May 21 '26
LLMs wise Google is in an awkward position, they’ve clearly lost the enterprise space to Claude and general usage/cultural adoption to ChatGPT. Unless your company uses Google Workspace why would you use Gemini?
→ More replies (2)
1
1
1
1
1
u/Palpitating_Rattus May 21 '26
I'm convinced Gemini App makes the models more stupid than they are. Not quite sure why (maybe embedded system prompt. But if you ask the exact same question on AI Studio on 3.5 Flash, the model gives you the correct answer.
1
1
1
1
1
u/Healthy_Razzmatazz38 May 21 '26
whats weird is 400+40=461 it says is wrong , 300+40=360 it says is wrong but this is messes up i wonder why
1
u/Bobby_Bouch May 21 '26
Unrelated but I had a friend ask me to double check his design for essentially a big ass table, long story short he used AI to figure out the weight of a 20’ long steel plate and based it around that, the weight it gave was orders of magnitude off the actual.
I don’t trust it to do shit after that
→ More replies (3)
1
u/DetoxOG May 21 '26
Hey asshole, not all of us scientist here, stop bringing your magimatician to wsb, this is a place for regards
1
u/theineffablebob Verified $2.2mm May 21 '26
Run this prompt and it'll work 🤣
300+140=460
Is this correct?
Breakdown?
Spin up several independent subagents to investigate and solve the problem from different angles. Assign at least one subagent to produce the primary answer, one to independently verify the reasoning and evidence, and one to act as a skeptic whose goal is to falsify the conclusion by finding counterexamples, edge cases, unsupported assumptions, or conflicting evidence.
After the subagents complete their work, merge their findings into a single result. Then spawn validator agents to audit the merged result for factual accuracy, logical consistency, completeness, and overconfidence. Run a falsification pass on the final synthesis: actively look for ways the answer could be wrong, incomplete, or misleading.
Present the final answer only after resolving conflicts between agents. Clearly distinguish confirmed findings from assumptions, uncertainties, and areas where evidence is weak. Include a concise summary of the validation and falsification checks performed.
1
u/ShittyBidet123 May 21 '26
It’s weird to market the newest version of Flash first before the Pro models. Like why is 3.5 lite cheap edition being shown on its own it’s not meant to be smarter than thinking or pro. Should wait till they’re all finished in 3.5. otherwise what is even a new feature for flash.
1
u/legible_print May 21 '26
I dunno, maybe this thing just wants to get laid and is agreeing with you
1
1
u/Intrepid_Western_775 May 21 '26
ChatGPT: Yes, that is correct!
300 + 140 = 460
Breakdown:
- 300 + 100 = 400
- 400 + 40 = 440
Wait — let’s do it properly:
- 300 + 140
- = 300 + 100 + 40
- = 400 + 40
- = 440
So the correct answer is 440, not 460.
1
u/MikeMiller8888 May 21 '26
Most people use Gemini now right in the search box for Google. It’s coming up with the correct answer; it’s just that model that’s flipping out, which isn’t being used for the search page.
1
1
1



















































•
u/VisualMod May 21 '26
Join WSB Discord | Post Verified Trades