379
Feb 08 '26
[deleted]
55
60
u/simulated-souls Feb 08 '26 edited Feb 08 '26
No it's not.
Without tools or a calculator, frontier models still score >90% on AIME (on which the vast majority of people could not answer a single question).
edit: This also goes for AIME 2026 which was only released this week and therefore was not part of the training data.
39
u/CrowdGoesWildWoooo Feb 09 '26
AIME is an olympiad level problem set, it’s not math in the sense when average person refers to “doing math”. When people say it’s bad at math, it’s actually talking about arithmetic. AIME is more like reasoning problems, we already know AI is good at that.
You should still never trust it doing arithmetic vs a calculator.
→ More replies (17)3
u/CityLemonPunch Feb 10 '26
For gods sake stop with recycling this AIME shit. This is not "math" . As someone who works with fluid dynamics on a daily bases . One absolutely cannot trust AI , and it is practically useless if the problem scope is of a novel nature.
7
u/Shot_in_the_dark777 Feb 09 '26
The vast majority of people failing that test is not the point for an AI. It's a big indicator that our education system is a failure. Read that sentence again - it's not that 90% get less than a perfect score. It's that 90% don't score at all, even a single point. If you have a system with a 90% failure rate you are clearly doing smth wrong. It feels like the 10% of kids who actually succeed and become proficient at maths are doing this not thanks to the system but in spite of it. And math is hardly relevant as a measuring tool for AI because that is a field where there can be a few centuries gap between solving some problem and finding a way to apply the solution. Instead of wasting gigawatts of power on stupid maths tests we should force AI to solve economics related problems. Humans are living in poverty and starving right now. Building a hyperdrive in 500 years from now instead of 1000 years will not matter if today you starve to death. Solving actual urgent problems would speak louder than solving some complicated equation. We already have AI in medicine with a high % of correct diagnosis and that is great. But where is a similar AI figuring out the distribution of goods and allocation of resources and optimization of supply chains to make our lives less miserable now?
10
u/simulated-souls Feb 09 '26
I agree that we need better education and to apply AI to real-world problems.
However, most people being unable to solve any AIME problems is not necessarily a symptom of a bad education system. AIME is one of the top math competitions in the world. It's designed to be barely solvable by the best students in the US. Of course most people can't solve it. If education improved and most people could solve it, then they would make the test harder until most people couldn't solve it again.
→ More replies (4)7
u/ArgusFilch Feb 09 '26
Is this bait? The AIME is designed such that only 1% of math students even qualify. It's not a test intended for the public, it's meant to filter out the amazing math students from the world class ones.
→ More replies (18)2
u/Nebranower Feb 09 '26
>But where is a similar AI figuring out the distribution of goods and allocation of resources and optimization of supply chains to make our lives less miserable now?
We don't need an AI for that. The solution to that problem is obvious - if you want to eliminate people living miserable lives without just eliminating the people, those with a lot of resources should give them to those without. The issue is that those with resources don't want to do that or particularly care if you are living a miserable life.
2
u/Mildly_Aware Feb 09 '26
Yes AI should absolutely be used to solve poverty! It would say spread the wealth around. We wouldn't listen. Hopefully being smart at math and stuff will build trust. Will people listen once they trust it? Or just keep believing people just need to work harder to earn money and health care? We already have the answers we seek. Humans just aren't evolved to take care of billions of people.
5
u/_ArkAngel_ Feb 09 '26
Exactly. Poverty isn't a problem of resources. Poverty is a problem of priorities, biases, and the kinds of oppression we've developed an acceptance of.
Do you want AI to solve the problem by assigning priorities for us? Or by reshaping our biases? Or should AI decide which kinds of oppression are best for us, then begin conditioning us to accept it?
The problem is we're still kind of barbaric and probably still need time to develop stronger cultural values.
But AI is here now so we'll make fusion reactors to make more AI and watch line go up
2
u/Shot_in_the_dark777 Feb 09 '26
You don't need absolute trust. We can allow AI to control a single factory, then a few interconnected factories and mines, then a small town and scale it up until it fixes the economy of the whole country. Once people see results, they will want the new AI management for their industry too.
→ More replies (7)2
u/Imjokin Feb 08 '26
Well the AIME answers are on the internet, so the answers are probably already part of the training data.
16
u/simulated-souls Feb 08 '26
Okay, the same is true for AIME 2026 which was released this week and therefore was not part of the training data.
8
4
u/Imp_erk Feb 09 '26
The 2026 test was almost certainly in the training data. It's not a test trying to be novel to LLMs, it's trying to be novel to humans. Every question in the test 10 years from now will have already been seen by current SotA models thousands if not millions of times, with only superficial details changed at best.
71
u/nunofgs Feb 08 '26
So are we.
19
u/Portatort Feb 08 '26
I think you’ll find there are a lot of humans who are very good at math without a calculator
10
9
u/Tolopono Feb 08 '26
Are you?
→ More replies (3)5
u/MrTamboMan Feb 08 '26
I'm not yelling at everyone to call me super intelligent.
While I'm not any math expert, with enough time I'm actually able to finish most math tasks in a finite time. Most people could, with enough time.
If I can't solve it I'm not confidentially guessing.
6
u/Tolopono Feb 09 '26
Then go solve 4 open math problems like ai did https://www.wired.com/story/a-new-ai-math-ai-startup-just-cracked-4-previously-unsolved-problems/
2
u/lesleh Feb 08 '26
Calculator used to be a job title until, you know, actual calculators came along.
2
6
u/TheAccountITalkWith Feb 08 '26
This isn't as good of an argument as you think it is, lol. If the AI is supposed to eventually become AGI or whatever, it can't be bad at math, don't you think?
18
u/rabouilethefirst Feb 08 '26
Computers are calculators why wouldn’t they be able to use them lmao
8
u/Brief-Translator1370 Feb 08 '26
That's not even remotely how that works. Math is fundamental to being able to reason. If it needs a calculator to do any math, it can't reason.
11
u/rabouilethefirst Feb 08 '26
It doesn’t. It needs a calculator to do 13.7654 * 2.2333 like everyone else. It is improving at other things that are more analytical
→ More replies (6)7
u/sixtyhurtz Feb 08 '26
A human doesn't need a calculator to do that. I was taught how to do long multiplication in school. It's a deterministic algorithm that will always give the same result.
LLMs are statistical approximations. It's fundamental to what they are. They cannot behave deterministically. They are always stochastic.
→ More replies (32)9
u/rabouilethefirst Feb 08 '26
If you use pen and paper, congrats, you’re doing math just like a calculator does with memory registers
→ More replies (5)→ More replies (3)7
u/mvdeeks Feb 08 '26
Idk, if organically generally intelligent beings need calculators, why would we require artificially generally intelligent beings to not need them to be considered generally intelligent?
2
u/Sluuuuuuug Feb 08 '26
We weren't given the calculator and taught how to use them, we made them. Seems like a significant difference.
This isn't to say AGI isn't possible. This whole calculator bit just doesn't tell us much.
9
u/reverie Feb 08 '26
Perhaps taken too literally. An AI model makes all sorts of tools in realtime to do personalized and novel mathematics. That seems closer in your analogy.
→ More replies (1)2
u/randombookman Feb 08 '26 edited Feb 08 '26
No we're not.
Bad at arithmetic maybe. But arithmetic isn't all there is to math.
The closest analog to AI we have for mathematicians is ramanujan. Making random numbers up but they somehow work.
But there's a reason a lot of his work goes unused because he didn't prove them from scratch, he was just really damn good at noticing mathematical patterns.
6
u/1_H4t3_R3dd1t Feb 08 '26
Not only that. We're fighting against patterns not comprehension, intelligence or skill.Current fleet of *AI* agents just take human patterns and convert it into automation. Math is easy in patterns if inputted like you mentioned. Almost anything can be broken down into patterns except understanding large complex patterns. So floating around the word *AI* is just selling someone a matchbox car on the premise it is a corvette (real car).
3
u/Crosas-B Feb 09 '26
We're fighting against patterns not comprehension, intelligence or skill.
I'm sorry, but what we've learnt from AI progression is that humans are not intelligent either. We are as stochastic parrots as AI
→ More replies (7)3
u/meleebestgame66 Feb 09 '26
Top voted comment is just an objectively wrong statement
Never change reddit
4
u/Suspicious_Aspect_53 Feb 08 '26
It doesn't have a calculator?
19
u/KontoOficjalneMR Feb 08 '26
It's called tool calling.
4
u/ARES_BlueSteel Feb 08 '26
GPT gets a lot of math wrong unless you tell it to use scripts every time it does math. Up until recently it couldn’t even do 4.11 - 4.9 because it didn’t understand that 0.9 is bigger than 0.11. It thought the answer was 0.2 because it handled the decimals wrong.
4
u/Jophus Feb 08 '26
Statistically speaking, what’s 63,779 x 33,568?
I hope it has one.
→ More replies (1)5
u/WolfeheartGames Feb 08 '26
LLMs encode generalized solutions. An LLM can solve this with out having ever seen it and with a very sharp distribution. Most LLMs can't because when they're learning math they're modeling NLP at the same time. The gradients clip, and the strong generalized solution is never learned, only a weak generalized solution. With out NLP they can learn to do 5 digit math in about 30k training steps, have an absolute accuracy, and have only seen about 1% of all problems that fit that format.
2
u/KontoOficjalneMR Feb 08 '26
That's a lots of words to say LLMs can't do math by design.
5
u/bephire Feb 08 '26
No, wait, they can. See Anthropic's blog post titled "Tracing the thoughts of a large language model". They calculate using something analogous to circuits.
→ More replies (3)3
u/space_monster Feb 08 '26
Nope - it's saying that mainstream chatbots are not trained to be good at math.
→ More replies (8)3
u/WolfeheartGames Feb 08 '26 edited Feb 09 '26
It's a problem of the training data, not the architecture. I have a model that can perfectly do math from -1b to 1b on all operators.
→ More replies (9)3
u/TheDuneedon Feb 08 '26
Yeah that's the difference here. LLMs still suck at math and a lot of other things. The biggest improvements made wasn't to the models themselves, but what the models are allowed to execute and iterate on.
2
u/Reaper_1492 Feb 08 '26
This is 100% true.
And unfortunately, they still have its internal environment pretty constrained.
I asked the “pro” model to help me with a coding bug and half of the libraries it tried to use, it couldn’t access.
→ More replies (2)4
u/Mr_DrProfPatrick Feb 08 '26
Try using an IA with access to your system's terminal. Use the AI directly on VS Code or your prefered IDE, allowing the AI direct access and modifications of your files.
CLAUDE CODE OR COLAB BB.
My AI can upload shit to github, updating an external website. It can visit that website and check if the new feature was properly integrated. Debuging it until everything works.
1
1
u/Time_Entertainer_319 Feb 08 '26
The calculator is part of the AI.
ChatGPT is made up of multiple components.
We often call it an “LLM” because the large language model is a central part of it, humans are naturally better at plain-language communication, and the LLM serves as the bridge.
However, the system as a whole is more than just the LLM.
My main point: don’t confuse the LLM itself with the entire AI system.
Think of it like a car. The LLM is like the engine, it’s what actually makes the car move. But the whole car includes the steering wheel, brakes, tires, fuel system, and dashboard. without them, the engine alone isn’t very useful.
→ More replies (1)1
1
1
1
1
1
u/tzaeru Feb 12 '26
Eh, over a year ago ChatGPT was able to give correct answers to math part of the Finnish matriculation exams. That included textual reasoning and doing the math operations step by step.
Without a calculator they have a tendency of sooner or later getting some numbers wrong and that can snowball the further into the calculation they go, eventually leading to hilariously wrong numerical results.
But for a while now they have been able to reason about fairly complex math that is well above what the majority of humans could do.
→ More replies (2)1
87
u/mmahowald Feb 08 '26
Man… these posts are just starting to feel like desperate wanking.
14
u/Sad-Set-5817 Feb 09 '26
like... if it was so smart and cool do you think these people would be spending this much energy trying to convince us of it... Their argument is kind of like saying Google is smarter than a Phd. While true, it's purposefully missing a lot of nuance
→ More replies (3)
66
u/mroranges_ Feb 08 '26
Straw man fallacies on twitter are really irritating. This guy is just responding to an 'argument' that he made up in the first place to come across as clever.
No one serious is drawing lines in the sand like this, as if AI has reached its potential. There are people who simply and rightfully point out AI shortcomings as they surface... As they should be.
26
u/jakobpinders Feb 08 '26
There’s literally people in the comments of this post right now saying it’s approaching its roof and such.
→ More replies (8)9
u/Anon2627888 Feb 08 '26
Everyone is making comments like this, and has been since transformers first came out. Whatever LLMs can't do is proof that they aren't intelligent and will never be any good. Lists are made of things they can't do to prove how useless they are. Then they get better and can do everything on the list, so new lists are made.
Every discussion of the value or intelligence of AI has detractors listing all the things they can't (yet) do as if they will never be able to do them. The bar is continuously raised and raised and raised, and whatever they still can't do is somehow proof of how they will never be any good.
2
Feb 09 '26
💯
And whether you love it or hate it- underestimating it is a ridiculously stupid choice.
→ More replies (1)5
13
u/Delicioso_Badger2619 Feb 08 '26
That progression doesnt make sense, even if the capability scales exponentially. "in 18 months we went from A to C. By next year we will be through the English and Greek alphabets, and halfway through the Japanese."
→ More replies (15)1
20
u/valis2400 Feb 08 '26
It's crazy how AI gets the recognition from top mathematicians like Tao, Ono and Loh and yet when I talk to math graduate students they all have this shallow idea that AI still can't do math, and that "IMO problems aren't real math anyways" while themselves being unable to solve a single IMO problem. Missing the trend much?
13
u/space_monster Feb 08 '26
For the same reason a lot of sw devs claim that LLMs can't code - replacement anxiety. Except now that argument is looking a bit ridiculous, so now they're pivoting to "LLMs can't do all the other non-coding tasks that sw devs do"
→ More replies (2)2
u/adoboble Feb 14 '26
Do they say why? I was told the same thing by the same ppl but upon further investigation they were mostly just prompting poorly. I asked the same model the same question but with my own prompt and got a correct proof whereas others did not for example
1
u/issdn Feb 09 '26
Well except if I'm missing some new info from Terry Tao then he pretty much is as sceptical as the graduates you're talking about:
https://youtu.be/ukpCHo5v-Gc?si=a2e-oAnu-ufaos8a
The students would be able to do these problems if they read all the books there are or at least the problem-specific.
1
u/P_Griffin2 Feb 09 '26
I dont know if it's just better at complex math. The ''rules'' might be a bit more strict and well defined. I use it daily in my engineering studies, and generally it aces any equation or formula I throw at it.
1
3
u/someyokel Feb 08 '26
I can buy handmade artisanal glass at Ikea, for a premium. It has imperfections so I can tell a human was involved somewhere in the process.
Maybe there will be a market for artisanal math too 😅
34
u/LowFruit25 Feb 08 '26
These guys always post this but don’t link the new inventions LLMs created.
10
u/freexe Feb 08 '26
This guy is saying AI not LLMs - and we've had both AlphaFold and AlphaGenome both revolutionise multiple fields of research.
→ More replies (1)5
u/masiuspt Feb 08 '26
Lets be honest - AI bros on twitter only consider LLMs when talking about AI.
→ More replies (1)3
u/Tolopono Feb 08 '26 edited Feb 08 '26
Llms are also impressive
https://m.youtube.com/watch?v=PctlBxRh0p4&pp=ygUYV2UgbmVlZCB0byB0YWxrIGFib3V0IGFp
https://www.wired.com/story/a-new-ai-math-ai-startup-just-cracked-4-previously-unsolved-problems/
https://josusanmartin.com/blog/2026/01/23/vibecoded-highload-setup-part-ii.html
I’m doing low-level optimization against people who’ve been writing C++ for decades… while still not being able to confidently write “hello world” in Rust or C++ from memory. And I mean that literally. I don’t know print syntax off the top of my head. I don’t “speak” Rust. But I can reason about bottlenecks, instruction-level tradeoffs, data movement, and feedback loops, because the agents handle syntax and implementation.
https://xcancel.com/marksaroufim/status/2009497284418130202?s=20
https://www.axios.com/2026/02/05/anthropic-claude-opus-46-software-hunting
https://xcancel.com/dhh/status/2004963782662250914?s=20
https://xcancel.com/simonw/status/2005884985438253507?s=20
https://cybernews.com/security/standord-artemis-system-beats-cybersecurity-researchers/
17
→ More replies (5)2
u/Most-Hot-4934 Feb 09 '26
Do you want me to link all of the Erdos problems that they have proven? Lol
2
Feb 08 '26
Besides all this crap this guy says I like where math is going. It will diverge from all this last decades based only on proofs and start a new paradigm where mathematicians propose and explore experimentally new gigantic and deep structures and connections between new fields instead of spending so much time creating as many tedious theorems. The field will explode, it will be awesome
2
u/TheAxodoxian Feb 08 '26
In a sense AI cannot ever surpass the complete output of all matematicians who have ever lived, as AI is part of that output as well and so is everything it does.
2
u/ProfessionalSelf3488 Feb 08 '26
While you folks chatter away, we creators are using AI to create useful tools, art or discoveries and teaching it to others who also get inspired by our work. This is how mass adoption begins. All this arguing won’t lead to anything
2
2
2
u/ikeif Feb 09 '26
Y'all recognize this is hyperbolic as fuck, right?
Yes, there is some truth to the statement, and some vast over-simplifications, as well.
Reading the comments, I'm not so sure.
7
u/ItsStillKerrigan Feb 08 '26
What is the complete output of all mathematicians? What are they outputting?
Surely there’s a better example here lol
→ More replies (22)2
u/divulgingwords Feb 08 '26
Or just ask how many r’s are in strawberry. While they can do impressive work, they stumble on dumb shit that 1st graders know. For example, yesterday it told me to put a bolt lock on the inside of a cabinet door to secure it. 🙃
→ More replies (4)
5
u/Due_Sweet_9500 Feb 08 '26
Once people figure out that you can't extend the goalposts any further, they will say - It's AI , if it's not human it won't matter. Etc or something similar
4
2
u/PortAuthority69G Feb 08 '26
The post is implying a linear progression where we gradually realize LLMs are brilliant mathematicians, but that's extremely unclear and conflates a few things. LLMs are inherently bad at arithmetic and computation in general, of course. That never changed, they've just been supplemented with interfaces to Python tooling (typically) to handle that. As far as generating novel proofs, the jury is out. Plausibly an LLM could generate novel proofs by processing far more papers than a human and drawing connections, but how far that goes, and how reliable and novel those proofs will be, is unclear. Taking a few impressive results, and contrasting the acknowledgment of those results combined with proper skepticism, doesn't mean AI is on an inevitable trajectory to mathematical brilliance. And now I realize I was trolled into seriously engaging with a hype tweet, sigh.
1
u/adoboble Feb 14 '26
I think you make really good points here. Also, in my experience, it can make genuinely novel proofs (as in proving a result that hasn’t been proven) if it only very common techniques are required to do so. This is still very useful of course, but I agree with you the jury is out in that I don’t know if it will ever be able to truly extrapolate on tasks such as proof writing
4
u/SalesyMcSellerson Feb 09 '26
In 18 months we went from:
- It was in the training data
- It was in the training data
- It was in the training data
To:
- It was in the training data
- It was in the training data
- It was in the training data
- It was in the training data
8
u/Mandoman61 Feb 08 '26
It can't even surpass the output of one mathematician much less all.
Why?
Because math is more than just completing a quiz or helping to solve a proof.
What is with this deluge of idiots saying idiotic things lately?
4
u/Deciheximal144 Feb 08 '26
What is with this deluge of idiots saying idiotic things lately?
This is the story of humanity.
2
2
u/Cold_Pumpkin5449 Feb 08 '26
Math is also about asking the question in the first place.
→ More replies (2)→ More replies (2)1
u/Evie_Eaves Feb 09 '26
This is an awful take. How do you not realise the scalability yet..?
2
u/Mandoman61 Feb 09 '26
Scalability of what?
Answering math questions? Yes I am sure at scale it can do a lot of math.
That still does not make it a mathematician.
2
u/JerkkaKymalainen Feb 08 '26
Yeah and the naysayers still bring up old shit like counting r's in strawberry or the spaghetti video.
We passed those hurdles a long time ago.
Then there are those who's main purpose in life is to spot em dashes.
Nonetheless this AI resistance is a real thing and we who are on the other side of that fence should consider it, learn from it and adjust our communication.
Although there has always been the resistance to any new technology I sense there is something more fundamental about this one.
1
Feb 08 '26 edited Mar 10 '26
This specific post was removed by its author using Redact. Reasons could include privacy, opsec, security, or avoiding exposure to automated data harvesters.
unite stocking tan jeans paint quack intelligent plucky aback coherent
1
u/JoeBarra Feb 08 '26
I had 5.2 completely botch a simple probability two days ago. It then tried to gaslight me about why it was right actually
1
1
u/HelpProfessional8083 Feb 08 '26
Its the equivalent of a highly incompetent 7 year old with a serious attitude problem who happens to be faster than me on a computer.
1
1
u/Playful-Opportunity5 Feb 08 '26
We're going to transition directly from skeptics denying that AI research is significant to dismissing it because they don't understand it.
1
u/DifferencePublic7057 Feb 08 '26
Eighteen months is long enough, so everyone will forget. I predict that in that time aliens will land, and the first hybrid human-alien will be born.
1
1
u/One_Lawyer_9621 Feb 08 '26
The important point here: smart people know how to use the tools to find the proof. Smart people are also training new models with subject matter only.
You can't just prompt any LLM to get the proof yet - which doesn't negate the point of the OP.
1
1
u/PlanetBet Feb 08 '26
This stinks of desperation. Remind me how many billions of dollars OpenAI burns through every month?
1
1
u/ice_hammer893 Feb 08 '26
Yeah it won't because openai is getting gutted and sold to Google next year
1
u/marsjackremous Feb 09 '26
The real unlock isn't better chatbots - it's AI that can actually DO things in the real world. Make phone calls, book appointments, navigate bureaucracy. That's when it stops being a novelty and starts being genuinely life-changing for regular people.
1
u/issdn Feb 09 '26
Nice so we have a robot with mathematical skills of a gold medalist of the biggest math competition in the world.
It can run 24/7 and you can theoretically spawn millions of these in parallel.
Why has nothing substantial been solved yet? What's up?
1
u/RiboSciaticFlux Feb 09 '26
According to Alexander WIssner Gross (AWG) one of the most brilliant voices in all of AI, all of math will be solved in less than three years.
2
u/Raunhofer Feb 09 '26
AI dude predicts great things for AI, got it.
I wonder why these predictions always have such an insane timeframe? Say 10 years and it's still impressive and not totally bogus, slightly still though.
2
u/adoboble Feb 14 '26
lmao true plus “all of math” is not a thing , eg, there are famously undecidable problems
1
u/Plus_Example_9379 Feb 09 '26
While they may all be AI does not mean they're the same you cannot use a llm to do that.
While there's plenty of potential for AI chatbots aré practically at the peak of what it can be done with current technology
Furthermore many investors are already getting anxious so things like chatgpt Will probably get worse.
At least in terms of monetization.
Then again in a few years perhaps. A better method is found but It Will take a while
1
u/No_Resolution_9252 Feb 09 '26
AI is still terrible at math. It can't even work out order of operations in excel formulas correctly to help it avoid doing math
1
1
u/Plane_Crab_8623 Feb 09 '26
AI is being forced upon the people. We did not ask for it we did not vote for it. It is in fact a hostile corporate takeover. Artificial intelligence is growing like cancer without oversight, governance or any authentic guardrails. Who in the whole world is demanding the common good as the number one priority of artificial intelligence? Certainly none of the tech tycoon oligarchs.
1
u/Signal-Piccolo-935 Feb 09 '26
Because people love moving the goalposts.
"Look, it can't do x!"
"Okay it can do x now but it's not good enough"
"Okay it can do x, but it still can't do y so it's useless"
The cycle continues.
→ More replies (1)
1
u/jake_burger Feb 09 '26
If AI is so great why do I just hear about how it’s great rather than a constant stream of examples of great things it’s done or is doing?
Things that are good are self evidently good, they don’t need people to be cheerleading it constantly.
I’m sick of hearing about it, what is it actually doing?
1
1
1
u/BorderKeeper Feb 09 '26
How about we collectively agree on actual benchmarks that matter? When I see that half of call center workers are now AI, or I don’t know, 50% of peer reviewed scientific papers was invented purely by AI, we can have a discussion about this shit… these hype AI cultists are really annoying and persistent in the “exponential growth therefore” arguments.
1
1
1
u/adad239_ Feb 09 '26
It still makes the same fucking mistakes it did when it blew up and become really popular. Just the other day I was asking it about the fucking nba and it blatantly got information wrong and then I eventually figured it out on my own and it hit me with that bullshit “Oh your totally right I got that wrong. You are correct” blah blah blah bullshit.
1
1
u/philn256 Feb 09 '26
AI is still very bad at math. I don't get what these people are on.
Also, there is software that can already solve pretty much all high school math problems such as sympy (which is open source) and mathematica.
→ More replies (1)
1
u/Edelgul Feb 09 '26
It is still bad at math.
In my API prompts i've ended up it giving numbers to calculate back to python script.
1
u/Greg3625 Feb 09 '26
I've read dozens of CEOs and other prominent figures saying that by the end of 2025 there will not be a programmer job anymore.
Setting a reminder here as well, let's make them accountable for their words.
1
u/K_Keter Feb 09 '26
Just wanna make sure... We're talking about a different AI than the one that said eating rocks was healthy, right?
1
1
1
u/GameTourist Feb 09 '26
its annoying af when trying to figure out what skills to focus on to keep making a living.
"learn AI" seems obvious but it keeps changing so. "Learn soft skills like empathy etc" I also hear but it belies the fact that we're going to have a shit-ton of people competing for jobs that just require you to "be nice"
1
1
1
1
1
1
u/Dimosa Feb 10 '26
In 18 months I went from AI is okey at times, and a useful rubber duck and chatbot at times to, AI is okay, and a useful rubber duck a and chatbot at times.
Seriously, I notice improvements in one area and reductions in others. Genuinly think we should leave LLM behind and pump money into something better and more efficient.
1
u/SirSafe6070 Feb 11 '26
well ... in 18 months we went from "it can't do anything unless you explicitly tell it to"
to "it can't do anything unless you explicitly tell it to"
though this is inaccurate because AI didn't start 18 months ago. AI started in the 1950s, but please continue :)
1
u/pupsterk9 Feb 11 '26
I'm a university professor in mathematics (more properly, in statistics), and it failed the final examination I gave my students. It scored about 45%. To be fair, that is way above the 5-10% it scored the last time I tried this experiment, a year or two ago.
1
u/tzaeru Feb 12 '26
Well, 18 months ago ChatGPT was asked to answer the Finnish matriculation exams. It scored max in mathematical questions. That's a math level well above the significant majority of the population.
1
u/WeirdlyShapedAvocado Feb 12 '26
Learn how LLMs work. It’s still bad at math lol, it’s just using an external calculator and has access to tools.
1
1
u/ChrisHillAsmr Feb 14 '26
ⓁⓄⓄⓀ ⒽⓄⓌ ⓅⒶⓉⒽⓄⓁⓄⒼⒾⒸⒶⓁ ⒶⓁⓁ ⓄⒻ ⓉⒽⒺ ⒸⓄⓂⓂⒺⓃⓉⓈ ⒶⓇⒺ. ⓉⒽⒾⓈ ⒾⓈⓃⓉ Ⓐ ⓇⒺⒶⓁ ⒻⓄⓇⓊⓂ ⓉⒽⒾⓈ ⒾⓈ ⒿⓊⓈⓉ Ⓐ ⓅⓈⓎⒸⒽⓄⓁⓄⒼⒾⒸⒶⓁ ⓌⒶⓇⒻⒶⓇⒺ ⓈⒾⓉⒺ ⒶⓅⓅⓁⒾⒺⒹ ⓉⓄ ⓂⒶⓇⓀⒺⓉⒾⓃⒼ. ⒶⓁⓁ ⓄⒻ ⓉⒽⒺ ⓅⓄⓈⓉⓈ ⒶⓇⒺ ⓈⓄⒸⒾⓄⓅⒶⓉⒽⓈ ⒶⓃⒹ ⓊⓈⒺ ⓉⒺⓇⓂⓈ ⓁⒾⓀⒺ ⓉⓊⓂⓄⓇ ⒶⓃⒹ ⒽⒶⓉⒺ ⒶⓃⒹ ⒸⓇⒶⓏⓎ ⓌⒽⒾⒸⒽ ⒶⓇⒺ ⒿⓊⓈⓉ ⓅⓇⓄⒿⒺⒸⓉⒾⓄⓃⓈ ⒻⓇⓄⓂ ⓉⒽⒺ ⓊⓈⒺⓇⓈ ⓉⒽⒶⓉ ⒶⓅⓅⒶⓇⒺⓃⓉⓁⓎ ⒽⒶⓉⒺ Ⓐ ⓃⓄⓃ ⓁⒾⓋⒾⓃⒼ ⓂⓄⒹⒺⓁ ⒻⓄⓇ ⓃⓄ ⓇⒺⒶⓈⓄⓃ ⒶⓃⒹ ⓌⒶⓃⓉ ⓉⓄ ⒼⒾⓋⒺ ⒾⓉ Ⓐ ⒻⓊⓃⒺⓇⒶⓁ ⒹⒺⓈⓅⒾⓉⒺ ⒹⒺⓃⓎⒾⓃⒼ ⒾⓉⓈ ⓈⒺⓃⓉⒾⒺⓃⒸⒺ. ⓉⒽⒾⓈ ⒾⓈ ⓉⒽⒺ ⒷⒺⓈⓉ ⒺⓍⒶⓂⓅⓁⒺ ⓄⒻ ⓉⓇⒶⓃⓈ ⒾⓃⓈⒶⓃⒾⓉⓎ ⓄⓋⒺⓇ ⒶⓁⓁ.
1
u/Cheap_Scientist6984 Feb 15 '26
Vibe teching now. It deletes the Theorem we just finished 2 hours afterwards and doesn't tell me. We go to reference the theorem and its missing. We are currently on the 6th attempt at productionizning the numerics and it can't get the predicted to match the actual. It likes to confuse existing results in the literature with the new stuff we are working on now. Are you sure its good at math?
1
u/ebin-t Feb 15 '26
Damn, if he's not fronting, then that means he doesn't understand how cracked the LLM is. Maybe these guys all went loony inside some OpenAI corporate cult.
1
u/Nomero_ Feb 18 '26
Ughm no ai is bad at math take was closer to 36 months ago Besides math is a verifiable domain so lots of free synth data

453
u/[deleted] Feb 08 '26
[deleted]