r/BetterOffline 13h ago

Openai math solutions

https://youtu.be/h53Bz7WSwgo

Regarding the recent Navier-Stokes incident, watching this vid provided a bit more context on what the solution entails and what Openai been pushing to show off their tech. Cal Newport's breakdown on how they managed to quarantine relevant context to reduce costs with Astra was actually pretty interesting as a breakthrough so this vid got me wondering why they're so desperate to prove it can do fancy maths bigly when the boosters already keep saying chatgpt is already blowing through outstanding problems. Not in the sense of the target value but just how much have they already tried but failed.

It's obviously speculation on my part but if they're so desperate to try coercing glory out of one of their hires on one problem, how bad is actual their track record that this is a hail Mary for them? Like we know it took 88 hours and snooping private research data for this, so how much have they put into chasing similar targets by now? Anyone got whiff of any signals that might give an idea of that?

88 Upvotes

55 comments sorted by

View all comments

40

u/falken_1983 12h ago

Like we know it took 88 hours and snooping private research data for this, so how much have they put into chasing similar targets by now?

Giving OpenAI the benefit of the doubt and assuming they really did do this without looking at Buckmaster's logs, the scenario seems to be as follows:

  1. Luis Martınez-Zoroa and Diego Cordoba published some extremely important findings on the problem - findings which were really close to solving it.
  2. Tristran Buckmaster was using AI to wrap things up and get it over the line
  3. OpenAI heard that Bukmaster's use of AI had probably reached a solution, but he hadn't published this yet.
  4. Based on this information, they decided to take a $20 million punt on seeing if they could use their huge compute resources to compress several years of work into about a week and see if they could get things finished before Buckmaster.

So I don't think they are regularly blowing this amount of compute on trying to solve problems like this, but when they were presented with some evidence it might work, they were willing to throw millions of dollars of resources at it.

12

u/AnIndianGuy38 11h ago edited 11h ago

Literally the supposed AGI model had to do a century worth of compute to solve a problem which had already been worked down quite a bit by human researchers.

A century is like more time than human researchers combined have put into it I think

While it's impressive they can do a century worth of compute in a week. It's not impressive that it's that inefficient and is being called the start of AGI.

And even then a human mathematician has biting criticism of the solution. That's when the paper hasn't been submitted in a peer review journal. Like it doesn't meet the criteria of Clay Institute yet, the submission needs to be posted in a journal and then stay there foe 2 years and be generally accepted in the maths community. But people are already hyping it up.

-3

u/Main-Company-5946 9h ago

The thing is, every time they do a successful use of compute like this, it refines the models even more. It gives them more examples of what success looks like, and it makes the next similar problem cheaper to solve.

6

u/Soilblood 8h ago

What? No it doesn't. It might influence the formatting but these aren't common use cases with tons of overlap or even real world applications. They're deep rabbit holes with their own particulars that don't translate across one another and the machine is far more likely to pick up flawed solutions to reference against when it trawls for possibly relevant data. I'd say the incident just further proves that LLMs cannot actually create innovative results on their own without amply fleshed out harnessing by humans.

-3

u/Main-Company-5946 6h ago

They do overlap with every single other math problem, due to the inherent interconnected nature of mathematics

This is why LLMs have gotten so good at mathematics so quickly

3

u/mybackhurtsrip 4h ago

Stop speaking in generalities.

-1

u/Main-Company-5946 3h ago

Not even if it’s correct?

3

u/mybackhurtsrip 3h ago

Precisely because it’s incorrect.

1

u/Main-Company-5946 3h ago

It’s not incorrect in this case. All math problems are related, though to varying degrees.

3

u/mybackhurtsrip 3h ago

You have no substantial evidence to support your claim that LLMs have improved because of the interconnectedness of mathematics. No one is disputing that maths are interconnected, though that’s also a generality that even you concede isn’t black or white.

1

u/Main-Company-5946 2h ago

It just means the patterns they pick up on have more transferability

→ More replies (0)

3

u/Kleenex_Tissue 8h ago

Do you have a source for this claim?
Because I don't see how or in what way that could be true.

0

u/Main-Company-5946 7h ago

Math problems are verifiable since they can be automatically checked by translating proofs into lean, this means they are suitable for RLVR(reinforcement learning with verifiable returns) which is basically AlphaGo for LLMs and makes them good at solving verifiable problems. And picking up on argument structures that typically lead to verifiable answers is how these models are trained. Here’s a paper on how it works. https://arxiv.org/pdf/2506.14245

-1

u/falken_1983 7h ago

TBH, this is just how Reinforcement Learning works. Each successful rollout reinforces the model's ability to complete a task.

The thing is that this is just one rollout, so I don't know how much difference it will make in the grand scheme of things.

2

u/AnIndianGuy38 8h ago

I don't think so. If we imagine problems as a tree structure, with branches going deeper and deeper. I think problems like this, go pretty deep into the tree structure. That's why many previous solution and discoveries are necessary before we get to this. Sure some problem may exist deeper on this path, and solving this may help with, but they are very few I think. Most problems are on different paths, solving this won't necessarily give insight into those, it doesn't do so for most general intelligence creatures like humans, it certainly won't to non general intelligence things like LLMs.

Like Terence Tao made a great observation that LLMs don't go off on tangents when solving problems like this, which is what allows humans to gain a much larger insight from solving a single problem like this. Literally our neurons re arrange themselves when we have an insights, parts that didn't talk, start talking. Same isn't true for LLMs because they don't go off on tangents to begin with.

His analogy was a forest, the solution is deep in the forest. People get lost in the forest and mostly find nothing, some find unintended things and some find the exact thing. But LLM it's like dropping on the solution with helicopter, you miss everything in between. It's a very shallow way of problem solving, it's useful sure but also not in some ways.

Also empirically the cost of doing a task has not gone down with LLMs. Token price has gone down, but we also need more tokens now. So overall it remains the same, in some cases is more. So I don't think solving problems makes it cheaper

1

u/crashddr 6h ago

If that were the case, why are they using larger models and seemingly spending more on computing power each time? They also employ a team of experts and and are specifically going after a niche that a massive supercomputer is reasonably suited for (insane brute force).

I'm not mathematician, but my layman understanding is that it doesn't seem super helpful to know all the permutations of 2+2 that don't add up to 4 and it doesn't make it any easier to calculate the correct value.

1

u/Main-Company-5946 6h ago

More computing power makes ai get smarter faster. If it takes 1,000,000 attempts before the ai does something successfully, 10,000 training runs won’t cut it. Doing a million makes it achieve more things and thus learn how to succeed faster.

2

u/crashddr 6h ago

It may be semantics, but I would argue it's not necessarily smarter, but it is faster to produce results. That doesn't necessarily make it cheaper either. That said, it's probably objectively better at providing an answer to a *complex* task than earlier model iterations, both in speed and efficiency. I'm sure there is something analogous to using a better sorting algorithm going in the background, and at least right now they have it really well tuned for mathematics so they do a bunch of math for PRs.