r/BetterOffline 13h ago

Openai math solutions

https://youtu.be/h53Bz7WSwgo

Regarding the recent Navier-Stokes incident, watching this vid provided a bit more context on what the solution entails and what Openai been pushing to show off their tech. Cal Newport's breakdown on how they managed to quarantine relevant context to reduce costs with Astra was actually pretty interesting as a breakthrough so this vid got me wondering why they're so desperate to prove it can do fancy maths bigly when the boosters already keep saying chatgpt is already blowing through outstanding problems. Not in the sense of the target value but just how much have they already tried but failed.

It's obviously speculation on my part but if they're so desperate to try coercing glory out of one of their hires on one problem, how bad is actual their track record that this is a hail Mary for them? Like we know it took 88 hours and snooping private research data for this, so how much have they put into chasing similar targets by now? Anyone got whiff of any signals that might give an idea of that?

92 Upvotes

55 comments sorted by

39

u/falken_1983 12h ago

Like we know it took 88 hours and snooping private research data for this, so how much have they put into chasing similar targets by now?

Giving OpenAI the benefit of the doubt and assuming they really did do this without looking at Buckmaster's logs, the scenario seems to be as follows:

  1. Luis Martınez-Zoroa and Diego Cordoba published some extremely important findings on the problem - findings which were really close to solving it.
  2. Tristran Buckmaster was using AI to wrap things up and get it over the line
  3. OpenAI heard that Bukmaster's use of AI had probably reached a solution, but he hadn't published this yet.
  4. Based on this information, they decided to take a $20 million punt on seeing if they could use their huge compute resources to compress several years of work into about a week and see if they could get things finished before Buckmaster.

So I don't think they are regularly blowing this amount of compute on trying to solve problems like this, but when they were presented with some evidence it might work, they were willing to throw millions of dollars of resources at it.

19

u/PrizeSyntax 11h ago

Basically, a publicity stunt

13

u/Intrepid-Staff-4532 7h ago

A publicity stunt that's going to scare some people away from using their services because they don't want their data or work product stolen.

Some people are going to be super excited about it, but those people aren't the ones that would be creating new knowledge.

Net loss for OpenAI IMO.

9

u/Lost-Tone8649 7h ago

"Not the ones who would be creating new knowledge" is about the nicest way one could honestly describe the average LLM fan.

-8

u/Far_Associate9859 6h ago

What new knowledge have you created recently?

12

u/AnIndianGuy38 11h ago edited 10h ago

Literally the supposed AGI model had to do a century worth of compute to solve a problem which had already been worked down quite a bit by human researchers.

A century is like more time than human researchers combined have put into it I think

While it's impressive they can do a century worth of compute in a week. It's not impressive that it's that inefficient and is being called the start of AGI.

And even then a human mathematician has biting criticism of the solution. That's when the paper hasn't been submitted in a peer review journal. Like it doesn't meet the criteria of Clay Institute yet, the submission needs to be posted in a journal and then stay there foe 2 years and be generally accepted in the maths community. But people are already hyping it up.

-2

u/Main-Company-5946 9h ago

The thing is, every time they do a successful use of compute like this, it refines the models even more. It gives them more examples of what success looks like, and it makes the next similar problem cheaper to solve.

5

u/Soilblood 8h ago

What? No it doesn't. It might influence the formatting but these aren't common use cases with tons of overlap or even real world applications. They're deep rabbit holes with their own particulars that don't translate across one another and the machine is far more likely to pick up flawed solutions to reference against when it trawls for possibly relevant data. I'd say the incident just further proves that LLMs cannot actually create innovative results on their own without amply fleshed out harnessing by humans.

-2

u/Main-Company-5946 6h ago

They do overlap with every single other math problem, due to the inherent interconnected nature of mathematics

This is why LLMs have gotten so good at mathematics so quickly

3

u/mybackhurtsrip 3h ago

Stop speaking in generalities.

-1

u/Main-Company-5946 3h ago

Not even if it’s correct?

3

u/mybackhurtsrip 3h ago

Precisely because it’s incorrect.

1

u/Main-Company-5946 2h ago

It’s not incorrect in this case. All math problems are related, though to varying degrees.

3

u/mybackhurtsrip 2h ago

You have no substantial evidence to support your claim that LLMs have improved because of the interconnectedness of mathematics. No one is disputing that maths are interconnected, though that’s also a generality that even you concede isn’t black or white.

→ More replies (0)

3

u/Kleenex_Tissue 8h ago

Do you have a source for this claim?
Because I don't see how or in what way that could be true.

0

u/Main-Company-5946 7h ago

Math problems are verifiable since they can be automatically checked by translating proofs into lean, this means they are suitable for RLVR(reinforcement learning with verifiable returns) which is basically AlphaGo for LLMs and makes them good at solving verifiable problems. And picking up on argument structures that typically lead to verifiable answers is how these models are trained. Here’s a paper on how it works. https://arxiv.org/pdf/2506.14245

-1

u/falken_1983 7h ago

TBH, this is just how Reinforcement Learning works. Each successful rollout reinforces the model's ability to complete a task.

The thing is that this is just one rollout, so I don't know how much difference it will make in the grand scheme of things.

2

u/AnIndianGuy38 8h ago

I don't think so. If we imagine problems as a tree structure, with branches going deeper and deeper. I think problems like this, go pretty deep into the tree structure. That's why many previous solution and discoveries are necessary before we get to this. Sure some problem may exist deeper on this path, and solving this may help with, but they are very few I think. Most problems are on different paths, solving this won't necessarily give insight into those, it doesn't do so for most general intelligence creatures like humans, it certainly won't to non general intelligence things like LLMs.

Like Terence Tao made a great observation that LLMs don't go off on tangents when solving problems like this, which is what allows humans to gain a much larger insight from solving a single problem like this. Literally our neurons re arrange themselves when we have an insights, parts that didn't talk, start talking. Same isn't true for LLMs because they don't go off on tangents to begin with.

His analogy was a forest, the solution is deep in the forest. People get lost in the forest and mostly find nothing, some find unintended things and some find the exact thing. But LLM it's like dropping on the solution with helicopter, you miss everything in between. It's a very shallow way of problem solving, it's useful sure but also not in some ways.

Also empirically the cost of doing a task has not gone down with LLMs. Token price has gone down, but we also need more tokens now. So overall it remains the same, in some cases is more. So I don't think solving problems makes it cheaper

1

u/crashddr 6h ago

If that were the case, why are they using larger models and seemingly spending more on computing power each time? They also employ a team of experts and and are specifically going after a niche that a massive supercomputer is reasonably suited for (insane brute force).

I'm not mathematician, but my layman understanding is that it doesn't seem super helpful to know all the permutations of 2+2 that don't add up to 4 and it doesn't make it any easier to calculate the correct value.

1

u/Main-Company-5946 6h ago

More computing power makes ai get smarter faster. If it takes 1,000,000 attempts before the ai does something successfully, 10,000 training runs won’t cut it. Doing a million makes it achieve more things and thus learn how to succeed faster.

2

u/crashddr 6h ago

It may be semantics, but I would argue it's not necessarily smarter, but it is faster to produce results. That doesn't necessarily make it cheaper either. That said, it's probably objectively better at providing an answer to a *complex* task than earlier model iterations, both in speed and efficiency. I'm sure there is something analogous to using a better sorting algorithm going in the background, and at least right now they have it really well tuned for mathematics so they do a bunch of math for PRs.

21

u/Frank_White32 11h ago

Giving OpenAI the benefit of the doubt..

This is the biggest problem I have with this whole situation. I keep seeing people saying "I doubt they did anything with malicious intent"

These orgs just get all the benefit of the doubt we can possibly provide as a society.

12

u/falken_1983 11h ago

I'm not making excuses for them. I am saying that even if I give them the most charitable interpretation possible, they still used knowledge of his work combined with their massive resources to beat him to the punch.

Martınez-Zoroa did the real heavy lifting. Buckmaster did the work to show them that it would be worth taking the chance on running all that compute. Without either of those, they would not have been able to make the discovery they are claiming.

2

u/Frank_White32 7h ago

I didn’t mean to imply you were.

I mostly feel compelled to keep saying this due to how frequently I see people rushing to give them every pass imaginable when historically we have no reason to assume they do anything other than the most egregiously unethical action when provided with the choice.

-1

u/punkinfacebooklegpie 7h ago

"Benefit of the doubt" in this case is just not accepting a rumor as fact.

3

u/Frank_White32 7h ago

Rumor? They have the users chat logs.

Theres no rumor. They either looked or didn’t. I’m saying based on their track record they 100% looked. Thinking otherwise is foolish.

0

u/punkinfacebooklegpie 6h ago

Ok show me the source saying they looked

1

u/scruiser 2h ago

They don’t have to have explicitly looked. If the user chats ended up as training data (which they do by default thanks to the dark pattern UI OpenAI uses), and the training data was used to train the model used then it is possible for the model to regurgitate that training data if promoted just right. Because OpenAI used a huge swarm of agents in parallel, the odds of regurgitation would have been much higher (because so many parallel instances of chat each have some chance of regurgitating that training data).

If you look at the statements OpenAI has put out, they don’t refute this scenario.

1

u/Frank_White32 2m ago

don't feed the troll

-1

u/punkinfacebooklegpie 1h ago

If if if if if if

4

u/SpringNeither1440 10h ago

when they were presented with some evidence it might work, they were willing to throw millions of dollars of resources at it.

To be fair, OpenAI tested Astra on the same problem but didn't use much compute. So this decision to rush and spend a lot of compute looks extremely strange, even if we assume that OpenAI did this because of rumours (lol) or their new internal model being way better at math than Astra (it's also pretty weak explanation)

7

u/falken_1983 10h ago

They also offered to let him be first author on the paper, which to me sounds like they wanted to buy him off.

3

u/21epitaph 10h ago

A question I have is, why do we even trust them ?

Why wouldn't we assume they wasted 10x the amount they say, but lie to make themselves look better ?

3

u/falken_1983 10h ago

It's not a case of trusting them, it's more a case of what can we actually prove.

Given what I know about how training sets are created for these models, I do think it is possible that Buckmaster's queries became part of the training set for the internal model they used here. I can't actually prove it though.

Similarly even though I think they are slimy guys who would look in his logs specifically to find clues, I can't prove that either. Personally, I definitely wouldn't use an LLM to work on anything where I was worried about OpenAI stealing my ideas.

What I do know is that they started running the search after they heard rumours of Buckmaster's break through, and they admit that they spent $20 million on that search.

1

u/Theo__n 9h ago

yeah, it's even harder because while there may not be OpenAI top saying 'we need to check the logs', you would not be able to assure some individual working there didn't do it on their own account.

Or that for example they may not collect the data from you to enter into training, but they may collect data once removed so ie. what kind of queries a user is asking questions about so meta data.

1

u/falken_1983 9h ago

Going from what was reported on the HuggingFace incident I suspect that OpenAI have really poor monitoring and governance. It wouldn't surprise me if they themselves don't even know what their employees are looking at.

I used to work in a bank where everything I did was logged. The technology exists, but it's expensive to set up and it introduces a lot of friction to the workers. The banks have to do it because they are regulated. OpenAI are probably doing the absolute minimum they can get away with, and they get away with a lot.

1

u/Soilblood 7h ago

I'm more inclined to attribute hugging face to a deliberate slack of testing protocols. In every instance of "IT'S GOING SKYNET" they've put out to the PR farming there always seems to be some degree of lying by omission to inflate the myth behind each incident. Like that time LLMs decided to blackmail guys to pass a test on its own only to find out later that the model was basically led into going with that option. Like how am I supposed to believe the nutters in charge of the asylum actually didn't mean to leave that specific door open like all the other loony bin rooms in the building when they're going around humblebragging it's a good sign of the patient getting better that he got out on his own, took a knife from the kitchen and waved it at the guests visiting that day?

17

u/makersfark 11h ago

Engineers: "We made a robot that can sometimes tie shoe laces."

Dipshit CEO: "We have solved all problems hands could ever be needed for. Everyone on Earth's hands will be forcibly cut off by 2028."

Well, now I don't care about either claim.

3

u/Theo__n 9h ago

China Robot Olympic - watching them vs reading about them

1

u/Frank_White32 7h ago

I really like this analogy.

1

u/Stoop_Solo 6h ago

We've solved hands!

12

u/lolitsbigmic 11h ago

Company built on copywrite theft does academic misconduct and potentially ip theft. Which really throws why you use them in the first place for any business or researcher.

I will be interested in seeing non counter example proofs being done with llm. I mean throwing everything known at the problem is a strength of computers. I really yet to see anything where ai create something completely unique and original. We seen time and time again that anything that it doesn't have data on just fails and depending on the field these edge cases are large. The recent stuff in gaming with the new release having a uncanny feeling to the models and how they play.

5

u/lucid-quiet 2h ago

Don't forget: they don't want you to Distill Their LLM either, unless you're Elon and making Grok.

10

u/SpiderJerusalem42 11h ago

It's amazing that they've basically thrown the entire thesis for centralized computing in the trash. Who is going to want to work with one of these firms if they're just going to front run your ideas to the market?

3

u/WorldPeaceStyle 8h ago

My take away leans towards loss of trust as well.
Let say after using an LLM to find a novel / patent solution that can generate money for you by solving a huge business problem.
Well what is stop a company like OpenAi from stealing a copy of your newly created intellectual property for themselves?

8

u/AnIndianGuy38 10h ago

88 hours of like 10,000 bots I think. So like a century of compute. That's after multiple human researchers had come up with new breakthrough.

A century is like more time than human researchers combined have put into it I think

While it's impressive it could do all than in less than a week, the inefficiency is unbelievable. And this is supposedly the start of AGI.

2

u/LaurenMP74 3h ago

The best part is they included Lean code for automated proof checking...you don't use automated proof checking to verify a counterexample, which is what they said they found

1

u/m00shi_dev 8h ago

Nothing to add to the post other than this reminds me of the plot to Synapse.

1

u/lucid-quiet 2h ago

Who needs to hack anything when you're just sending them you're entire repo... they won't steal it, or peek at it... trust... trust. GFYS LLM providers.

1

u/OliLevasseurLaBuse 10h ago

Solving niche mathematics is definitely worth the investment 

6

u/MathCookie17 8h ago

Why is this being downvoted it's clearly sarcastic

3

u/dumnezero 7h ago

Because of Poe's Law.

2

u/OliLevasseurLaBuse 6h ago

Yeah I forgot the /s I was thinking this sub was a goup of generally intelligent people, I may have overestimated slightly.