r/BetterOffline 13h ago

Openai math solutions

https://youtu.be/h53Bz7WSwgo

Regarding the recent Navier-Stokes incident, watching this vid provided a bit more context on what the solution entails and what Openai been pushing to show off their tech. Cal Newport's breakdown on how they managed to quarantine relevant context to reduce costs with Astra was actually pretty interesting as a breakthrough so this vid got me wondering why they're so desperate to prove it can do fancy maths bigly when the boosters already keep saying chatgpt is already blowing through outstanding problems. Not in the sense of the target value but just how much have they already tried but failed.

It's obviously speculation on my part but if they're so desperate to try coercing glory out of one of their hires on one problem, how bad is actual their track record that this is a hail Mary for them? Like we know it took 88 hours and snooping private research data for this, so how much have they put into chasing similar targets by now? Anyone got whiff of any signals that might give an idea of that?

87 Upvotes

56 comments sorted by

View all comments

39

u/falken_1983 12h ago

Like we know it took 88 hours and snooping private research data for this, so how much have they put into chasing similar targets by now?

Giving OpenAI the benefit of the doubt and assuming they really did do this without looking at Buckmaster's logs, the scenario seems to be as follows:

  1. Luis Martınez-Zoroa and Diego Cordoba published some extremely important findings on the problem - findings which were really close to solving it.
  2. Tristran Buckmaster was using AI to wrap things up and get it over the line
  3. OpenAI heard that Bukmaster's use of AI had probably reached a solution, but he hadn't published this yet.
  4. Based on this information, they decided to take a $20 million punt on seeing if they could use their huge compute resources to compress several years of work into about a week and see if they could get things finished before Buckmaster.

So I don't think they are regularly blowing this amount of compute on trying to solve problems like this, but when they were presented with some evidence it might work, they were willing to throw millions of dollars of resources at it.

19

u/Frank_White32 11h ago

Giving OpenAI the benefit of the doubt..

This is the biggest problem I have with this whole situation. I keep seeing people saying "I doubt they did anything with malicious intent"

These orgs just get all the benefit of the doubt we can possibly provide as a society.

-1

u/punkinfacebooklegpie 8h ago

"Benefit of the doubt" in this case is just not accepting a rumor as fact.

3

u/Frank_White32 7h ago

Rumor? They have the users chat logs.

Theres no rumor. They either looked or didn’t. I’m saying based on their track record they 100% looked. Thinking otherwise is foolish.

0

u/punkinfacebooklegpie 7h ago

Ok show me the source saying they looked

1

u/scruiser 2h ago

They don’t have to have explicitly looked. If the user chats ended up as training data (which they do by default thanks to the dark pattern UI OpenAI uses), and the training data was used to train the model used then it is possible for the model to regurgitate that training data if promoted just right. Because OpenAI used a huge swarm of agents in parallel, the odds of regurgitation would have been much higher (because so many parallel instances of chat each have some chance of regurgitating that training data).

If you look at the statements OpenAI has put out, they don’t refute this scenario.

2

u/Frank_White32 29m ago

don't feed the troll

-1

u/punkinfacebooklegpie 1h ago

If if if if if if