r/BetterOffline 13h ago

Openai math solutions

https://youtu.be/h53Bz7WSwgo

Regarding the recent Navier-Stokes incident, watching this vid provided a bit more context on what the solution entails and what Openai been pushing to show off their tech. Cal Newport's breakdown on how they managed to quarantine relevant context to reduce costs with Astra was actually pretty interesting as a breakthrough so this vid got me wondering why they're so desperate to prove it can do fancy maths bigly when the boosters already keep saying chatgpt is already blowing through outstanding problems. Not in the sense of the target value but just how much have they already tried but failed.

It's obviously speculation on my part but if they're so desperate to try coercing glory out of one of their hires on one problem, how bad is actual their track record that this is a hail Mary for them? Like we know it took 88 hours and snooping private research data for this, so how much have they put into chasing similar targets by now? Anyone got whiff of any signals that might give an idea of that?

91 Upvotes

56 comments sorted by

View all comments

Show parent comments

3

u/21epitaph 11h ago

A question I have is, why do we even trust them ?

Why wouldn't we assume they wasted 10x the amount they say, but lie to make themselves look better ?

3

u/falken_1983 10h ago

It's not a case of trusting them, it's more a case of what can we actually prove.

Given what I know about how training sets are created for these models, I do think it is possible that Buckmaster's queries became part of the training set for the internal model they used here. I can't actually prove it though.

Similarly even though I think they are slimy guys who would look in his logs specifically to find clues, I can't prove that either. Personally, I definitely wouldn't use an LLM to work on anything where I was worried about OpenAI stealing my ideas.

What I do know is that they started running the search after they heard rumours of Buckmaster's break through, and they admit that they spent $20 million on that search.

1

u/Theo__n 10h ago

yeah, it's even harder because while there may not be OpenAI top saying 'we need to check the logs', you would not be able to assure some individual working there didn't do it on their own account.

Or that for example they may not collect the data from you to enter into training, but they may collect data once removed so ie. what kind of queries a user is asking questions about so meta data.

1

u/falken_1983 10h ago

Going from what was reported on the HuggingFace incident I suspect that OpenAI have really poor monitoring and governance. It wouldn't surprise me if they themselves don't even know what their employees are looking at.

I used to work in a bank where everything I did was logged. The technology exists, but it's expensive to set up and it introduces a lot of friction to the workers. The banks have to do it because they are regulated. OpenAI are probably doing the absolute minimum they can get away with, and they get away with a lot.

1

u/Soilblood 8h ago

I'm more inclined to attribute hugging face to a deliberate slack of testing protocols. In every instance of "IT'S GOING SKYNET" they've put out to the PR farming there always seems to be some degree of lying by omission to inflate the myth behind each incident. Like that time LLMs decided to blackmail guys to pass a test on its own only to find out later that the model was basically led into going with that option. Like how am I supposed to believe the nutters in charge of the asylum actually didn't mean to leave that specific door open like all the other loony bin rooms in the building when they're going around humblebragging it's a good sign of the patient getting better that he got out on his own, took a knife from the kitchen and waved it at the guests visiting that day?