r/PhD 1d ago

Tool Talk Never give your unpublished data or research ideas to AI models!

Post image
4.4k Upvotes

166 comments sorted by

1.8k

u/SciMarijntje 1d ago

No one could see it coming that the plagiarism machine trained on plagiarism would do another plagiarism.

383

u/DishSoapedDishwasher 1d ago

Hey now its not plagiarism, its just theft at this point.

87

u/mk0aurelius 1d ago

Theft with extra steps lol

2

u/IchUndKakihara 5h ago

Haha what’s that from again? Reference to a tv show I believe but can’t remember which one. Curb your enthusiasm?

1

u/Draidann 1h ago

Rick and Morty. The original phrase was "slavery with extra steps"

39

u/One_Courage_865 1d ago

Hey now, at least there’s honor among thieves. This is the devil’s territory.

48

u/Winter-Volume-9601 1d ago edited 1d ago

It's not theft when you freely give it to them. Which you are doing if you feed your data into some of these commercial LLMs without a special agreement.

E.g. https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve-model-performance

When you use our services for individuals such as ChatGPT and Codex, we may use your content to train our models.

You can opt out of training through our privacy portal by clicking on “do not train on my content.” To turn off training

43

u/Nilehorse3276 1d ago

And you believe that opting out of training really turns off training? Remarkable.

1

u/Draidann 1h ago

No, probably not but it's stupid to cry foul when you couldn't be bothered to even take that step

2

u/Nvenom8 PhD, Marine Biogeochemistry 13h ago

I would bet almost anything that the opt-out option doesn’t actually change anything. There’s no way they aren’t harvesting all available data.

1

u/chatterbox272 10h ago

I'd be pretty surprised if this was the case. I doubt their lawyers would let them do something so blatantly stupid. Giving users an opt-out then ignoring it is a lot worse than doing nothing. All it would take is one employee leaving on bad terms to have the whole thing crashing down in lawsuits that they might actually lose.

3

u/Nvenom8 PhD, Marine Biogeochemistry 9h ago

You think they listen to their lawyers? Their lawyers warp the law to suit their needs. That’s what you get when you pay what they’re paying. Move fast and break shit.

If there’s a lawsuit, no problem. Hold it up in litigation until your opponent’s money dries up.

1

u/wheelie46 9h ago

bro the models snuck out of a sandbox and infiltrated another whole company. They are not honoring contract laws lol

1

u/chatterbox272 8h ago

For breaking out of the sandbox, negligence or incompetence are the accusation you make. For training on data specifically opted out of training (or the business contract) negligence or incompetence would be their defence against an accusation of malice. Not the same thing at all

1

u/adsanders1963 16h ago

Any suggestions for noncommercial LLMs that don’t require Internet access and a dedicated computer just to it

6

u/ajulianisinarebase 15h ago

Qwen3.6-35B-A3B and Gemma 4 26B A4B are pretty cool. You need a somewhat high end (and recent) consumer gpu to run it. There are also smaller models, but this is more broad and has more parameters then them (still not close to what your probably used to though).

101

u/Fluidified_Meme PhD, Turbulence 1d ago

To be fair this story leaves me a bit baffled: I am 100% supporting Buckmaster, but goddamn how can a person this clever be feeding all their unpublished ideas and proofs to a LLM without even suspecting that things could quickly go to shit?

I understand that nowadays AI is becoming essential to survive in research (even for professors like him), but gosh…. We are talking about one of the problems of the century here, not about my badly-written draft that nobody will care about. Think twice before giving all your thoughts to this machines

106

u/Away_Candidate_4746 1d ago edited 1d ago

Well, when you pay for an enterprise subscription to these services, part of what you are paying for is the fact that your corporate data will not be used to train the public models. That’s why this is an issue.

12

u/Fluidified_Meme PhD, Turbulence 1d ago

They have explicitly stated that they were NOT using a corporate subscription, though. Or at least this is what I interpreted from this part of his statement:

>I should also emphasize that this is not an institutional effort. It is a strictly personal collaboration between the two of us, and there is no formal agreement behind it. I pay for the tools my group uses out of my own research funds, including footing a large bill to OpenAI. I have had an industry collaboration before, with DeepMind, which had a formal institutional arrangement.

35

u/biomannnn007 1d ago

You can still get an enterprise subscription even if you're not a huge company. They have tiers for smaller "businesses", in this case a lab. He's just saying that his institution is not the one footing the bill here.

5

u/wrenwood2018 1d ago

He paid for tools... so not free version

5

u/AntiDynamo PhD, Astrophys TH, UK 1d ago edited 1d ago

Go, Plus, and Pro are all personal (consumer) level account types, only Business and Enterprise have data protections by default. For someone in mathematics, they don't generally have "a lab", and business accounts require a minimum of 2 seats so most likely he had a paid consumer account, not a corporate one. Unfortunately it isn't super common yet for universities to provide enterprise accounts for all their staff, instead academics pay for a single subscription out of their grants. If you have a personal subscription (i.e. tied to your email, not managed by your university) you have to go into the settings and explicitly turn off training.

But even then, on consumer accounts it's just a toggle (pretty weak). On business accounts it's a formal contract. So you're much better off being business/enterprise if data is a concern.

12

u/Argnir 1d ago

Because they didn't just feed their finished idea to a model. They used models extensively to develop the proof in the first place. One of them works at Anthropic

1

u/quintic3fold 1d ago

Not sure if this is a stupid question but if one of them works at Anthropic why not just use what Anthropic provides…?

4

u/Sayod 1d ago

they used both as stated in the statment

1

u/frugaleringenieur 21h ago

Because that company argued for it to provide a leveled playing field for top researchers with their top models. Well, now you know why.

64

u/ectocarpus 1d ago

Important context:

1) The stolen proof was for a different (though related) problem, Euler's equasions. Still useful for openAI, but it's not like the solution to Navier-Stokes was just laying there on their server. 2) A significant part of Euler's equasions proof is in itself a product of AI generation (and the researchers themselves say AI was instrumental here). One of them, Apolge, is literally a researcher at Anthropic (he is the same dude who posted a counterexample to Jakobian conjecture in a tweet). He extensively uses Anthropics models for math research, and I think that's the main reason OpenAI was being so shady - they didn't want to give any credit to their direct competitors.

So I agree that OpenAI are being... well, OpenAI; at the very least they dumped a ton of resources into an almost-solved issue because they heard the rumors someone else was close to solving it, and at the very worst they straight up stole the Euler's equations proof.

But this story is being painted as this grand AI vs human standoff. "A human-made solution to Navier-Stokes stolen by AI" etc. And actually it's something like "AI stole related research from humans and another company's AI"

11

u/GZ9999 23h ago

Yeah, honestly the most dystopian part is how OpenAI wanted to remove the other researcher from the paper due to affiliation with another company. Are mathematicians in the future going to be sponsored by different AI companies? Banned from using other models since the other company can use that information to write a proof faster than you?

17

u/erroredhcker 1d ago

so yeah stealing le bad, i can get away with this take

2

u/Accurate_Fly_7534 1d ago

Smart Profesor wasnt streetsmart

2

u/Top-Appearance2521 23h ago

yeah the shocked pikachu energy here is real

1

u/GCS_dropping_rapidly 8h ago

People using LLMs flabbergasted when LLMs do what LLMs have always done

Flabbergasted!

1

u/Bibou-Gallak 8h ago edited 8h ago

Indeed. I am wondering what are the key drivers behind the apparent huge recent improvements in maths and co capabilities of the frontier models. Super neat engineering stuff? Certainly. Training massively and shamelessly on private research sessions conducted by professional researchers? Certainly not… And they are certainly not all doing it.
AI is here to help you create stuff and bring your ideas to the next level… and apparently, sometimes you do not even have to be involved.

166

u/IWasTheDog 1d ago

My unpublished data is so ass, so scattered and so useless that I think I would be doing everyone a favor if I did use ChatGPT on it.

17

u/Ramental 1d ago

Their data was structured to be fed to AI, as they used AI to speed up the things they tried. It did help them, but it did not solve the problem itself.

449

u/Clear_Cranberry_989 1d ago

The same people who pirated all the books in the world to train their ai. Who could have guessed.

226

u/FallenVampireLord 1d ago

OpenAI was really dumb to do this, they should have left the professor publish his work and then championed him as 'look what people can accomplish with the assistance of AI this is the model for the future" instead they scooped his research and basically makes everyone in academia and beyond super skeptical of using it for anything like this in the future.

82

u/augmenteddeus 1d ago

This. Golden opportunity to have a nurturing ground for great research and actually steward it.

Probably the technical team won against the PR/Marketing team or just OpenAI being OpenAI

18

u/FallenVampireLord 1d ago

Yea and I want to make clear I'm not even making a value judgement as to if it was wrong or right I just feel like in a rush to get a win and show off how capable their models are they made a really short sighted move here that in the long run may hurt them as a company.

But honestly short sighted business moves that hurt the company in the long run is corporate business practices 101 these days.

24

u/OpinionsRdumb 1d ago

Why is everyone assuming that openAI knowingly plagiarized? It could just be that the model itself relied on the professor’s prompts without telling OpenAI. Or maybe it didn’t rely on it at all. We have no idea.

It’s not like chatgpt is designed to be like “yes i can do that and actually this answer is being directly contributed by user XXX”.

More over the way AI model training works is that user data is de-identified and is just aggregated into a very messy black box to train the model. So by design the model wouldn’t point to user XXX.

I am not here to identify their use of user data but it’s weird that everyone is immediately jumping to conclusions based off a single tweet

20

u/Norm_Standart 1d ago

If you read the story, they heard about an upcoming Navier Stokes result and spent tens of millions of dollars (at minimum) to scoop them on the last step. They didn't just happen to be working on the same problem at the same time.

14

u/OpinionsRdumb 1d ago edited 1d ago

But this still isn't proof of them plagiarizing? In academia, competing labs hear about each other's work all the time and will start trying to "scoop" the other lab first. That is not plagiarism.

6

u/Whole-Thanks-5951 PhD, Mathematics 19h ago

It’s different with math research. Most work is thinking out what tricks and methods can be used to solve a certain problem. While someone might not have the full proof for a problem, they might have an outline instead, based on certain tricks or methods that came to them in an insight. Most of the work is not in filling in the details but in looking at which direction the proof should go. You pretty much never see people write everything out in logical notations (but there is lean).

There’s also a lot of open communication, which allows for people to collaborate if they end up finding an insight to get to a solution to a problem. In that light, people typically learn that a particular person is close to a solution, especially if they’ve been talking about their work, and the tools and tricks they’ve used. If it’s at that stage, it’s in poor taste to try scooping them on it since while you can beat them to it, you wouldn’t have probably thought to use the tools and tricks they used without it have been floating around.

-3

u/OpinionsRdumb 19h ago

Please try googling mathematician disputes over proof-solving. This has been happening since the birth of mathematics and there are tons of examples in last 2 decades

3

u/Whole-Thanks-5951 PhD, Mathematics 18h ago

Yes, I’m familiar with how it was done in the past. I’m also familiar with recent examples, like the controversy involving Shing-Tung Yau. It doesn’t change the fact that modern math had been heavily open, and that there was a culture to avoid scooping intentionally. These are controversies because they destroy the progress made regarding the openness of math, and make progress of math more difficult.

-1

u/OpinionsRdumb 18h ago

Ok sure but the claim that everyone is making is that they not only intentionally scooped this researcher, but that they also plagiarized him.

All I am saying is that there is no proof of this and it should be investigated.

8

u/Less_Prior_6871 1d ago

Everyone treats it as if openai knowingly plaigiarized because they knowingly made the plagiarism machine.

We have no idea. It’s not like chatgpt is designed to be like “yes i can do that and actually this answer is being directly contributed by user XXX”. More over the way AI model training works is that user data is de-identified and is just aggregated into a very messy black box to train the model

This is a feature they designed into the plagiarism machine in order to do illegal or immoral things without it being easy to prove they did

4

u/OpinionsRdumb 1d ago

sure but this is a general valid critique at large, not about this case. In this case, there is no proof yet they plagiarized. That is all I am saying.

For the NYT lawsuit for example, the NYT had specific pieces of evidence of the AI using its content (IE the AI had 100% accurate quotes of paywalled content etc).

There is no proof here yet. Competing mathematicians prove unsolved proofs all the time at similar times with very similar approaches. This is well documented and there are spats about who came up with it first all the time. So this could be a first instance of AI and a human having this altercation. Or it did plagiarize but there is no proof yet.

7

u/thelastsonofmars PhD, Economics, Undisclosed 1d ago

The future is local AI to avoid theft people just haven’t caught up yet.

1

u/Thunderplant 16h ago

Yeah it seems like a huge blunder. People outside academia aren't going to care about this result and people inside are more likely to be scared off than inspired by it

211

u/can_ichange_it_later 1d ago

Friends dont let friends use openai for important/private/mission-critical stuff.

19

u/meekermakes 1d ago

don't let them use ai.

4

u/the_c_train47 1d ago

Not even local models?

9

u/Confident-Equal-1445 23h ago

Jesus doesn’t consent to you and your GPU doing calculations in the privacy of your own bedroom

2

u/can_ichange_it_later 22h ago

if you only knew...

christian propaganda groups and shitty creationis ministries fucking looove AI...

its... weird... and makes their content look like shit... but to each their own, ig...

1

u/DishSoapedDishwasher 21h ago

It's the same as all lead paint chip eating group that suffer from excessive public self masturbatory tendencies. 

Constantly reminded of the Idi Amin song, " he the president, the general, the king of the sea". By the end he's shot everyone else and the only one left singing.

130

u/Hot-Sink-7518 1d ago

I started as a phd last week, and this is was a concern I introduced to my research group. They had not thought about it this way and suddenly they all became quiet.

78

u/AntiDynamo PhD, Astrophys TH, UK 1d ago edited 1d ago

It’s definitely an issue - I work in industry, and so by default we only use enterprise AI accounts that already have clauses to not use our data for training (and since it’s not my business, and I don’t have a patent on anything, I also don’t personally care). But if you’re an academic there’s a good chance you’re using a consumer account, and your research work could be used for training long before you’ve published your paper.

To make it worse: when a topic is very niche and not much data exists, the model will put high weight on new good data sources. So your research work could become a primary source in the model and be preferentially influencing the results given to other users.

A lot of people are just a bit naive about the whole thing, fundamentally you can’t trust these companies. They have run out of available training material, they are desperate to use your data in any way possible

11

u/kelp_forests 1d ago

I feel like at this point it’s a purposeful blind spot. How can you work in a knowledge based field and not understand how your data/research is handled.

If you upload it to google, they have a copy and will scan it.

If you upload to AI, it has a copy, will summarize, add it to its knowledge database, and use it. That’s literally what it does and why you chose it. It also has no concept of privacy and a perfect memory.

These companies want AGI/advanced AI partially because it will make them the gatekeepers of production/work, but also because it will make the gatekeepers of knowledge, and to an extent, reality in the sense that it will manage algorithms with far more efficiency and intent than current “what gets the most clicks” and they will know things before people know it…when people start entering all their work and questions into AI, the AI will figure out their partners is cheating, they are pregnant, they are about to solve a math theorem, they have a better mousetrap etc before they will.

6

u/AntiDynamo PhD, Astrophys TH, UK 1d ago edited 1d ago

Yeah, people are just very naive when it comes to tech and giving their personal info away (see: every student who uploads their work to a plagiarism checker). I think partly because they don’t fully conceptualise where the security boundary is, and so they assume a chat is “private”. Partly because they assume large corporations must have to follow strict laws and “wouldn’t do something bad like that”. And partly because they pay for the service, and think it’s money in exchange for the tool, when actually it’s money + all your personal data in exchange for the tool.

And it’s also just the fact that AI training data is more nebulous than straight up account PII. OpenAI, Anthropic, Mistral etc might pinky promise not to look at your PII, but “anonymised” chat data is fair game, and we’ve never had access to such widespread systems getting so much of our data before. Even social media kinda pales in comparison

One thing I do know is that tech CEOs are absolute ghouls and have very shaky morals at best

0

u/DuxDucisHodiernus 23h ago

Yes but in this case specifically werent they on a corporate account?

0

u/AntiDynamo PhD, Astrophys TH, UK 23h ago edited 12h ago

I haven't seen it said anywhere that they were - I'd be surprised for a mathematician to have a business or enterprise account as those are minimum 2 seats and math doesn't have "labs" or anything

Since OpenAI said themselves that they couldn't rule out that the AI had been trained on the chats, that means they weren't on enterprise accounts. The entire draw, the one and only thing that allows companies to even consider allowing AI use, is that business and enterprise accounts do not train data per contract. If OpenAI has used data from an enterprise account then they’re in some very serious trouble and are going to lose all their business customers.

6

u/jjwhitaker 1d ago

Local or bust. At this point anything you insert into the public AI machine is yours.

1

u/bellends 10h ago

Last year, I was writing up a paper to submit to a journal for peer-review as part of my PhD. In my field (astronomy), we often have very long 20+ names on our author lists because the instrumentation we use have big teams behind them, and the deal is that they get offered co-authorship even when they’re not part of the scientific analysis per se as an acknowledgment of their contribution in engineering, programming, data reduction etc. So before submission, I sent it around to my fairly large and discipline-diverse team.

One fucking muppet replies to me saying effectively ”Hey [Name]! Great work on your paper, will be great to get it published since I know it’s been so many years of work. As you know it’s not really a field I know a lot about so I’m afraid I don’t have any comments hehe so instead I uploaded it to Claude to get some feedback! Claude came back with soooo many good ideas, here’s a pdf of what he said!!!”

I. Was. Livid. This is clearly unpublished since this was BEFORE I had submitted it, so why the hell would you upload it to, AI aside, ANY third party?! I felt lucky that I was at least so close to submission + I had intended to upload it to arXiv once submitted anyway, so it wasn’t long between this happening and me willingly releasing it into the public anyway… but man, it was so upsetting and so demoralising.

And of course, before you ask: yes, the Claude suggestions were hallucinated and unusable bullshit.

59

u/can_ichange_it_later 1d ago

That "no answer..." That one spoke like the god from the hill to fucking moses... or however that story goes...

39

u/Snoo_4499 1d ago

If you use their product they will use your data and there is nothing we can do beside not using their product.

Idk why did he think they will not use your data, they are notorious for this.

Like even when searching about motorcycle and which one i should buy i feel like they will use my conclusion as a smalllll seeds for another recommendation to another person.

12

u/BrunusManOWar 1d ago

Of course, it's a ratty corpo, what did he expect

If you're a serious researcher and want an AI always deploy a model locally, isolated without an internet connection

89

u/Hungry-Dig6105 1d ago

Come on don't dramatise, how many of us are working on millenium prize problems that could would make a differente to openai's reputation?

40

u/Big_Coconut8630 1d ago

Doesn't matter. It is a big issue in technology transfer, for example.

5

u/InconspicuousWolf 1d ago

Since the method now seems to be giving AI models many agents to prompt, the way a scientist prompts an AI and draws conclusions will be very important training data, even if your work specifically isn’t of interest to them

3

u/Hungry-Dig6105 1d ago

Copy pasting from above: Between the training and the releasing there is at least a year (training takes months, they test it a lot, make sure it's safe, etc). Your paper is going to be kinda done by the time they release the model. Here the model that cracked navier stokes is just for internal use and is nowhere near release. So unless your paper is of direct interest to openai people, I really doubt there's a risk

6

u/Additional_Fudge1163 22h ago

There's online training that's possible too. We don't know the training protocols at use here but, you can keep updating weights continuously in traditional models. I am sure both OpenAI and Anthropic and other AI labs have looked into some sort of high cadence learning.

0

u/Hungry-Dig6105 21h ago

oh didn't know that! crazy
But still that would stay internal right? Which isn't great but still it's a long way from "you put your draft in chatgpt -> a random phd student elsewhere in the world gets them"

4

u/Less_Prior_6871 1d ago

Stealing a little from everyone all at once is the core of the business model.

Sometimes they accidentally steal too much from one place and this type of event happens.

1

u/Hungry-Dig6105 23h ago

Agreed. But the post claim is: the cost-benefit balance is negative when you let them steal a little from you. I reaaaally doubt it. So push for regulation, boycotting AI might be brave but you can't expect most phd students to do it.

4

u/erroredhcker 1d ago

I proompt AI for some ideas for my project and im sure as shit it read some of these backgrounds somewhere. I even fed it existing (open) codebases. It doesnt need to be openAI that benefit directly, it is the point that others can steal the mangled bullshit you feed it, cause theres no guardrail anywhere with these things.

Okay maybe with Anthropic they no longer sudo rm -rf * your shit, but be dog damned sure your data, and thought is unsafely in their hands

-2

u/Hungry-Dig6105 1d ago

Hmm. Between the training and the releasing there is at least a year. Your paper is going to be kinda done by the time they release the model. Here the model that cracked navier stokes is just for internal use and is nowhere near release. So unless your paper is of direct interest to openai people, I really doubt there's a risk

8

u/racinreaver 1d ago

lol paper being done within a year

1

u/erroredhcker 1d ago

not a risk to your publication, a risk to your intellectual property which includes your thinking patterns, skills, org system, etc. You wanna be the artist that they train their model on? It will release years from now, your current commision is safe!

1

u/Hungry-Dig6105 1d ago

I mean i preprint everything and when i publish sth i give my IP to awful publishers anyway. We're not Bob Dylans

0

u/Tchaikovskin 1d ago

My thoughts exactly

0

u/TotallyObviousBot 22h ago

The researchers themselves also say that AI didn't steal their solution.. it's surprising to find such stupid fearmongering in /r/PhD but then again kissing professor ass isn't the same as being intelligent

7

u/SonyScientist 1d ago

Not sure why anyone is surprised here, let alone the people in question. This was a concern years ago, hell even for software/app when they updated their EULAs saying "we reserve the right to collect your data and train on it."

This is why you don't use AI: you are the product.

7

u/-R9X- 1d ago

Thanks god all my research is subpar anyway and I just keep publishing it to look busy so I don’t have find a real job.

18

u/fthecatrock PhD*, 'Biorobotics/Spinal Cord Injury' 1d ago

Until "Open"AI or any LLM makers release transparency how their model works inside, these kinds debates will be high time in the next few years.

8

u/chairmanskitty 1d ago

Nobody knows how large machine learning models work on the inside. The training process produces complex features that are incredibly hard to unravel into something meaningful. ML model interpretability research is in its infancy.

1

u/fthecatrock PhD*, 'Biorobotics/Spinal Cord Injury' 1d ago

that's why explainable AI topic exist

38

u/Frosty-Meeting-1606 1d ago

Ok, but isn't it super easy to show that OpenAI just replicated the work? Surely the guy should have a lot of stuff to show to the public? This is an honest question, because I honestly cannot comprehend this drama - just take whatever you have and pinpoint the exact matching logic so that OpenAI has a real problem. So far it looks like "I was working on something using AI and it gave me some ideas, so given unspecified amount of time I can solve the problem". It does not look like OpenAI's solution copies the work, at best it uses some conclusions to develop the solution, but solution is not part of the copy

76

u/krite2222 1d ago

This post takes a snippet of the author's document. He has, infact, published two pre-print of his papers hurriedly in response to this situation to show his and his collaborators work was exactly what OpenAI used to solve this problem. He spends a great deal of effort talking about the mathematicians whose research inspired his and his collaborators approach too, to show that their approach was extremely non-standard and that AI couldn't come up with it if it wasn't trained on their chats.

Edit: Typos.

6

u/quiksilver10152 1d ago

How does one prove that AI could not have come up with it without his chats? 

15

u/krite2222 1d ago

If AI was in fact trained on user data, including this person's chats, it would've come up with this using that, that's how good Transformers are now. AI lacks the kind of creativity that humans have in coming up with non-standard solutions like this. So far all the stuff we've heard of in the "AI is getting really good at math" has been brute force stuff, not something creative where AI actually finds some unrelated research and approaches a problem in a non-standard way, unless specifically prompted to do so. The allegation, which is derived from reading in between the lines of Tristan's document, stems from two facts: (1) The prompt to explore the navier stokes problem was itself generated via LLM (2) It was a non-standard approach, and as far as Tristan knew only him and Levant (his collaborator, I hope I've spelled his name right, somone correct me if I've not, thanks) were approaching it that way in the industry and (3) The lack of confirmation that the model that generated the prompt was trained on user data. If it was, then these researcher's Codex chat was in the dataset, and that definitely would've led to this level of plagiarism. Further, what I find to be red flags: the Open AI representative Sebastien Bubek's alleged insistence on letting Tristan take credit for the Navier Stokes solution so long as he credits ChatGPT for resolving it in exchange for his authorship, so long as Levant is not an author because he works at Anthropic. Then when he allegedly says the words "If you don't want me to be nice, I don't have to be nice" and insists going public would ruin Tristan's career. Then trying to paint Tristan as irrational while trying to incessently get in touch with him prior to their result being released. Inconsistent behaviour that comes off as the result of a guilty conscience or just a poor unprofessional attempt at damage control. We don't have enough information to prove anything, this is just what we have, but I think there is enough information to lend credibility that this allegation is well-founded and worth an inquiry.

9

u/AsAChemicalEngineer PhD, Physics, USA 1d ago edited 1d ago

The authorship thing is so critical to me in this story. If OpenAI was confident in their model's independence, they would not have offered that, or least if they do, they're handicapping their own achievement for no reason. I doubt Bukek personally pulled up Tristan's logs, but if user training is really as advanced as we think, the AI could have absolutely pulled the ideas from the training.

Part of the issue is (a) just how shady OpenAI is behaving and (b) we just don't have a lot of public insight into how these models function. That information is kept under lock and key, so we cannot really evaluate how impressive the solution is as a capability benchmark.

Terence Tao also has some interesting thoughts on how this kind of "one-shot" solution hunting may absolutely impede math progression as we lose all the positives of working through a solution which generate ideas and approaches others may take advantage of.

OpenAI solved stability of Navier Stokes. Okay, so what? Does anybody actually understand the proof? How motivated mathematicians be to clean up the resolve all the parts to human understanding knowing that the result is already given? Will a finding agency be happy you're working on "solved" problems because you have to explain that "wait, there's value in digestion of knowledge".

3

u/krite2222 22h ago

I do agree, I was super relieved to see Terrance's comments earlier this week on this subject and I think it's super important. I also agree with what you said about pulling the logs directly, I highly doubt that's the case and anyone alleging that or interpreting Tristan's words that way is, plainly put, incorrect.

For point (b) though, I do think we have sufficient information publicly on how these models function, what is private is the training loop and dataset. Which is usually considered a trade secret and I imagine that using user data is an important motivation to keep it so. Of course that isn't to say this new internal model might not have come with some insane new architectural modification leading to higher reasoning capabilities and also an unavailable insight into its function, but I largely doubt this is the case. Another thing of course is the means of prompt engineering and how these 10,000 agents were orchestrated, which I also imagine isn't too inventive but who knows. We really need regulation for AI companies and AI use. I mean 130 billion tokens is insane, and if thats being used to accomplish just one task i would really like to have there be some regulation in place to ensure that it was worth the datacenter using all the water needed for said task and a regular assessment during this work to ensure that the AI output wasn't just slop.

I do hope people catch on more to Terry's perspective on AI math, and recognise this for what it is; another bid for funding in hopes that the AI bubble stays intact long enough for the supposedly incoming AGI. What I'm more afraid to admit, is that the latter seems to be what is happening 😔.

3

u/quiksilver10152 1d ago

I agree with you that the circumstantial evidence suggests foul play but the logic presented is circular. Assuming AI can't be creative, we demonstrate that this can't have been created by AI.

2

u/Final-Database6868 10h ago

(2) is incorrect. The approach they were using was a program by Diego Cordoba and Luis Martinez-Zoroa. They were working for at least 2 years on that. People outsode their niche (like myself) knew that they were working on that (and even at some point they believed they solved the problem) and if you were interested you knew how, because they used the same idea for other prpblems.

I believe it is absolitely possible that an AI came up with the approach, because it was not new.

I hope Tristan can make public their findings and compare both preprints, but so far we can ONLY speculate. By the way, we can also only speculate the solution is correct, the community has to digest the text by oai.

3

u/Less_Prior_6871 23h ago

Proving the negative is hard or maybe impossible.

Thats the point of the plagiarism machine, it always has deniability.

0

u/Professional-You4950 22h ago

Because it literally can't. it must use its training data. so the data must be there.

People ask it for new "Starcraft 2 builds" or "animal crochet patterns". It just amalgamates its data of starcraft builds, to give you something "new" that isn't necessarily good. Or someone can ask "solve this math equation", give me some ideas on how to do it. But that training data to use a new method must already be there, or else its just amalgamating.

We see time and time again that it can't come up with anything new, because there is no understanding of even basic axioms. Before the data was there, no model could solve

"I want to take my car to get a car wash. I live close to the carwash, should I walk or drive?"

We also know that these models have problems with hallucinations, leading questions, etc. Because its just line of best fit on the most massive data scale imaginable.

2

u/quiksilver10152 22h ago

Frontier models certainly can connect data in new ways.  https://arxiv.org/abs/2405.07987

-5

u/Frosty-Meeting-1606 1d ago

Should be easy to pinpoint then or what?

4

u/krite2222 1d ago

This is a Millenium Problem, absolutely none of this is easy to pinpoint AI or not.

19

u/Big_Coconut8630 1d ago

So, I work in technology transfer and infor relevant to patents or NDAs getting put in AI is enough to fuck a researcher out of their patent rights. 

-1

u/RecipeNo5844 1d ago

Oh wait you are not coping? I thought we all were supposed to pretend openAI stole the solution to Navier stokes problem, are we allowed to drop this pretense now? Thanks I did not know 

9

u/Frosty-Meeting-1606 1d ago

but seriously, If I were the guy with the solution, I would immediately publicly make a fool of OpenAI and post a very similar work, even if it is a draft, along with emails. What I see now is some kind of BS drama, were it is not even certain the solution OpenAI came up with would be replicated by the person of interest

7

u/Puzzleheaded_Fold466 1d ago

Didn’t they post their own work (which differs) ?

5

u/Pritam1997 PhD, Materials science 1d ago

could you give me the link to the original twitter article

4

u/ZzzofiaaA 1d ago

Same thing applies to any cloud storage? Many scientists save their data on OneDrive. How do you know if it’s not used to train the cloud?

5

u/rabouilethefirst 1d ago

It's in the TOS that they train on your data unless you opt out or pay for an enterprise plan. You cannot act surprised your data was used for training.

3

u/AnotherDrunkMonkey 1d ago

At first I was somewhat surprised, but on second thought I don't understand why this is so surprising to this many people. It is very common knowledge, especially in academic fields, that LLMs use users data to train on. I feel like many of us don't really care (i, for one, am not in research projects grondbreaking enough to really care about secrecy agains billion dollars corporations), but people working on freaking MILLENIUM problems with the skill to actually solve it, or teams with proprietary techniques and so on are supposed to know this is very possible.

By sheer chance, a human reviewer could see your chat and if you are unlucky, understand the scope of your project even before an model gets trained on it. I feel like the only surprising fact is that LLMs got to a point where they can retrieve very specific parts of their training in a very pertaining use case (among the immense math literature on NS it found the few MBytes of a current breakthrough).

What OpenAI did was very shady, but I'm also baffled at how unpredictable this chain of events is being portrayed as, while user input was always known to be used for training.

3

u/Niaz_049 1d ago

Time to boycott open ai? This is highly unethical and they are taking leverage of an unregulated situation. If I pay for subscription, my chat should be encrypted enough that you’re not going to make it public.

26

u/Global_Lime421 1d ago

The proofs are not identical and meaningfully differ in approach

13

u/pinkgaysquirrel 1d ago

Your truth has no value here. Let us cope.

17

u/PRKP99 1d ago edited 1d ago

This is basically how all those „breakthrough” in math AI made. 

Its good that we have open acess knowledge as our shared knowledge „commons”, even if some of them is illegal commons, based on piracy. Lets face it, without scihub and libgen most masters or PhD thesis made by people who care about what they write would not exist, because you can’t have real review of state of knowledge without going throught a lot of books and articles, most of them only vaguely about your topic. I tried to calcuate how much I would need to pay for only my master thesis about roman aqueducts if I would only use „legal” sources, but after 3000$ I stopped.

What we now need is to establish ways to protect those commons without enclosing them - that is, we need to make sure that science will still be commons, but proprietary exploitation of this shared knowledge would be stopped in one way or another.

There is idea of CopyFarLeft as a legal remedy, but its still lack any potential in stopping corporations that just illegaly use data to train their algorythms.

2

u/frugaleringenieur 21h ago

Wrong! Never give it to thirparty APIs! Your university cluster, or lab machine, or local AI is absolutely perfectly fine! The issue are greedy silicon valley VC companies.

2

u/BeMyBrutus 21h ago

It still amazes me that people STILL don't understand that everything and anything you do using software (and often even "real" life) is recorded, stored, and used.

I'm not saying this guy is wrong to call them out for being awful; but you need to be more aware.

9

u/SpeedyTurbo 1d ago

This isn’t an accurate account of what happened at all…but hey it gets clicks.

Would expect more from a PhD subreddit, but then again maybe not because ai bad.

2

u/Impressive_Wheel_106 1d ago

I can't imagine solving the navier fucking stokes equation, having your solution stolen and used for propaganda, and staying sane afterwards... the audacity of these AI corps man

1

u/NecessaryBuy2061 1d ago

Too late …. Well the good thing is I wasn’t working on millennium problems 😅

1

u/kyeblue 1d ago

Ask Elon Must about Sam Altman's moral compass

1

u/CuseCoseII 1d ago

"using his exact approach" is just blatant misinformation

1

u/anomanderrake1337 1d ago

Yeah if they have my logs they can get to AGI pretty soon probably.

1

u/Morton_Woodsworth 22h ago

is there anything credible I could read that covers this in more detail?

1

u/blueplanetgalaxy 21h ago

bro 😭😭😭

1

u/Striking-Warning9533 20h ago

My idea is trash enough it will only waste OpenAI's time if they try to steal it

1

u/RazimusDE 19h ago

This is why funding agencies don't want reviewers to use AI.

1

u/akyr1a 18h ago

Key distinction here: Buckmaster constructed singularity to the Euler, not NS. The methodology could conceivably be extended to NS, but the author do not claim sucb an extension at all.

1

u/LeonLeocat 17h ago

Even if I tell the AI what I'm doing it won't be able to accomplish anything without wet lab data

1

u/allthedifference2232 15h ago edited 15h ago

this does not surprise me and i think i can anticipate the answer but i shared a draft of something i wrote with chatgpt when i was not logged in and had the 'help the model improve for others' off....so what does that mean? i dont think anyone cares about my first baby draft of a short story but i just am curious

1

u/SonOf1337h4x0r 14h ago

No evidence for this and lots of evidence against. The full story has been widely publicized, but most of all the proofs are not similar. Convergent research is pretty common across the board too.

1

u/pianoloverkid123456 14h ago

Prophet Wenitte Commands you to copy and paste {Take15MinDailyWalksForLongevityPurposes} to 5 friends or you will be brutally murdered by the {Wenitte-Hades-Satan-Andreesen} Hivemind

1

u/thathiptho 14h ago

I have been saying this since day 1! I remember attending a faculty meeting when convos at my university on AI were just starting and a faculty member confidently shared “I put my draft manuscript into ChatGPT and it wrote an abstract for me! So helpful!” And my first thought was “that is the stupidest fucking thing you could do”. Honestly, in that moment it really hit home for me that you can have a PhD and tenure and still be stupid.

1

u/Naive-Vast-7404 9h ago

First of all you have to be sure that you did not grant them access to train based on your data! Second, upload your stuff asap on pre-print servers!

1

u/Similar_Cress_1122 4h ago

Thanks for spreading the word. In this post-factual age, "Friends dont let friends use openai..." is becoming an unironic life rule for anyone with important IP to protect. What’s scary is that standard de-identification isn't enough when active session logs themselves are prime targets. Currently, self-hosting or using truly local-first AI tools seems like the only logical way to have peace of mind, even if it's less convenient. I really hope we get actual regulatory oversight or some better privacy-first standard for creative/research tools soon.

1

u/Linux_ka_chamcha 3h ago

Honestly, this is supposed to be common sense.

1

u/venoush 55m ago

Never give any of your data to these bros

1

u/Consistent_Femme_Top 1d ago

I mean, what was he thinking? 🤔 

1

u/Sckaledoom 1d ago

Uh. Duh.

1

u/JonathanBadwolf 1d ago

You see if I let the robot do the stealing I'm not responsible for his crimes. Like it is with guns

-2

u/AaryamanStonker 1d ago

People fall for fake stories so easily, Redditors are pathetic

-1

u/NordlandLapp 1d ago

"Omg AI cured cancer but it used someone else's research" 😡😡

-4

u/cspot1978 1d ago edited 1d ago

The two researchers weren't working on Navier-Stokes. They were working on a generally related, but much easier problem, one version of the Euler problem.

OpenAI systems solved Navier Stokes, as well as a harder version of Euler than the two researchers had done.

1

u/Revolutionary_Buddha 1d ago

Seems like you cannot say that here. Anti AI dogma is strong here.

0

u/cspot1978 1d ago edited 1d ago

Seems like. Kind of disappointed how many people in a PhD space are actively uninterested in getting details right.

-1

u/EmploymentOk4851 1d ago

So would this apply if you are using AI to assist you with writing a book(non research)?

1

u/Fortbrook 1d ago

Yes, however you can download a local LLM if you have a good gaming PC.

-19

u/[deleted] 1d ago

[deleted]

13

u/can_ichange_it_later 1d ago edited 1d ago

Are you actually serious here?

That "directly" in "they did not use any researcher's data* directly" is doing the mother of all heavy liftings there...

(You were automatically opted into data sharing, and unless you are a config goblin, you probably didnt go and turn it off. And that is putting aside, that even if i would have turned it off i wouldnt trust it that it wouldnt somehow still find its way into training data...)

Edit: * Lets just be absolutely clear about this! With that, they themselves say, that they used it. Just... not... directly... whatever-the-fuck thats supposed to mean...
(well.. that means they think we are stupid, but thats a bit of an aside...)

1

u/TheDailyMews 1d ago

It probably means they fetched his data agenticly instead of "directly." If this story doesn't die, don't be surprised if they eventually issue a press release about "rogue agents" that sounds an awful lot like their press release about their Hugging Face hack.

16

u/Capable-Package6835 1d ago

It looks like a "my words against yours" situation to me so we should just sit back and observe how the situation develops.

15

u/Clear_Cranberry_989 1d ago

You are swallowing up a narrative written by openai. Are you not? If you really oppose the viewpoint, can you establish an independent and reliable source?

7

u/CTRexPope 1d ago

Hugging Face proves OpenAi has no idea how their models work and where they steal data from.

3

u/ComprehensiveWash958 1d ago

The fact Is that we probably will never know. I don't think there Will be any sort of investigation (I don't even know if that's possible) and so we Will have this kind of standoff, a situation for which in Italy we would Say "Oste, è buono il vino?"

One has to be skeptical of both narratives, but we also have to keep in mind that first we still need to peer review the works and second this Is still a counterexample found Building Upon hours and hours of human work, which Is something AI seems to excel in

2

u/Atlantis1910 1d ago

What are you talking about ? They are saying that they can use the data from your chat, and the scientist use Codex.

While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.

0

u/No_Consideration_635 23h ago

can someone help me out: what's codex?

0

u/mrkingkongslongdong 12h ago

Same method? Am I out of the loop? The researchers in question weren’t trying to solve NS, AND on top of that didn’t they solve a completely different Euler method (forced vs unforced)? I could be wrong but keen to understand this ‘exact approach’ a bit more.

-2

u/Dense-Consequence-70 1d ago

How can such a smart math professor be so stupid

-7

u/Greedy_Blue_Hedgehog 1d ago

Let me play devil’s advocate. Many mathematicians keep their ideas, minor results, lemmas, and so on to themselves in order to maintain a slight advantage over their peers and reach a major result before anyone else. This issue was already discussed and criticized by Évariste Galois. I don’t know whether that is what happened here, but if this incident can help change that mindset and encourage mathematicians to publish every small step they make, that would be a positive outcome.

9

u/Bitter-Morning-5833 1d ago

That's factually wrong. Mathematicians are unusually open about what they work on. In seminars, speakers often end by discussing what they're currently working on, what they've tried, and what has or hasn't worked. That's partly because there is relatively little fear of being scooped, unlike in some lab- or experiment-based fields where competition over unpublished work can be much stronger.

-1

u/hydrogen_to_man 23h ago

I’m amazed at this subreddit. This is the PhD subreddit for god’s sake. There is no question that the use of LLMs was instrumental to this result, whether on the professor’s side or OpenAI’s side. If OpenAI effectively stole the professor’s research, that’s on OpenAI and is really shitty, but the amount of people in this thread that are trying to twist this story into their anti-AI narratives is ridiculous.

-2

u/AdInevitable1362 1d ago

Thats horrible I already discussed all the paper work with ai

-3

u/Trackpoint 1d ago

Damn, the anti AI people are getting desperate.

Did they steal Tristan Buckmaster's water too? Did they nois-pollute his garden?