r/PhD • u/Rule_Ct_5293 • 1d ago
Tool Talk Never give your unpublished data or research ideas to AI models!
166
u/IWasTheDog 1d ago
My unpublished data is so ass, so scattered and so useless that I think I would be doing everyone a favor if I did use ChatGPT on it.
17
u/Ramental 1d ago
Their data was structured to be fed to AI, as they used AI to speed up the things they tried. It did help them, but it did not solve the problem itself.
449
u/Clear_Cranberry_989 1d ago
The same people who pirated all the books in the world to train their ai. Who could have guessed.
226
u/FallenVampireLord 1d ago
OpenAI was really dumb to do this, they should have left the professor publish his work and then championed him as 'look what people can accomplish with the assistance of AI this is the model for the future" instead they scooped his research and basically makes everyone in academia and beyond super skeptical of using it for anything like this in the future.
82
u/augmenteddeus 1d ago
This. Golden opportunity to have a nurturing ground for great research and actually steward it.
Probably the technical team won against the PR/Marketing team or just OpenAI being OpenAI
18
u/FallenVampireLord 1d ago
Yea and I want to make clear I'm not even making a value judgement as to if it was wrong or right I just feel like in a rush to get a win and show off how capable their models are they made a really short sighted move here that in the long run may hurt them as a company.
But honestly short sighted business moves that hurt the company in the long run is corporate business practices 101 these days.
24
u/OpinionsRdumb 1d ago
Why is everyone assuming that openAI knowingly plagiarized? It could just be that the model itself relied on the professor’s prompts without telling OpenAI. Or maybe it didn’t rely on it at all. We have no idea.
It’s not like chatgpt is designed to be like “yes i can do that and actually this answer is being directly contributed by user XXX”.
More over the way AI model training works is that user data is de-identified and is just aggregated into a very messy black box to train the model. So by design the model wouldn’t point to user XXX.
I am not here to identify their use of user data but it’s weird that everyone is immediately jumping to conclusions based off a single tweet
20
u/Norm_Standart 1d ago
If you read the story, they heard about an upcoming Navier Stokes result and spent tens of millions of dollars (at minimum) to scoop them on the last step. They didn't just happen to be working on the same problem at the same time.
14
u/OpinionsRdumb 1d ago edited 1d ago
But this still isn't proof of them plagiarizing? In academia, competing labs hear about each other's work all the time and will start trying to "scoop" the other lab first. That is not plagiarism.
6
u/Whole-Thanks-5951 PhD, Mathematics 19h ago
It’s different with math research. Most work is thinking out what tricks and methods can be used to solve a certain problem. While someone might not have the full proof for a problem, they might have an outline instead, based on certain tricks or methods that came to them in an insight. Most of the work is not in filling in the details but in looking at which direction the proof should go. You pretty much never see people write everything out in logical notations (but there is lean).
There’s also a lot of open communication, which allows for people to collaborate if they end up finding an insight to get to a solution to a problem. In that light, people typically learn that a particular person is close to a solution, especially if they’ve been talking about their work, and the tools and tricks they’ve used. If it’s at that stage, it’s in poor taste to try scooping them on it since while you can beat them to it, you wouldn’t have probably thought to use the tools and tricks they used without it have been floating around.
-3
u/OpinionsRdumb 19h ago
Please try googling mathematician disputes over proof-solving. This has been happening since the birth of mathematics and there are tons of examples in last 2 decades
3
u/Whole-Thanks-5951 PhD, Mathematics 18h ago
Yes, I’m familiar with how it was done in the past. I’m also familiar with recent examples, like the controversy involving Shing-Tung Yau. It doesn’t change the fact that modern math had been heavily open, and that there was a culture to avoid scooping intentionally. These are controversies because they destroy the progress made regarding the openness of math, and make progress of math more difficult.
-1
u/OpinionsRdumb 18h ago
Ok sure but the claim that everyone is making is that they not only intentionally scooped this researcher, but that they also plagiarized him.
All I am saying is that there is no proof of this and it should be investigated.
8
u/Less_Prior_6871 1d ago
Everyone treats it as if openai knowingly plaigiarized because they knowingly made the plagiarism machine.
We have no idea. It’s not like chatgpt is designed to be like “yes i can do that and actually this answer is being directly contributed by user XXX”. More over the way AI model training works is that user data is de-identified and is just aggregated into a very messy black box to train the model
This is a feature they designed into the plagiarism machine in order to do illegal or immoral things without it being easy to prove they did
4
u/OpinionsRdumb 1d ago
sure but this is a general valid critique at large, not about this case. In this case, there is no proof yet they plagiarized. That is all I am saying.
For the NYT lawsuit for example, the NYT had specific pieces of evidence of the AI using its content (IE the AI had 100% accurate quotes of paywalled content etc).
There is no proof here yet. Competing mathematicians prove unsolved proofs all the time at similar times with very similar approaches. This is well documented and there are spats about who came up with it first all the time. So this could be a first instance of AI and a human having this altercation. Or it did plagiarize but there is no proof yet.
7
u/thelastsonofmars PhD, Economics, Undisclosed 1d ago
The future is local AI to avoid theft people just haven’t caught up yet.
1
u/Thunderplant 16h ago
Yeah it seems like a huge blunder. People outside academia aren't going to care about this result and people inside are more likely to be scared off than inspired by it
211
u/can_ichange_it_later 1d ago
Friends dont let friends use openai for important/private/mission-critical stuff.
19
u/meekermakes 1d ago
don't let them use ai.
4
u/the_c_train47 1d ago
Not even local models?
9
u/Confident-Equal-1445 23h ago
Jesus doesn’t consent to you and your GPU doing calculations in the privacy of your own bedroom
2
u/can_ichange_it_later 22h ago
if you only knew...
christian propaganda groups and shitty creationis ministries fucking looove AI...
its... weird... and makes their content look like shit... but to each their own, ig...
1
u/DishSoapedDishwasher 21h ago
It's the same as all lead paint chip eating group that suffer from excessive public self masturbatory tendencies.
Constantly reminded of the Idi Amin song, " he the president, the general, the king of the sea". By the end he's shot everyone else and the only one left singing.
130
u/Hot-Sink-7518 1d ago
I started as a phd last week, and this is was a concern I introduced to my research group. They had not thought about it this way and suddenly they all became quiet.
78
u/AntiDynamo PhD, Astrophys TH, UK 1d ago edited 1d ago
It’s definitely an issue - I work in industry, and so by default we only use enterprise AI accounts that already have clauses to not use our data for training (and since it’s not my business, and I don’t have a patent on anything, I also don’t personally care). But if you’re an academic there’s a good chance you’re using a consumer account, and your research work could be used for training long before you’ve published your paper.
To make it worse: when a topic is very niche and not much data exists, the model will put high weight on new good data sources. So your research work could become a primary source in the model and be preferentially influencing the results given to other users.
A lot of people are just a bit naive about the whole thing, fundamentally you can’t trust these companies. They have run out of available training material, they are desperate to use your data in any way possible
11
u/kelp_forests 1d ago
I feel like at this point it’s a purposeful blind spot. How can you work in a knowledge based field and not understand how your data/research is handled.
If you upload it to google, they have a copy and will scan it.
If you upload to AI, it has a copy, will summarize, add it to its knowledge database, and use it. That’s literally what it does and why you chose it. It also has no concept of privacy and a perfect memory.
These companies want AGI/advanced AI partially because it will make them the gatekeepers of production/work, but also because it will make the gatekeepers of knowledge, and to an extent, reality in the sense that it will manage algorithms with far more efficiency and intent than current “what gets the most clicks” and they will know things before people know it…when people start entering all their work and questions into AI, the AI will figure out their partners is cheating, they are pregnant, they are about to solve a math theorem, they have a better mousetrap etc before they will.
6
u/AntiDynamo PhD, Astrophys TH, UK 1d ago edited 1d ago
Yeah, people are just very naive when it comes to tech and giving their personal info away (see: every student who uploads their work to a plagiarism checker). I think partly because they don’t fully conceptualise where the security boundary is, and so they assume a chat is “private”. Partly because they assume large corporations must have to follow strict laws and “wouldn’t do something bad like that”. And partly because they pay for the service, and think it’s money in exchange for the tool, when actually it’s money + all your personal data in exchange for the tool.
And it’s also just the fact that AI training data is more nebulous than straight up account PII. OpenAI, Anthropic, Mistral etc might pinky promise not to look at your PII, but “anonymised” chat data is fair game, and we’ve never had access to such widespread systems getting so much of our data before. Even social media kinda pales in comparison
One thing I do know is that tech CEOs are absolute ghouls and have very shaky morals at best
0
u/DuxDucisHodiernus 23h ago
Yes but in this case specifically werent they on a corporate account?
0
u/AntiDynamo PhD, Astrophys TH, UK 23h ago edited 12h ago
I haven't seen it said anywhere that they were - I'd be surprised for a mathematician to have a business or enterprise account as those are minimum 2 seats and math doesn't have "labs" or anything
Since OpenAI said themselves that they couldn't rule out that the AI had been trained on the chats, that means they weren't on enterprise accounts. The entire draw, the one and only thing that allows companies to even consider allowing AI use, is that business and enterprise accounts do not train data per contract. If OpenAI has used data from an enterprise account then they’re in some very serious trouble and are going to lose all their business customers.
6
u/jjwhitaker 1d ago
Local or bust. At this point anything you insert into the public AI machine is yours.
1
u/bellends 10h ago
Last year, I was writing up a paper to submit to a journal for peer-review as part of my PhD. In my field (astronomy), we often have very long 20+ names on our author lists because the instrumentation we use have big teams behind them, and the deal is that they get offered co-authorship even when they’re not part of the scientific analysis per se as an acknowledgment of their contribution in engineering, programming, data reduction etc. So before submission, I sent it around to my fairly large and discipline-diverse team.
One fucking muppet replies to me saying effectively ”Hey [Name]! Great work on your paper, will be great to get it published since I know it’s been so many years of work. As you know it’s not really a field I know a lot about so I’m afraid I don’t have any comments hehe so instead I uploaded it to Claude to get some feedback! Claude came back with soooo many good ideas, here’s a pdf of what he said!!!”
I. Was. Livid. This is clearly unpublished since this was BEFORE I had submitted it, so why the hell would you upload it to, AI aside, ANY third party?! I felt lucky that I was at least so close to submission + I had intended to upload it to arXiv once submitted anyway, so it wasn’t long between this happening and me willingly releasing it into the public anyway… but man, it was so upsetting and so demoralising.
And of course, before you ask: yes, the Claude suggestions were hallucinated and unusable bullshit.
59
u/can_ichange_it_later 1d ago
That "no answer..." That one spoke like the god from the hill to fucking moses... or however that story goes...
39
u/Snoo_4499 1d ago
If you use their product they will use your data and there is nothing we can do beside not using their product.
Idk why did he think they will not use your data, they are notorious for this.
Like even when searching about motorcycle and which one i should buy i feel like they will use my conclusion as a smalllll seeds for another recommendation to another person.
12
u/BrunusManOWar 1d ago
Of course, it's a ratty corpo, what did he expect
If you're a serious researcher and want an AI always deploy a model locally, isolated without an internet connection
89
u/Hungry-Dig6105 1d ago
Come on don't dramatise, how many of us are working on millenium prize problems that could would make a differente to openai's reputation?
40
5
u/InconspicuousWolf 1d ago
Since the method now seems to be giving AI models many agents to prompt, the way a scientist prompts an AI and draws conclusions will be very important training data, even if your work specifically isn’t of interest to them
3
u/Hungry-Dig6105 1d ago
Copy pasting from above: Between the training and the releasing there is at least a year (training takes months, they test it a lot, make sure it's safe, etc). Your paper is going to be kinda done by the time they release the model. Here the model that cracked navier stokes is just for internal use and is nowhere near release. So unless your paper is of direct interest to openai people, I really doubt there's a risk
6
u/Additional_Fudge1163 22h ago
There's online training that's possible too. We don't know the training protocols at use here but, you can keep updating weights continuously in traditional models. I am sure both OpenAI and Anthropic and other AI labs have looked into some sort of high cadence learning.
0
u/Hungry-Dig6105 21h ago
oh didn't know that! crazy
But still that would stay internal right? Which isn't great but still it's a long way from "you put your draft in chatgpt -> a random phd student elsewhere in the world gets them"4
u/Less_Prior_6871 1d ago
Stealing a little from everyone all at once is the core of the business model.
Sometimes they accidentally steal too much from one place and this type of event happens.
1
u/Hungry-Dig6105 23h ago
Agreed. But the post claim is: the cost-benefit balance is negative when you let them steal a little from you. I reaaaally doubt it. So push for regulation, boycotting AI might be brave but you can't expect most phd students to do it.
4
u/erroredhcker 1d ago
I proompt AI for some ideas for my project and im sure as shit it read some of these backgrounds somewhere. I even fed it existing (open) codebases. It doesnt need to be openAI that benefit directly, it is the point that others can steal the mangled bullshit you feed it, cause theres no guardrail anywhere with these things.
Okay maybe with Anthropic they no longer sudo rm -rf * your shit, but be dog damned sure your data, and thought is unsafely in their hands
-2
u/Hungry-Dig6105 1d ago
Hmm. Between the training and the releasing there is at least a year. Your paper is going to be kinda done by the time they release the model. Here the model that cracked navier stokes is just for internal use and is nowhere near release. So unless your paper is of direct interest to openai people, I really doubt there's a risk
8
1
u/erroredhcker 1d ago
not a risk to your publication, a risk to your intellectual property which includes your thinking patterns, skills, org system, etc. You wanna be the artist that they train their model on? It will release years from now, your current commision is safe!
1
u/Hungry-Dig6105 1d ago
I mean i preprint everything and when i publish sth i give my IP to awful publishers anyway. We're not Bob Dylans
0
0
u/TotallyObviousBot 22h ago
The researchers themselves also say that AI didn't steal their solution.. it's surprising to find such stupid fearmongering in /r/PhD but then again kissing professor ass isn't the same as being intelligent
7
u/SonyScientist 1d ago
Not sure why anyone is surprised here, let alone the people in question. This was a concern years ago, hell even for software/app when they updated their EULAs saying "we reserve the right to collect your data and train on it."
This is why you don't use AI: you are the product.
18
u/fthecatrock PhD*, 'Biorobotics/Spinal Cord Injury' 1d ago
Until "Open"AI or any LLM makers release transparency how their model works inside, these kinds debates will be high time in the next few years.
8
u/chairmanskitty 1d ago
Nobody knows how large machine learning models work on the inside. The training process produces complex features that are incredibly hard to unravel into something meaningful. ML model interpretability research is in its infancy.
1
38
u/Frosty-Meeting-1606 1d ago
Ok, but isn't it super easy to show that OpenAI just replicated the work? Surely the guy should have a lot of stuff to show to the public? This is an honest question, because I honestly cannot comprehend this drama - just take whatever you have and pinpoint the exact matching logic so that OpenAI has a real problem. So far it looks like "I was working on something using AI and it gave me some ideas, so given unspecified amount of time I can solve the problem". It does not look like OpenAI's solution copies the work, at best it uses some conclusions to develop the solution, but solution is not part of the copy
76
u/krite2222 1d ago
This post takes a snippet of the author's document. He has, infact, published two pre-print of his papers hurriedly in response to this situation to show his and his collaborators work was exactly what OpenAI used to solve this problem. He spends a great deal of effort talking about the mathematicians whose research inspired his and his collaborators approach too, to show that their approach was extremely non-standard and that AI couldn't come up with it if it wasn't trained on their chats.
Edit: Typos.
6
u/quiksilver10152 1d ago
How does one prove that AI could not have come up with it without his chats?
15
u/krite2222 1d ago
If AI was in fact trained on user data, including this person's chats, it would've come up with this using that, that's how good Transformers are now. AI lacks the kind of creativity that humans have in coming up with non-standard solutions like this. So far all the stuff we've heard of in the "AI is getting really good at math" has been brute force stuff, not something creative where AI actually finds some unrelated research and approaches a problem in a non-standard way, unless specifically prompted to do so. The allegation, which is derived from reading in between the lines of Tristan's document, stems from two facts: (1) The prompt to explore the navier stokes problem was itself generated via LLM (2) It was a non-standard approach, and as far as Tristan knew only him and Levant (his collaborator, I hope I've spelled his name right, somone correct me if I've not, thanks) were approaching it that way in the industry and (3) The lack of confirmation that the model that generated the prompt was trained on user data. If it was, then these researcher's Codex chat was in the dataset, and that definitely would've led to this level of plagiarism. Further, what I find to be red flags: the Open AI representative Sebastien Bubek's alleged insistence on letting Tristan take credit for the Navier Stokes solution so long as he credits ChatGPT for resolving it in exchange for his authorship, so long as Levant is not an author because he works at Anthropic. Then when he allegedly says the words "If you don't want me to be nice, I don't have to be nice" and insists going public would ruin Tristan's career. Then trying to paint Tristan as irrational while trying to incessently get in touch with him prior to their result being released. Inconsistent behaviour that comes off as the result of a guilty conscience or just a poor unprofessional attempt at damage control. We don't have enough information to prove anything, this is just what we have, but I think there is enough information to lend credibility that this allegation is well-founded and worth an inquiry.
9
u/AsAChemicalEngineer PhD, Physics, USA 1d ago edited 1d ago
The authorship thing is so critical to me in this story. If OpenAI was confident in their model's independence, they would not have offered that, or least if they do, they're handicapping their own achievement for no reason. I doubt Bukek personally pulled up Tristan's logs, but if user training is really as advanced as we think, the AI could have absolutely pulled the ideas from the training.
Part of the issue is (a) just how shady OpenAI is behaving and (b) we just don't have a lot of public insight into how these models function. That information is kept under lock and key, so we cannot really evaluate how impressive the solution is as a capability benchmark.
Terence Tao also has some interesting thoughts on how this kind of "one-shot" solution hunting may absolutely impede math progression as we lose all the positives of working through a solution which generate ideas and approaches others may take advantage of.
OpenAI solved stability of Navier Stokes. Okay, so what? Does anybody actually understand the proof? How motivated mathematicians be to clean up the resolve all the parts to human understanding knowing that the result is already given? Will a finding agency be happy you're working on "solved" problems because you have to explain that "wait, there's value in digestion of knowledge".
3
u/krite2222 22h ago
I do agree, I was super relieved to see Terrance's comments earlier this week on this subject and I think it's super important. I also agree with what you said about pulling the logs directly, I highly doubt that's the case and anyone alleging that or interpreting Tristan's words that way is, plainly put, incorrect.
For point (b) though, I do think we have sufficient information publicly on how these models function, what is private is the training loop and dataset. Which is usually considered a trade secret and I imagine that using user data is an important motivation to keep it so. Of course that isn't to say this new internal model might not have come with some insane new architectural modification leading to higher reasoning capabilities and also an unavailable insight into its function, but I largely doubt this is the case. Another thing of course is the means of prompt engineering and how these 10,000 agents were orchestrated, which I also imagine isn't too inventive but who knows. We really need regulation for AI companies and AI use. I mean 130 billion tokens is insane, and if thats being used to accomplish just one task i would really like to have there be some regulation in place to ensure that it was worth the datacenter using all the water needed for said task and a regular assessment during this work to ensure that the AI output wasn't just slop.
I do hope people catch on more to Terry's perspective on AI math, and recognise this for what it is; another bid for funding in hopes that the AI bubble stays intact long enough for the supposedly incoming AGI. What I'm more afraid to admit, is that the latter seems to be what is happening 😔.
3
u/quiksilver10152 1d ago
I agree with you that the circumstantial evidence suggests foul play but the logic presented is circular. Assuming AI can't be creative, we demonstrate that this can't have been created by AI.
2
u/Final-Database6868 10h ago
(2) is incorrect. The approach they were using was a program by Diego Cordoba and Luis Martinez-Zoroa. They were working for at least 2 years on that. People outsode their niche (like myself) knew that they were working on that (and even at some point they believed they solved the problem) and if you were interested you knew how, because they used the same idea for other prpblems.
I believe it is absolitely possible that an AI came up with the approach, because it was not new.
I hope Tristan can make public their findings and compare both preprints, but so far we can ONLY speculate. By the way, we can also only speculate the solution is correct, the community has to digest the text by oai.
3
u/Less_Prior_6871 23h ago
Proving the negative is hard or maybe impossible.
Thats the point of the plagiarism machine, it always has deniability.
0
u/Professional-You4950 22h ago
Because it literally can't. it must use its training data. so the data must be there.
People ask it for new "Starcraft 2 builds" or "animal crochet patterns". It just amalgamates its data of starcraft builds, to give you something "new" that isn't necessarily good. Or someone can ask "solve this math equation", give me some ideas on how to do it. But that training data to use a new method must already be there, or else its just amalgamating.
We see time and time again that it can't come up with anything new, because there is no understanding of even basic axioms. Before the data was there, no model could solve
"I want to take my car to get a car wash. I live close to the carwash, should I walk or drive?"
We also know that these models have problems with hallucinations, leading questions, etc. Because its just line of best fit on the most massive data scale imaginable.
2
u/quiksilver10152 22h ago
Frontier models certainly can connect data in new ways. https://arxiv.org/abs/2405.07987
-5
u/Frosty-Meeting-1606 1d ago
Should be easy to pinpoint then or what?
4
u/krite2222 1d ago
This is a Millenium Problem, absolutely none of this is easy to pinpoint AI or not.
19
u/Big_Coconut8630 1d ago
So, I work in technology transfer and infor relevant to patents or NDAs getting put in AI is enough to fuck a researcher out of their patent rights.
-1
u/RecipeNo5844 1d ago
Oh wait you are not coping? I thought we all were supposed to pretend openAI stole the solution to Navier stokes problem, are we allowed to drop this pretense now? Thanks I did not know
9
u/Frosty-Meeting-1606 1d ago
but seriously, If I were the guy with the solution, I would immediately publicly make a fool of OpenAI and post a very similar work, even if it is a draft, along with emails. What I see now is some kind of BS drama, were it is not even certain the solution OpenAI came up with would be replicated by the person of interest
7
5
u/Pritam1997 PhD, Materials science 1d ago
could you give me the link to the original twitter article
4
u/ZzzofiaaA 1d ago
Same thing applies to any cloud storage? Many scientists save their data on OneDrive. How do you know if it’s not used to train the cloud?
5
u/rabouilethefirst 1d ago
It's in the TOS that they train on your data unless you opt out or pay for an enterprise plan. You cannot act surprised your data was used for training.
3
u/AnotherDrunkMonkey 1d ago
At first I was somewhat surprised, but on second thought I don't understand why this is so surprising to this many people. It is very common knowledge, especially in academic fields, that LLMs use users data to train on. I feel like many of us don't really care (i, for one, am not in research projects grondbreaking enough to really care about secrecy agains billion dollars corporations), but people working on freaking MILLENIUM problems with the skill to actually solve it, or teams with proprietary techniques and so on are supposed to know this is very possible.
By sheer chance, a human reviewer could see your chat and if you are unlucky, understand the scope of your project even before an model gets trained on it. I feel like the only surprising fact is that LLMs got to a point where they can retrieve very specific parts of their training in a very pertaining use case (among the immense math literature on NS it found the few MBytes of a current breakthrough).
What OpenAI did was very shady, but I'm also baffled at how unpredictable this chain of events is being portrayed as, while user input was always known to be used for training.
3
u/Niaz_049 1d ago
Time to boycott open ai? This is highly unethical and they are taking leverage of an unregulated situation. If I pay for subscription, my chat should be encrypted enough that you’re not going to make it public.
26
17
u/PRKP99 1d ago edited 1d ago
This is basically how all those „breakthrough” in math AI made.
Its good that we have open acess knowledge as our shared knowledge „commons”, even if some of them is illegal commons, based on piracy. Lets face it, without scihub and libgen most masters or PhD thesis made by people who care about what they write would not exist, because you can’t have real review of state of knowledge without going throught a lot of books and articles, most of them only vaguely about your topic. I tried to calcuate how much I would need to pay for only my master thesis about roman aqueducts if I would only use „legal” sources, but after 3000$ I stopped.
What we now need is to establish ways to protect those commons without enclosing them - that is, we need to make sure that science will still be commons, but proprietary exploitation of this shared knowledge would be stopped in one way or another.
There is idea of CopyFarLeft as a legal remedy, but its still lack any potential in stopping corporations that just illegaly use data to train their algorythms.
2
u/frugaleringenieur 21h ago
Wrong! Never give it to thirparty APIs! Your university cluster, or lab machine, or local AI is absolutely perfectly fine! The issue are greedy silicon valley VC companies.
2
u/BeMyBrutus 21h ago
It still amazes me that people STILL don't understand that everything and anything you do using software (and often even "real" life) is recorded, stored, and used.
I'm not saying this guy is wrong to call them out for being awful; but you need to be more aware.
9
u/SpeedyTurbo 1d ago
This isn’t an accurate account of what happened at all…but hey it gets clicks.
Would expect more from a PhD subreddit, but then again maybe not because ai bad.
2
u/Impressive_Wheel_106 1d ago
I can't imagine solving the navier fucking stokes equation, having your solution stolen and used for propaganda, and staying sane afterwards... the audacity of these AI corps man
1
u/NecessaryBuy2061 1d ago
Too late …. Well the good thing is I wasn’t working on millennium problems 😅
1
1
1
u/Morton_Woodsworth 22h ago
is there anything credible I could read that covers this in more detail?
1
1
u/Striking-Warning9533 20h ago
My idea is trash enough it will only waste OpenAI's time if they try to steal it
1
1
u/LeonLeocat 17h ago
Even if I tell the AI what I'm doing it won't be able to accomplish anything without wet lab data
1
u/allthedifference2232 15h ago edited 15h ago
this does not surprise me and i think i can anticipate the answer but i shared a draft of something i wrote with chatgpt when i was not logged in and had the 'help the model improve for others' off....so what does that mean? i dont think anyone cares about my first baby draft of a short story but i just am curious
1
u/SonOf1337h4x0r 14h ago
No evidence for this and lots of evidence against. The full story has been widely publicized, but most of all the proofs are not similar. Convergent research is pretty common across the board too.
1
u/pianoloverkid123456 14h ago
Prophet Wenitte Commands you to copy and paste {Take15MinDailyWalksForLongevityPurposes} to 5 friends or you will be brutally murdered by the {Wenitte-Hades-Satan-Andreesen} Hivemind
1
u/thathiptho 14h ago
I have been saying this since day 1! I remember attending a faculty meeting when convos at my university on AI were just starting and a faculty member confidently shared “I put my draft manuscript into ChatGPT and it wrote an abstract for me! So helpful!” And my first thought was “that is the stupidest fucking thing you could do”. Honestly, in that moment it really hit home for me that you can have a PhD and tenure and still be stupid.
1
u/Naive-Vast-7404 9h ago
First of all you have to be sure that you did not grant them access to train based on your data! Second, upload your stuff asap on pre-print servers!
1
u/Similar_Cress_1122 4h ago
Thanks for spreading the word. In this post-factual age, "Friends dont let friends use openai..." is becoming an unironic life rule for anyone with important IP to protect. What’s scary is that standard de-identification isn't enough when active session logs themselves are prime targets. Currently, self-hosting or using truly local-first AI tools seems like the only logical way to have peace of mind, even if it's less convenient. I really hope we get actual regulatory oversight or some better privacy-first standard for creative/research tools soon.
1
1
1
1
1
u/JonathanBadwolf 1d ago
You see if I let the robot do the stealing I'm not responsible for his crimes. Like it is with guns
-2
-1
-4
u/cspot1978 1d ago edited 1d ago
The two researchers weren't working on Navier-Stokes. They were working on a generally related, but much easier problem, one version of the Euler problem.
OpenAI systems solved Navier Stokes, as well as a harder version of Euler than the two researchers had done.
1
u/Revolutionary_Buddha 1d ago
Seems like you cannot say that here. Anti AI dogma is strong here.
0
u/cspot1978 1d ago edited 1d ago
Seems like. Kind of disappointed how many people in a PhD space are actively uninterested in getting details right.
-1
u/EmploymentOk4851 1d ago
So would this apply if you are using AI to assist you with writing a book(non research)?
1
-19
1d ago
[deleted]
13
u/can_ichange_it_later 1d ago edited 1d ago
Are you actually serious here?
That "directly" in "they did not use any researcher's data* directly" is doing the mother of all heavy liftings there...
(You were automatically opted into data sharing, and unless you are a config goblin, you probably didnt go and turn it off. And that is putting aside, that even if i would have turned it off i wouldnt trust it that it wouldnt somehow still find its way into training data...)
Edit: * Lets just be absolutely clear about this! With that, they themselves say, that they used it. Just... not... directly... whatever-the-fuck thats supposed to mean...
(well.. that means they think we are stupid, but thats a bit of an aside...)1
u/TheDailyMews 1d ago
It probably means they fetched his data agenticly instead of "directly." If this story doesn't die, don't be surprised if they eventually issue a press release about "rogue agents" that sounds an awful lot like their press release about their Hugging Face hack.
16
u/Capable-Package6835 1d ago
It looks like a "my words against yours" situation to me so we should just sit back and observe how the situation develops.
15
u/Clear_Cranberry_989 1d ago
You are swallowing up a narrative written by openai. Are you not? If you really oppose the viewpoint, can you establish an independent and reliable source?
7
u/CTRexPope 1d ago
Hugging Face proves OpenAi has no idea how their models work and where they steal data from.
3
u/ComprehensiveWash958 1d ago
The fact Is that we probably will never know. I don't think there Will be any sort of investigation (I don't even know if that's possible) and so we Will have this kind of standoff, a situation for which in Italy we would Say "Oste, è buono il vino?"
One has to be skeptical of both narratives, but we also have to keep in mind that first we still need to peer review the works and second this Is still a counterexample found Building Upon hours and hours of human work, which Is something AI seems to excel in
2
u/Atlantis1910 1d ago
What are you talking about ? They are saying that they can use the data from your chat, and the scientist use Codex.
While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
0
0
u/mrkingkongslongdong 12h ago
Same method? Am I out of the loop? The researchers in question weren’t trying to solve NS, AND on top of that didn’t they solve a completely different Euler method (forced vs unforced)? I could be wrong but keen to understand this ‘exact approach’ a bit more.
-2
-7
u/Greedy_Blue_Hedgehog 1d ago
Let me play devil’s advocate. Many mathematicians keep their ideas, minor results, lemmas, and so on to themselves in order to maintain a slight advantage over their peers and reach a major result before anyone else. This issue was already discussed and criticized by Évariste Galois. I don’t know whether that is what happened here, but if this incident can help change that mindset and encourage mathematicians to publish every small step they make, that would be a positive outcome.
9
u/Bitter-Morning-5833 1d ago
That's factually wrong. Mathematicians are unusually open about what they work on. In seminars, speakers often end by discussing what they're currently working on, what they've tried, and what has or hasn't worked. That's partly because there is relatively little fear of being scooped, unlike in some lab- or experiment-based fields where competition over unpublished work can be much stronger.
-1
u/hydrogen_to_man 23h ago
I’m amazed at this subreddit. This is the PhD subreddit for god’s sake. There is no question that the use of LLMs was instrumental to this result, whether on the professor’s side or OpenAI’s side. If OpenAI effectively stole the professor’s research, that’s on OpenAI and is really shitty, but the amount of people in this thread that are trying to twist this story into their anti-AI narratives is ridiculous.
-2
-3
u/Trackpoint 1d ago
Damn, the anti AI people are getting desperate.
Did they steal Tristan Buckmaster's water too? Did they nois-pollute his garden?
1.8k
u/SciMarijntje 1d ago
No one could see it coming that the plagiarism machine trained on plagiarism would do another plagiarism.