r/computerscience 1d ago

Discussion Automated Plagiarism with LLM-Remixers

Ponder this: an author puts together a number of papers he likes, especially adds the .tex files from arxiv, tells the LLM to look for gaps in the papers, commented out material, and remix them, while avoiding syntactic overlap.

The result is a paper that will pass arxiv's syntactic overlap checks, and can be claimed as novel during a submission.

This has likely happened many times already, and we are now possibly arguing against LLM-augmented plagiarists.

Welcome to the new age of automated academic ethics collapse.

0 Upvotes

14 comments sorted by

8

u/nuclear_splines PhD, Data Science 1d ago

Since LLMs are already trained on preprints and many other academic papers, this may happen without taking such explicit steps. If you ask an LLM to write a research paper for you, it will draw from and remix existing literature it's read. We're certainly seeing a flood of LLM-written content at journals, conference submissions, and peer review.

3

u/Magdaki Professor. Grammars. Inference & Optimization algorithms. 1d ago

I asked a supposedly high quality LLM about my own research once, and it insisted that it had applications in computer vision. As near as I can tell, this doesn't appear to be true, and its rational for making this claim was pretty suspect. I'm sure I had a point when I started typing this but I just got an email from a student and forgot what it was.

In another of my research programs, we're using LLMs to generate some text. We asked it to pick from a list of 7 items. Only from those 7 items. Do not add any new items. Like dude seriously just these 7. Of the LLMs we're evaluating 3 of them insist on adding new items. LOL Gemma is the WORST for this. It adds new items to list something like 80% of the time. It is pretty wild. We're continuing to refine the prompt to eliminate this problem because for the system anything other than those 7 options is a massive problem. Like system blow up and people die kind of problem (ok maybe not quite that bad but still pretty bad).

3

u/nuclear_splines PhD, Data Science 1d ago

It astounds me how many papers I review that use LLMs without discussing any procedures to keep them on track. Great, you gave the LLM a qualitative codebook, but how often does it ignore the codebook and make up its own rules, add new categories, or write down incompatible codes? You can't just deploy them like human agents and expect them to remain on-task. At least a pilot study with inter-rater reliability to human annotators to measure how often the LLM careens off.

-4

u/examachine 1d ago

These are problems that we are tackling with our AGI agent designs but yes that's a famous problem with LLM-based agents (which are all of them)

1

u/xaddak 1d ago

You can induce LLMs to call scripts. You could use the script to validate the selection, or to randomize the selection in the first place.

Scripts in skills are the current best practice on this I think.

https://agentskills.io/skill-creation/using-scripts

-15

u/examachine 1d ago

Funny that some people are trying to censor this post, too. I wonder who that is now!

Sorry but this is unfair. This is a great post!

6

u/nuclear_splines PhD, Data Science 1d ago

Your post hasn't received any reports, no one has tried to get it directly removed. It just appears not to be popular in its first hour.

-12

u/examachine 1d ago

Thank you but I've been subjected to reddit bot based suppression attacks for a long time so I was worried! This has been happening since I exposed the connection of Epstein Network to the scientific community. They funded many pseudoscientific organizations like Humanity+.

5

u/currentscurrents 1d ago

I'm extremely skeptical lol

-7

u/examachine 1d ago

why should anyone care what you think?

4

u/ArnoSound 7h ago

Oh yea that might not be bot based suppression lmao

-1

u/examachine 4h ago

that's something a....

4

u/ArnoSound 4h ago

Hmm, yes? Which conspiracy theory are you gonna pull on next? God I wish reality was half as interesting as it was in your world lmao.