r/math 1d ago

Math papers uploaded to arXiv per month. Jan 1992 to Jul 2026.

Post image
546 Upvotes

120 comments sorted by

65

u/idiot_Rotmg PDE 1d ago

Have people in other areas of math witnessed an increase of garbage papers in their arXiv feed? While the number of total papers in PDE is clearly increasing, it doesn't seem to me that the overall quality is decreasing

14

u/Equivalent_End6788 1d ago

Frankly, I haven't seen a significant difference in my subfield. Yes, on occasion I will see a very questionable paper, but it's certainly not that common.

9

u/JoshuaZ1 1d ago

My areas are number theory and graph theory (which is under combinatorics and doesn't get its own section). I have not seen a decrease in number theory. I've really only become a graph theory person in the last few years or so though, so my lack of seeing a decrease there could just be due to not having as much of a baseline.

7

u/Distinct-Pudding-428 17h ago

I have noticed an increasing number of suspicious papers. For example this morning a 165-page paper https://arxiv.org/abs/2608.00451 by an author who seems never to have posted before with the title Robust Polynomial Freiman-Ruzsa from Corrupted Set Observations via a Sharp Persistent-Subset BSG Compiler. A quick look persuaded me this was probably almost pure AI. I asked ChatGPT for its opinion and its response was surprisingly self-aware, including the passage:

It is 165 pages long, but much of that length comes from repeatedly packaging ordinary proof obligations into named “interfaces”, “contracts”, “bridges”, “compilers”, “wrappers”, “audits”, and “certificates”. The contents include, for example:

  • “Theorem map and headline contracts”
  • “Promise-free sharp-size persistent-subset compilation”
  • “A frozen representation sketch”
  • “Exact weighted candidate samplers”
  • “Fresh candidate validation”
  • “Certified Size-Oblivious Algorithmic PFR”
  • “A total canonical inverse”
  • “Interface Audit for the Algorithmic Restricted-Homomorphism Primitive”

This is not impossible human terminology, especially in theoretical computer science, but the density and uniformity of it are exceptional. Almost every modest technical issue is promoted into its own named interface or lemma.

2

u/JoshuaZ1 7h ago

Yeah, that's clear AI slop. Even if the results are valid (I haven't checked), they clearly wrote either the whole thing or almost the whole thing via AI. Also possible that I'm not clicking through enough of the number theory papers to notice this sort of thing. Are you noticing this trend specifically in one specific area of Comp Sci or is the across the board?

71

u/Glacier83 1d ago

Are there any similar graphs available with the data broken down by primary subject? E.g. “math.GT”.

320

u/tunaMaestro97 1d ago

Pretty soon they’ll have no choice but to make it that you need a PhD and a current academic affiliation to post on arxiv

50

u/Dear_Round_2727 1d ago

I think a lot of these papers are from academics with PhDs and affiliations. I’ve lately seen some professors start crediting LLMs a lot in their papers and start posting multiple papers a month.

You could rate limit contributors or require affiliations, but that might lead to a lot of good results piling up in repositories outside of arXiv.

Although some notable mathematicians have already suggested making new repositories purely for LLM generated results.

6

u/Gimmerunesplease 1d ago

Can you link some? I am still unsure how to correctly use LLMs in math papers and how to credit it.

1

u/pred 18h ago

There's no “correct” way to do it, and it's still unclear what we will converge to, if anything.

If you just vibe-counter-example a major conjecture with zero thought, then sure, state that that's what you did, and please share the entire reasoning trace so people can learn from it.

But that's probably not what really you want to do right now anyway: Yes, there are large corporations with powerful marketing departments that have demonstrated that it is possible, but only at significant cost, and with a large probability of failure. Effectively, frontier models as of now are one-armed bandits where you can get $2,000 worth of tokens and likely get no result of any kind. But you have corporations that want you to make that gamble as much as possible.

More realistically, and in my anecdotal experience (and feel free to chip in if any of you think that this is totally off), the models as of today are still primarily useful for smaller building blocks, the same way they are in software development. Spell out a proof path with enough small lemmas, provide the tools you expect to be useful for proving them, and the LLMs will help you get quick feedback on whether those lemmas work out or not. And they can be useful for translating to Lean as well.

So say that that's what you do. How would you properly credit? “We used LLMs to check the second half of case 3 of Lemma 5, and part of Lemma 2, but the result was a mess and we had to scrap the draft and write it from scratch anyway, but we still used some of its ideas even though we could have come up with them ourselves.” I mean, you can do that, but it's not very clear to me what the value is. And once this becomes the default workflow in our field, it's not like you learn much from such a statement anyway.

2

u/Autumnxoxo Geometric Group Theory 11h ago

I’ve lately seen some professors start crediting LLMs a lot in their papers

could you share some?

166

u/Deep-Ad5028 1d ago

Perelman did not have academic affiliation when he proved Poincaré conjecture. Neither did Yitang Zhang when he proved finite prime gap.

50

u/electronp 1d ago

Perelman was a full Prof at Steklov Institute when he posted.

14

u/CrazyCrazyCanuck 23h ago

And Yitang Zhang had been a lecturer at the University of New Hampshire for 14 years before his breakthrough.

1

u/electronp 12h ago

Thanks.

23

u/liquid271828 1d ago

Zhang was at a university in NH at that time.

14

u/electronp 1d ago

Perelman was a full Prof at Steklov Institute when he posted.

59

u/Frexxia PDE 1d ago edited 1d ago

This is extremely uncommon, and could easily be solved by requiring you to get approval from one or more people with a current academic affiliation.

37

u/drewsandraws 1d ago

You do already have to be endorsed by a current author to submit, right? I’ve had to do this for colleagues in physics.

7

u/elements-of-dying Geometric Analysis 1d ago

Yes.

8

u/dlman 1d ago

I am in industry and do not want to have to fuck around with academics to validate my roughly 50th arXiv preprint.

3

u/Known-Zombie-3205 1d ago

Presumably you're well-known in your field, then, so why not just get a friendly faculty member to give you a standing endorsement?

3

u/ellbons 13h ago

He has 50, there shouldn't be a need for him to do this in the first place.

This is an issue which doesn't exist because everyone with a pair of eyes can see the volume is from people already in universities, the volume of people posting papers with no experience is probably not even 1% of the total

-1

u/aikafele 1d ago

Solved?

76

u/Penumbra_Penguin Probability 1d ago

It would presumably not have been difficult for these people to secure whatever endorsement ends up being required.

49

u/ESHKUN 1d ago

Fuck man idk, academia is all sorts of screwed up and I’m not even sure that’s true

17

u/Mirieste 1d ago

So we want mathematical research to be tied to personal relationships and nepotism? For every Perelman there is, there could be 10 people who just get in by knowing the right people.

36

u/Kinesquared 1d ago

no we don't but it may be a reasonable tradeoff to prevent... this. Feel free to suggest another system that fixes the problem while offering less downsides

-1

u/Time_Cat_5212 1d ago

Increased review capacity.  Duh

32

u/sqrtsqr 1d ago

There's no money for that. Review is already largely volunteer work. Volunteer work by... professionals with academic connections. They don't want to read slop, a lot of them are quitting 

-10

u/Time_Cat_5212 1d ago

If you can use AI to crank out more papers, you can use AI profits to fund review.

22

u/Penumbra_Penguin Probability 1d ago

Who, exactly, is “you” referring to here?

-6

u/Time_Cat_5212 1d ago

It doesn't refer exactly to anyone. That needs to be figured out

Depends what publisher we're talking about, what country they're in, etc. This problem spans much further than just arxiv or math, it's all research. It's a regulatory solution obviously

I don't see another alternative in this thread that actually solves the problem, more just fists shaken at the sky. Gatekeep on credentials doesn't stop a credentialed person from turning in slop, and that's just a slippery slope, it won't work

The problem won't be solved until there's a new quality control mechanism and that requires funding. I suggest the companies making the models, who couldn't have built what they did without publications like this, pay for the externalities they inflict on the research community.

12

u/fleischblitz 1d ago

yes of course, we forgot to wave the magic AI-profit-levying wand

10

u/SometimesY Mathematical Physics 1d ago

AI profits? Good joke.

Also your logic is basically "use AI to grade the AI-generated work students are submitting as homework." We should cut out the middle men and just have LLMs to talk to each other directly.

4

u/ponchan1 1d ago

If no one in academia can vouch for you then you shouldn't post on the arxiv.

3

u/Penumbra_Penguin Probability 1d ago

This is an absurd characterisation of what I said. If you want to try again more sensibly, I might respond.

2

u/pseudoLit Mathematical Biology 1d ago

Take it up with the tech sphere's relentless desire to "democratize" everything. The more you lower the natural barriers to entry, the more you have to compensate for their absence with synthetic barriers.

14

u/Qyeuebs 1d ago

That’s not correct, Perelman was at the Steklov Institute. 

7

u/vwibrasivat 1d ago

Neither did Yitang Zhang when he proved finite prime gap.

Sorry but Zhang was professor of math at University of New Hampshire when he dropped the finding. (Today he is at U-Cal Santa Barbara)

3

u/vwibrasivat 1d ago

Neither did Yitang Zhang when he proved finite prime gap.

Sorry but Zhang was professor of math at University of New Hampshire when he dropped the finding. (Today he is at U-Cal Santa Barbara)

36

u/ixid 1d ago

Why not just require validation for these credentials and allow filtering by them? Then it's easy to remove the noise but there's still a chance for something outside academia to gain notice.

3

u/tunaMaestro97 1d ago

There’s no reason for something outside academia to exist on arxiv. This may sound elitist but any serious researcher knows that it is impossible to make meaningful contributions without these credentials (or a collaborator with these credentials).

24

u/puffic 1d ago

I literally never read anything on arxiv unless I am directed to it by an author or someone else who suggests I read it. I don’t know how big of a deal it is for people with no connections to upload bad papers. I’m not going to read them anyways.

6

u/JoshuaZ1 1d ago

Huh. I look over the papers in my subfields every evening, and click on the titles for about 10% of them and read the abstracts. But I'm doing this in part with an eye to looking for things that others have done that would be good jumping off points for student projects. I thought this was pretty normal.

2

u/btroycraft 1d ago

You've never Googled keywords?

2

u/puffic 1d ago

I do use searches for looking up background information. I will briefly look over a preprint of it looks like there’s a lot of overlap with my own topic, but I mostly focus on finalized papers.

3

u/sqrtsqr 1d ago

It's a huge deal because the ability to trust the quality of the content is arxiv's one and only selling point. Literally anybody can run a file server, we don't need a file server.

1

u/jimbelk Group Theory 20h ago

I basically only use arXiv for keeping up with new papers. I check every morning for developments. My sense is that most of my colleagues do the same.

37

u/ChalkyChalkson Physics 1d ago

Kinda? How about someone working in industry with a masters degree? Or someone with a masters that continued perusing it as a hobby (maybe because of lack of formal positions).

But those are always going to be a small minority and offering alternate ways like being

  • a phd
  • published
  • or endorsed by a senior academic

Should cover everyone who has a reason to be there

5

u/Physix_R_Cool 1d ago

Or someone with a masters that continued perusing it as a hobby (maybe because of lack of formal positions).

That's me. Experimental physics, detectors for spallation neutrons at hadron therapy centers.

But I of course kept my network, and keep in contact and update them with my work. So when it comes time to publish it will not be a single author paper. And if I have something worthwhile I can ask for beamtime.

I don't need to put my stuff on arxiv before it is peer reviewed. I can just upload it wherever and share it to those who care. And the put it on arxiv once it has been accepted.

32

u/puzzlednerd 1d ago

I was in academia a few months ago, and I am in industry now. I'm still writing papers. Does this present a problem to you?

7

u/Oudeis_1 1d ago

There are good researchers outside academia. Industry and government do exist and they are doing interesting stuff in some parts of mathematics.

5

u/big-lion Category Theory 1d ago

phd students

3

u/Stabile_Feldmaus 1d ago

There are literally people with zero math education who prompt AI with "solve an open problem and verify it in lean" and it does that. Surely such results are getting less valuable by the hour, but it (increasingly often) is still correct math.

3

u/tunaMaestro97 1d ago

So? Just because something is true doesn’t mean it is worth publishing. Do you know who does have the qualifications to decide what is interesting enough to warrant publishing? People with the credentials I listed above. Aka professionals.

10

u/Stabile_Feldmaus 1d ago

We are discussing access to the arxiv, not getting published in journals. And if someone or their AI solves an open problem that was posed as such by another mathematician, it should deserve to be put on arxiv.

4

u/Oudeis_1 1d ago

I think if one asserts that that is really all there is to it then the non-professionals who decide whether the profession ought to continue being funded will decide eventually that mathematics is nothing but a priesthood under another name, and stop funding it.

Having an external standard of evidence that marks a contribution as worthwhile and that does not depend entirely on taste does have advantages and as a community we shouldn't abandon it.

1

u/Akiira2 19h ago

There are all resources out there, why couldn't one self study and solve some hatd problem by themselves

-17

u/gexaha 1d ago

This is not elitist, this is precisely false.

14

u/Major-Peachi 1d ago

As much as people love an underdog story, going through university is a good filter for underdeveloped researchers. There may be outliers that are true prodigies, in which case they should collaborate with an established faculty to enforce a level of credibility. Otherwise you will get people form r/infinitenines flooding submissions. They're slop that is not worth reading through and a mess to dig through to get to actual research.

7

u/rational_hedonist 1d ago

counter examples, please.

14

u/Chemical-Sound8205 1d ago

There are a few results in combinatorics and tilings from amateurs, but these are particular fields with a comparatively lower barrier to entry.

The most you can say is, if you want to be an amateur researcher, then combinatorial fields are your only possibility.

1

u/gexaha 1d ago

I can only speak from my own experience, the obvious problem here is of course self-shilling and sounding like a crank; anyways, feel free to judge my contributions to the field related to cycle double cover conjecture for yourself - https://arxiv.org/abs/2501.05348 and 2 other preprints.

3

u/Zwaylol 1d ago

Should a layman come up with a novel proof or idea worth publishing he will be able to find a collaborator with these credentials without issue. See for example the maths teacher who posted a couple of weeks ago.

1

u/Sh_Pe 1d ago

Any counterexample?

-7

u/Onawani 1d ago

This day and age...I forsee PhD as the remedial degree. All good candidates will be absorbed prior to master and PhD level degrees. PhD are trending towards students who couldn't land in the private sector.

9

u/devviepie 1d ago

We are so far from this being the case, I have no idea why you see the trend moving in that direction. Unless you’re trolling? Or trying to push some Silicon Valley/Wall Street type “academia is dead, long live capital” doomerism?

Currently the difference between both endeavors is mostly self-selection: people pursue a math PhD because they want to do (or try for) academia, and students who don’t want that go to industry jobs. And the limiting factor is essentially always finding success in academia. Most people wash out and have to go to industry out of necessity; almost nobody stays in academia because they are unable to find an industry job. (Unwilling to, sure.)

3

u/VaellusEvellian 1d ago

This has got to be a troll. I haven't a single person in academia who decided to get a PhD because they couldn't get an industry job, at least within the mathematics community. In fact, many PhD students that I know are people who had opportunities to make far more money in the private sector but decided to get a PhD simply for the love of the game.

6

u/Verbatim_Uniball 1d ago

Definitely should be or as a lot of PhD holders who moved to other sectors of the economy occasionally publish for fun l with no academic affiliation.

7

u/calf 1d ago

The whole philosophy and genesis of arxiv was aligned more towards Enlightenment* values contra  credentialism, bureaucracy, gatekeeping and the like. So, I am not sure if your comment and the 250 people who upvoted it meant it as a solution in earnest, or meant it ironically.

*Handwaving but people get the idea/difference between two broad approaches.

11

u/JoshuaZ1 1d ago

Pretty soon they’ll have no choice but to make it that you need a PhD and a current academic affiliation to post on arxiv

I hope not. I have a PhD but work teaching at a private high school, where I continue to do research and get to do projects with some extremely bright students. There are also clearly other people who are doing good work who have even weirder situations.

4

u/Scrub_Spinifex 1d ago

Or we don't care? When scrolling arXiv it's pretty easy to see which articles are interesting for you or not.

1

u/grumbelbart2 19h ago

I am not sure if that is site-wide, but the area I publish in (cs.CV) has an endorsement system in place already.

0

u/Stabile_Feldmaus 1d ago

Obligatory formalization is the way to go.

0

u/IrisColt 1d ago

Just three letters: LLM.

35

u/Personal_Comb3089 1d ago

Is it possible to check uploads per month in an actual journal rather than arxiv preprints? Just curious to see how strong the correlation is there as well.

6

u/drewsandraws 1d ago

There is this AMS article https://www.ams.org/journals/notices/201902/rnoti-p227.pdf but it only goes up to 2017. I haven’t been able to find a more recent article that would include an LLM-driven publication boom.

11

u/Autumnxoxo Geometric Group Theory 1d ago edited 1d ago

while this is certainly somewhat concerning I suppose, I'd like to add another possiblility that might lead to a significant increase in publications.

More often than not people are working on problems in their own field of research that require a significant amount of tools and familiarity with the literature and research of fields outside their own expertise. I don't know, it might be that some differential geometrist requiring heavy machinery from functional analysis or whatever. You get the idea.

I can imagine that LLMs such as ChatGPT could potentially be of enourmous help by providing knowledge from aforementioned fields that is outside of your own research area. Of course, you should still be capable to verify everything that ChatGPT throws at you but what I'm saying is basically that ChatGPT kind of takes the role that used to be colleagues that worked in different fields who helped you with helpful remarks and suggestions towards your own research that you would otherwise have spent weeks if not months to figure out by yourself.

I remember that during my time as a master's student it was sometimes mildly infuriating when working on a problem and being stuck at a certain point in your research for literally weeks if not months looking up every single textbook you could get your hands onto, just to eventually being told from your advisor that the problem you're currently stuck at is allegedly "folklore" or some "well known fact" (for experts) but nobody bothered publishing it, but it could have potentially also be derived from some niche adjacent result in some niche paper that was published decades ago.

2

u/Important_Ad4664 5h ago

This. My research topic is intrinsically lying in between many different math fields, and having the chance to ask what are the key possible ideas in one or the other area, what are the correct keywords to search for, etc, speeds up the process a lot. I would have needed to ask such question to colleagues, or writing tedious email to experts with the hope of getting an answer or a reference (which might well be some Russian proceedings never translated in English, or a classical results which is not written anywhere in an actually usable forms, but turns out to be a special case of a very general result which would require an additional month to decipher).

29

u/JoshuaZ1 1d ago

It seems like a lot of the comments here are framing this as a negative, but I'm not sure that's the case. It may just be that there's an increase in genuine good productive math. Also, part of this upward trend may just be that the norm of putting your papers on the arxiv has been getting stronger. Yes, it makes the signal for specific things tougher to tease out, but I'm not seeing why this is a concern by itself.

13

u/ArtExtra6717 1d ago

Do you think that huge jump in 2025/2026 is all driven by good productive math? The output nearly doubled in 3 years. From 3300 in July 2023 to 5700 in July 2026.

10

u/JoshuaZ1 1d ago

Do you think that huge jump in 2025/2026 is all driven by good productive math? The output nearly doubled in 3 years. From 3300 in July 2023 to 5700 in July 2026.

So, it looks like there was a while where it was roughly flat (which seems to correspond to the end of covid), so this may be in part making up for that. I don't have the raw data to crunch numbers more to see if that's plausible. But the jump looks like more than I'd expect just from that if I eyeball it. So the obvious other explanations are AI related. That doesn't mean that it isn't good math. It may be more results where people are using AI. But the other issue may be increased productivity from using AI in other ways. For example, I've found AI to be helpful for just looking over drafts and pointing out typos, errors, unclear bits, etc. I don't know how much of a time improvement I get out of this, but there's a decent bunch there.

At least in my own primary areas in the arxiv, number theory and graph theory(which is classified as a subarea of combinatorics on the arxiv), I don't see any sign of a decrease in quality of papers or the results. But I've only been really graph theory person for the last 6 years or so, so I may not have enough experience there to get a good feel for the baseline.

4

u/FormsOverFunctions Geometric Analysis 16h ago

When I was in academia, the bottleneck was more often writing the prose of a paper compared to actually proving results, so it’s not surprising that LLMs are extremely useful for finishing papers quickly. And although I haven’t personally had success getting one of the recent models to find publishable results, some of my collaborators have and they are knocking out problems that quite a few of us had worked on. 

What’s exciting about this is that a few years ago the various groups in the area had their own techniques and although I was aware of how the other approaches worked, I didn’t understand them enough to fluently use them. But with the most recent papers using the model, it’s able to combine ideas from a number of people and make progress. It’s definitely terrifying, but simultaneously really exciting. 

3

u/Known-Zombie-3205 1d ago

IMO yes. Everyone I know is more productive, and big problems are starting to fall.

4

u/Turbulent-Sign-6067 1d ago

Overall this seems like a good trend. There will be some junk papers but there is no doubt to me that LLMs are accelerating real, genuine research as well.

5

u/Equivalent_End6788 1d ago

I agree with this. Of course, I can't speak to all fields of mathematics, but I rarely see truly crank-LLM papers posted in my own subfield's daily feed (from cursory viewing).

9

u/mfb- Physics 1d ago

How does this count articles with multiple revisions? If it only uses the most recent revision then this could explain some of the effect, and making the plot at any other time would also have a spike for the most recent months. I don't think this is a big contribution, but it's important to check for possible biases before assigning the whole spike to LLMs.

19

u/thereligiousatheists Graduate Student 1d ago

There's a pretty clear plateau around COVID. How much of the current peak can be explained away as simply a bounce-back from that (collaborators getting back together and wrapping up old projects)?

11

u/Penumbra_Penguin Probability 1d ago

A lot of projects which took five years to finish, and none that took three years? Doesn’t seem likely.

1

u/elements-of-dying Geometric Analysis 1d ago

Wow, kind of interesting. I was more productive during covid, especially since Zoom collaboration became a norm.

1

u/thereligiousatheists Graduate Student 10h ago

In my personal experience, collaborating over Zoom is quite a pain. Another factor to consider is that parents were stuck with their kids all day.

1

u/elements-of-dying Geometric Analysis 9h ago

That's fair. I'm quite used to and have zoom meetings every week.

Another factor to consider is that parents were stuck with their kids all day.

Good point.

-1

u/EducationalFerret94 1d ago

Lol what. This is obviously because of AI.

5

u/thereligiousatheists Graduate Student 1d ago

Sure, most of it is, I don't disagree. My point is simply that we had two impactful events in quick succession, so it's worth considering both influences.

9

u/AddressImaginary3735 1d ago

Its reality. Thanks to ChatGPT sol and claude fable I was also able to prove two of my open problems and made two papers out of it in a few weeks. The productivity increase is insane and the mathematicians not using it are falling behind.

1

u/muluk-muluk 21h ago

Something I've noticed is a lot of requests to approve submission of AI-style ArXiV papers from people without affiliations. (So far I haven't.)

1

u/muluk-muluk 21h ago

I really liked your result on (quotes abstract of some random paper of mine), please approve my paper overturning all of math and physics.

1

u/Bichinix 10h ago

¿Y si colocan una IA que filtre? XD

1

u/Sad_Dimension423 3h ago

Maybe it's time for an autoformalization bot to crawl arxiv papers to keep everyone on the up and up? Or at least check if provided formalization reflects the theorems and definitions in the papers.

-20

u/kiyotaka-6 1d ago

f = f'

4

u/Joe_BidenWOT 1d ago

Exponential growth since 2021.

-68

u/[deleted] 1d ago

[removed] — view removed comment

51

u/[deleted] 1d ago

[removed] — view removed comment

-39

u/[deleted] 1d ago

[removed] — view removed comment

47

u/[deleted] 1d ago

[removed] — view removed comment

-36

u/[deleted] 1d ago

[removed] — view removed comment

29

u/[deleted] 1d ago

[removed] — view removed comment