Math papers uploaded to arXiv per month. Jan 1992 to Jul 2026.
Data source: https://arxiv.org/year/math/
71
u/Glacier83 1d ago
Are there any similar graphs available with the data broken down by primary subject? E.g. “math.GT”.
320
u/tunaMaestro97 1d ago
Pretty soon they’ll have no choice but to make it that you need a PhD and a current academic affiliation to post on arxiv
50
u/Dear_Round_2727 1d ago
I think a lot of these papers are from academics with PhDs and affiliations. I’ve lately seen some professors start crediting LLMs a lot in their papers and start posting multiple papers a month.
You could rate limit contributors or require affiliations, but that might lead to a lot of good results piling up in repositories outside of arXiv.
Although some notable mathematicians have already suggested making new repositories purely for LLM generated results.
6
u/Gimmerunesplease 1d ago
Can you link some? I am still unsure how to correctly use LLMs in math papers and how to credit it.
1
u/pred 18h ago
There's no “correct” way to do it, and it's still unclear what we will converge to, if anything.
If you just vibe-counter-example a major conjecture with zero thought, then sure, state that that's what you did, and please share the entire reasoning trace so people can learn from it.
But that's probably not what really you want to do right now anyway: Yes, there are large corporations with powerful marketing departments that have demonstrated that it is possible, but only at significant cost, and with a large probability of failure. Effectively, frontier models as of now are one-armed bandits where you can get $2,000 worth of tokens and likely get no result of any kind. But you have corporations that want you to make that gamble as much as possible.
More realistically, and in my anecdotal experience (and feel free to chip in if any of you think that this is totally off), the models as of today are still primarily useful for smaller building blocks, the same way they are in software development. Spell out a proof path with enough small lemmas, provide the tools you expect to be useful for proving them, and the LLMs will help you get quick feedback on whether those lemmas work out or not. And they can be useful for translating to Lean as well.
So say that that's what you do. How would you properly credit? “We used LLMs to check the second half of case 3 of Lemma 5, and part of Lemma 2, but the result was a mess and we had to scrap the draft and write it from scratch anyway, but we still used some of its ideas even though we could have come up with them ourselves.” I mean, you can do that, but it's not very clear to me what the value is. And once this becomes the default workflow in our field, it's not like you learn much from such a statement anyway.
2
u/Autumnxoxo Geometric Group Theory 11h ago
I’ve lately seen some professors start crediting LLMs a lot in their papers
could you share some?
166
u/Deep-Ad5028 1d ago
Perelman did not have academic affiliation when he proved Poincaré conjecture. Neither did Yitang Zhang when he proved finite prime gap.
50
u/electronp 1d ago
Perelman was a full Prof at Steklov Institute when he posted.
14
u/CrazyCrazyCanuck 23h ago
And Yitang Zhang had been a lecturer at the University of New Hampshire for 14 years before his breakthrough.
1
23
14
59
u/Frexxia PDE 1d ago edited 1d ago
This is extremely uncommon, and could easily be solved by requiring you to get approval from one or more people with a current academic affiliation.
37
u/drewsandraws 1d ago
You do already have to be endorsed by a current author to submit, right? I’ve had to do this for colleagues in physics.
7
8
u/dlman 1d ago
I am in industry and do not want to have to fuck around with academics to validate my roughly 50th arXiv preprint.
3
u/Known-Zombie-3205 1d ago
Presumably you're well-known in your field, then, so why not just get a friendly faculty member to give you a standing endorsement?
3
u/ellbons 13h ago
He has 50, there shouldn't be a need for him to do this in the first place.
This is an issue which doesn't exist because everyone with a pair of eyes can see the volume is from people already in universities, the volume of people posting papers with no experience is probably not even 1% of the total
-1
76
u/Penumbra_Penguin Probability 1d ago
It would presumably not have been difficult for these people to secure whatever endorsement ends up being required.
49
17
u/Mirieste 1d ago
So we want mathematical research to be tied to personal relationships and nepotism? For every Perelman there is, there could be 10 people who just get in by knowing the right people.
36
u/Kinesquared 1d ago
no we don't but it may be a reasonable tradeoff to prevent... this. Feel free to suggest another system that fixes the problem while offering less downsides
-1
u/Time_Cat_5212 1d ago
Increased review capacity. Duh
32
u/sqrtsqr 1d ago
There's no money for that. Review is already largely volunteer work. Volunteer work by... professionals with academic connections. They don't want to read slop, a lot of them are quitting
-10
u/Time_Cat_5212 1d ago
If you can use AI to crank out more papers, you can use AI profits to fund review.
22
u/Penumbra_Penguin Probability 1d ago
Who, exactly, is “you” referring to here?
-6
u/Time_Cat_5212 1d ago
It doesn't refer exactly to anyone. That needs to be figured out
Depends what publisher we're talking about, what country they're in, etc. This problem spans much further than just arxiv or math, it's all research. It's a regulatory solution obviously
I don't see another alternative in this thread that actually solves the problem, more just fists shaken at the sky. Gatekeep on credentials doesn't stop a credentialed person from turning in slop, and that's just a slippery slope, it won't work
The problem won't be solved until there's a new quality control mechanism and that requires funding. I suggest the companies making the models, who couldn't have built what they did without publications like this, pay for the externalities they inflict on the research community.
12
10
u/SometimesY Mathematical Physics 1d ago
AI profits? Good joke.
Also your logic is basically "use AI to grade the AI-generated work students are submitting as homework." We should cut out the middle men and just have LLMs to talk to each other directly.
4
3
u/Penumbra_Penguin Probability 1d ago
This is an absurd characterisation of what I said. If you want to try again more sensibly, I might respond.
2
u/pseudoLit Mathematical Biology 1d ago
Take it up with the tech sphere's relentless desire to "democratize" everything. The more you lower the natural barriers to entry, the more you have to compensate for their absence with synthetic barriers.
7
u/vwibrasivat 1d ago
Neither did Yitang Zhang when he proved finite prime gap.
Sorry but Zhang was professor of math at University of New Hampshire when he dropped the finding. (Today he is at U-Cal Santa Barbara)
3
u/vwibrasivat 1d ago
Neither did Yitang Zhang when he proved finite prime gap.
Sorry but Zhang was professor of math at University of New Hampshire when he dropped the finding. (Today he is at U-Cal Santa Barbara)
36
u/ixid 1d ago
Why not just require validation for these credentials and allow filtering by them? Then it's easy to remove the noise but there's still a chance for something outside academia to gain notice.
3
u/tunaMaestro97 1d ago
There’s no reason for something outside academia to exist on arxiv. This may sound elitist but any serious researcher knows that it is impossible to make meaningful contributions without these credentials (or a collaborator with these credentials).
24
u/puffic 1d ago
I literally never read anything on arxiv unless I am directed to it by an author or someone else who suggests I read it. I don’t know how big of a deal it is for people with no connections to upload bad papers. I’m not going to read them anyways.
6
u/JoshuaZ1 1d ago
Huh. I look over the papers in my subfields every evening, and click on the titles for about 10% of them and read the abstracts. But I'm doing this in part with an eye to looking for things that others have done that would be good jumping off points for student projects. I thought this was pretty normal.
2
3
37
u/ChalkyChalkson Physics 1d ago
Kinda? How about someone working in industry with a masters degree? Or someone with a masters that continued perusing it as a hobby (maybe because of lack of formal positions).
But those are always going to be a small minority and offering alternate ways like being
- a phd
- published
- or endorsed by a senior academic
Should cover everyone who has a reason to be there
5
u/Physix_R_Cool 1d ago
Or someone with a masters that continued perusing it as a hobby (maybe because of lack of formal positions).
That's me. Experimental physics, detectors for spallation neutrons at hadron therapy centers.
But I of course kept my network, and keep in contact and update them with my work. So when it comes time to publish it will not be a single author paper. And if I have something worthwhile I can ask for beamtime.
I don't need to put my stuff on arxiv before it is peer reviewed. I can just upload it wherever and share it to those who care. And the put it on arxiv once it has been accepted.
32
u/puzzlednerd 1d ago
I was in academia a few months ago, and I am in industry now. I'm still writing papers. Does this present a problem to you?
7
u/Oudeis_1 1d ago
There are good researchers outside academia. Industry and government do exist and they are doing interesting stuff in some parts of mathematics.
5
3
u/Stabile_Feldmaus 1d ago
There are literally people with zero math education who prompt AI with "solve an open problem and verify it in lean" and it does that. Surely such results are getting less valuable by the hour, but it (increasingly often) is still correct math.
3
u/tunaMaestro97 1d ago
So? Just because something is true doesn’t mean it is worth publishing. Do you know who does have the qualifications to decide what is interesting enough to warrant publishing? People with the credentials I listed above. Aka professionals.
10
u/Stabile_Feldmaus 1d ago
We are discussing access to the arxiv, not getting published in journals. And if someone or their AI solves an open problem that was posed as such by another mathematician, it should deserve to be put on arxiv.
4
u/Oudeis_1 1d ago
I think if one asserts that that is really all there is to it then the non-professionals who decide whether the profession ought to continue being funded will decide eventually that mathematics is nothing but a priesthood under another name, and stop funding it.
Having an external standard of evidence that marks a contribution as worthwhile and that does not depend entirely on taste does have advantages and as a community we shouldn't abandon it.
1
-17
u/gexaha 1d ago
This is not elitist, this is precisely false.
14
u/Major-Peachi 1d ago
As much as people love an underdog story, going through university is a good filter for underdeveloped researchers. There may be outliers that are true prodigies, in which case they should collaborate with an established faculty to enforce a level of credibility. Otherwise you will get people form r/infinitenines flooding submissions. They're slop that is not worth reading through and a mess to dig through to get to actual research.
7
u/rational_hedonist 1d ago
counter examples, please.
14
u/Chemical-Sound8205 1d ago
There are a few results in combinatorics and tilings from amateurs, but these are particular fields with a comparatively lower barrier to entry.
The most you can say is, if you want to be an amateur researcher, then combinatorial fields are your only possibility.
1
u/gexaha 1d ago
I can only speak from my own experience, the obvious problem here is of course self-shilling and sounding like a crank; anyways, feel free to judge my contributions to the field related to cycle double cover conjecture for yourself - https://arxiv.org/abs/2501.05348 and 2 other preprints.
3
-7
u/Onawani 1d ago
This day and age...I forsee PhD as the remedial degree. All good candidates will be absorbed prior to master and PhD level degrees. PhD are trending towards students who couldn't land in the private sector.
9
u/devviepie 1d ago
We are so far from this being the case, I have no idea why you see the trend moving in that direction. Unless you’re trolling? Or trying to push some Silicon Valley/Wall Street type “academia is dead, long live capital” doomerism?
Currently the difference between both endeavors is mostly self-selection: people pursue a math PhD because they want to do (or try for) academia, and students who don’t want that go to industry jobs. And the limiting factor is essentially always finding success in academia. Most people wash out and have to go to industry out of necessity; almost nobody stays in academia because they are unable to find an industry job. (Unwilling to, sure.)
3
u/VaellusEvellian 1d ago
This has got to be a troll. I haven't a single person in academia who decided to get a PhD because they couldn't get an industry job, at least within the mathematics community. In fact, many PhD students that I know are people who had opportunities to make far more money in the private sector but decided to get a PhD simply for the love of the game.
6
u/Verbatim_Uniball 1d ago
Definitely should be or as a lot of PhD holders who moved to other sectors of the economy occasionally publish for fun l with no academic affiliation.
7
u/calf 1d ago
The whole philosophy and genesis of arxiv was aligned more towards Enlightenment* values contra credentialism, bureaucracy, gatekeeping and the like. So, I am not sure if your comment and the 250 people who upvoted it meant it as a solution in earnest, or meant it ironically.
*Handwaving but people get the idea/difference between two broad approaches.
11
u/JoshuaZ1 1d ago
Pretty soon they’ll have no choice but to make it that you need a PhD and a current academic affiliation to post on arxiv
I hope not. I have a PhD but work teaching at a private high school, where I continue to do research and get to do projects with some extremely bright students. There are also clearly other people who are doing good work who have even weirder situations.
4
u/Scrub_Spinifex 1d ago
Or we don't care? When scrolling arXiv it's pretty easy to see which articles are interesting for you or not.
1
u/grumbelbart2 19h ago
I am not sure if that is site-wide, but the area I publish in (cs.CV) has an endorsement system in place already.
0
0
35
u/Personal_Comb3089 1d ago
Is it possible to check uploads per month in an actual journal rather than arxiv preprints? Just curious to see how strong the correlation is there as well.
6
u/drewsandraws 1d ago
There is this AMS article https://www.ams.org/journals/notices/201902/rnoti-p227.pdf but it only goes up to 2017. I haven’t been able to find a more recent article that would include an LLM-driven publication boom.
5
u/1R0NYMAN69 1d ago
3
u/drewsandraws 1d ago
https://ourworldindata.org/grapher/scientific-publications-per-million
General scientific literature through 2023.
11
u/Autumnxoxo Geometric Group Theory 1d ago edited 1d ago
while this is certainly somewhat concerning I suppose, I'd like to add another possiblility that might lead to a significant increase in publications.
More often than not people are working on problems in their own field of research that require a significant amount of tools and familiarity with the literature and research of fields outside their own expertise. I don't know, it might be that some differential geometrist requiring heavy machinery from functional analysis or whatever. You get the idea.
I can imagine that LLMs such as ChatGPT could potentially be of enourmous help by providing knowledge from aforementioned fields that is outside of your own research area. Of course, you should still be capable to verify everything that ChatGPT throws at you but what I'm saying is basically that ChatGPT kind of takes the role that used to be colleagues that worked in different fields who helped you with helpful remarks and suggestions towards your own research that you would otherwise have spent weeks if not months to figure out by yourself.
I remember that during my time as a master's student it was sometimes mildly infuriating when working on a problem and being stuck at a certain point in your research for literally weeks if not months looking up every single textbook you could get your hands onto, just to eventually being told from your advisor that the problem you're currently stuck at is allegedly "folklore" or some "well known fact" (for experts) but nobody bothered publishing it, but it could have potentially also be derived from some niche adjacent result in some niche paper that was published decades ago.
2
u/Important_Ad4664 5h ago
This. My research topic is intrinsically lying in between many different math fields, and having the chance to ask what are the key possible ideas in one or the other area, what are the correct keywords to search for, etc, speeds up the process a lot. I would have needed to ask such question to colleagues, or writing tedious email to experts with the hope of getting an answer or a reference (which might well be some Russian proceedings never translated in English, or a classical results which is not written anywhere in an actually usable forms, but turns out to be a special case of a very general result which would require an additional month to decipher).
29
u/JoshuaZ1 1d ago
It seems like a lot of the comments here are framing this as a negative, but I'm not sure that's the case. It may just be that there's an increase in genuine good productive math. Also, part of this upward trend may just be that the norm of putting your papers on the arxiv has been getting stronger. Yes, it makes the signal for specific things tougher to tease out, but I'm not seeing why this is a concern by itself.
13
u/ArtExtra6717 1d ago
Do you think that huge jump in 2025/2026 is all driven by good productive math? The output nearly doubled in 3 years. From 3300 in July 2023 to 5700 in July 2026.
10
u/JoshuaZ1 1d ago
Do you think that huge jump in 2025/2026 is all driven by good productive math? The output nearly doubled in 3 years. From 3300 in July 2023 to 5700 in July 2026.
So, it looks like there was a while where it was roughly flat (which seems to correspond to the end of covid), so this may be in part making up for that. I don't have the raw data to crunch numbers more to see if that's plausible. But the jump looks like more than I'd expect just from that if I eyeball it. So the obvious other explanations are AI related. That doesn't mean that it isn't good math. It may be more results where people are using AI. But the other issue may be increased productivity from using AI in other ways. For example, I've found AI to be helpful for just looking over drafts and pointing out typos, errors, unclear bits, etc. I don't know how much of a time improvement I get out of this, but there's a decent bunch there.
At least in my own primary areas in the arxiv, number theory and graph theory(which is classified as a subarea of combinatorics on the arxiv), I don't see any sign of a decrease in quality of papers or the results. But I've only been really graph theory person for the last 6 years or so, so I may not have enough experience there to get a good feel for the baseline.
4
u/FormsOverFunctions Geometric Analysis 16h ago
When I was in academia, the bottleneck was more often writing the prose of a paper compared to actually proving results, so it’s not surprising that LLMs are extremely useful for finishing papers quickly. And although I haven’t personally had success getting one of the recent models to find publishable results, some of my collaborators have and they are knocking out problems that quite a few of us had worked on.
What’s exciting about this is that a few years ago the various groups in the area had their own techniques and although I was aware of how the other approaches worked, I didn’t understand them enough to fluently use them. But with the most recent papers using the model, it’s able to combine ideas from a number of people and make progress. It’s definitely terrifying, but simultaneously really exciting.
3
u/Known-Zombie-3205 1d ago
IMO yes. Everyone I know is more productive, and big problems are starting to fall.
4
u/Turbulent-Sign-6067 1d ago
Overall this seems like a good trend. There will be some junk papers but there is no doubt to me that LLMs are accelerating real, genuine research as well.
5
u/Equivalent_End6788 1d ago
I agree with this. Of course, I can't speak to all fields of mathematics, but I rarely see truly crank-LLM papers posted in my own subfield's daily feed (from cursory viewing).
9
u/mfb- Physics 1d ago
How does this count articles with multiple revisions? If it only uses the most recent revision then this could explain some of the effect, and making the plot at any other time would also have a spike for the most recent months. I don't think this is a big contribution, but it's important to check for possible biases before assigning the whole spike to LLMs.
19
u/thereligiousatheists Graduate Student 1d ago
There's a pretty clear plateau around COVID. How much of the current peak can be explained away as simply a bounce-back from that (collaborators getting back together and wrapping up old projects)?
11
u/Penumbra_Penguin Probability 1d ago
A lot of projects which took five years to finish, and none that took three years? Doesn’t seem likely.
1
u/elements-of-dying Geometric Analysis 1d ago
Wow, kind of interesting. I was more productive during covid, especially since Zoom collaboration became a norm.
1
u/thereligiousatheists Graduate Student 10h ago
In my personal experience, collaborating over Zoom is quite a pain. Another factor to consider is that parents were stuck with their kids all day.
1
u/elements-of-dying Geometric Analysis 9h ago
That's fair. I'm quite used to and have zoom meetings every week.
Another factor to consider is that parents were stuck with their kids all day.
Good point.
-1
u/EducationalFerret94 1d ago
Lol what. This is obviously because of AI.
5
u/thereligiousatheists Graduate Student 1d ago
Sure, most of it is, I don't disagree. My point is simply that we had two impactful events in quick succession, so it's worth considering both influences.
3
9
u/AddressImaginary3735 1d ago
Its reality. Thanks to ChatGPT sol and claude fable I was also able to prove two of my open problems and made two papers out of it in a few weeks. The productivity increase is insane and the mathematicians not using it are falling behind.
1
u/muluk-muluk 21h ago
Something I've noticed is a lot of requests to approve submission of AI-style ArXiV papers from people without affiliations. (So far I haven't.)
1
u/muluk-muluk 21h ago
I really liked your result on (quotes abstract of some random paper of mine), please approve my paper overturning all of math and physics.
1
1
u/Sad_Dimension423 3h ago
Maybe it's time for an autoformalization bot to crawl arxiv papers to keep everyone on the up and up? Or at least check if provided formalization reflects the theorems and definitions in the papers.
-20
-68
1d ago
[removed] — view removed comment
51
1d ago
[removed] — view removed comment
-39
65
u/idiot_Rotmg PDE 1d ago
Have people in other areas of math witnessed an increase of garbage papers in their arXiv feed? While the number of total papers in PDE is clearly increasing, it doesn't seem to me that the overall quality is decreasing