As an academic and regular submitter to arXiv, this is an eminently sensible policy. I like that it's per-submitter and not per-author, that means that the big labs with large collaborations shouldn't be terribly put out (more authors = more submission budget).
I hope something like this could also be adopted for some of our larger conferences - the absolute limits on co-authorship are what seem to cause the most grumbles.
> In September of 2016, arXiv received 9,869 submissions. In September of 2024, arXiv received 20,569 submissions. This September, arXiv received 40,363 submissions, which in turn generated almost 9,000 support tickets for arXiv staff and moderators.
> arXiv now limits submitters to up to two submissions per calendar month, with a limit of three total active submissions at any given time.
Arxiv papers are not supposed to signify anything. The only reason that people may derive some kind of career benefits is because Google Scholar indexes it and counts it towards the h-index etc. All that metric-based career-optimization going on is quite broken anyway, so this is fixing things from the wrong end as well. (The better end: stop hiring and promoting people based on Arxiv paper counts, so there is no incentive to flood it)
The hosting cost side I can understand, though they received quite some funding recently, but I assume it's not going towards actually running the site, but to who knows what broader impacts and so on.
Moderation for Arxiv is a silly idea. It already shouldn't be taken as a quality signal that something is able to be up on Arxiv. Obviously they should remove illegal stuff, but beyond that, requiring moderation is a misunderstanding of their role and reason for popularity.
If hosting is too expensive, I guess an alternative aggregator could also arise. With just metadata and a hash of the pdf that can be hosted anywhere, and as the sumbitter, if you move the file, you can change the URL.
The main reason for Arxiv's existence is the timestamping and the easy referencing. (Though I admit that the stable hosting is also a pretty important part, but they mention moderation effort as the reason, not the hosting costs.)
Academia is losing sight of the forest for the trees, can't see more than an arm's length ahead of their noses.
> Moderation for Arxiv is a silly idea. It already shouldn't be taken as a quality signal that something is able to be up on Arxiv.
It sounds like you're advocating for arXiv to become viXra. On https://vixra.org/all/2610 the second paper is currently "Correct Interpretation of the Great Discoveries in Particle Physics: I. Reconsidering the Higgs Boson via Vedic Vortex Structure", followed by a paper arguing that classical electrodynamics is bogus, "Gauss’s Flux Theorem Does not Hold in Time-Varying Electric Fields", and then "On the Boundary Problem in the Origins of Matter, Life, and Consciousness".
I don't think arXiv will continue to receive "quite some funding" if this is what it contains.
Agree that it basically shouldn’t be necessary. But unfortunately Arxiv isn’t able to change the bad incentive structure that the broader hiring ecosystem has set up, right?
If academics and their committees actually deeply considered the scientific qualities of applicants, gaming the numbers would not matter. So I just find it hard to blame AI slop instead of the broken process that academia has settled on.
Part of the reason that the value of Arxiv papers has gone up, is that it's pretty well known that the conference review system is already quite broken and random and lazy and superficial. So being rejected from there is not a good reason to ignore a paper, so conference-rejected Arxiv-only but good papers are a regular occurrence. If conference review had better signal, arxiv could have zero signal.
The other issue is the time delay. Conferences have stopped being an actual place to learn about new work, the way it had been 10+ years ago. Now it's all outdated stuff, and as a specialist you already know the important papers months before the actual conference, so the conference is a networking event basically. The real exchange of ideas moved to Arxiv and Github and social media, with its own problems, hype, algorithmic engagement optimization incentives etc.
But the pressure that shifted away from conferences to arxiv is continuing. The old guard was pearl clutching already by the erosion of conference rubber stamps as this big prestige. The newer ones are now trying to hold back the tide at the Arxiv line. But it will keep on moving faster and faster. And what is going to matter in the end is not how things used to be, but what delivers actual value. If the slop is slop, it will dwindle. If it starts to be actually good, all this will just turn into basically a dock-unions-opposing-automation story.
Nice writeup and the policy is a step in the right direction. I suspect simple rate-limit measures like this will be a big improvement.
I suspect a logical conclusion Arxiv and elsewhere may be an identity management system with an aggressive filter, and shared blacklists. I suspect that classifying people as spam/slop-submitters, then banning them (or whatever identity they used; name, email, name + organization etc), applying incremental rate limits over the general one, may be required.
If the queue size grows beyond volunteer capacity then volunteer time becomes a scarce resource. In my experience (in other organizations), more likely than out and out block listing is aggressive prioritization. It takes sensible design to determine which papers get to consume moderators' time first.
Then, if someone posts rejected papers regularly and is deprioritized, it naturally follows that they may try to submit three more... and get stuck naturally. There is no need to ban them. They will still get a response, and they will get a fair shake - when there is time.
It doesn't seem humanly possible to produce >2 "high SNR" papers per month. If you have a larger batch of related papers, it's easy to spread submissions over multiple months. You should probably be doing that anyway, because comments on one paper may affect the others.
> It doesn't seem humanly possible to produce >2 "high SNR" papers per month.
The situation right now is pretty crazy, for example this established TCS professor [1] has at least 4 arXiv submissions co-authored by him in September [2]. This includes the recent breakthrough on the matroid secretary problem, which has a pretty interesting AI story of its own, if you haven't seen it yet [3].
Of course prolific professors can work around this by working with younger researchers who upload the work, but this just makes the rate limit a solo author bottleneck, which feels a bit weird.
[3]: Concurrent Discovery Disclosure: The proof of the main result in this manuscript was obtained in a conversation with ChatGPT-6 Astra on Tuesday, September 15, 2026 at 1:02 AM PDT. We then prepared this manuscript for public release, with the intent of uploading it on the morning of Thursday, September 17, 2026. In the early morning hours of September 17, while finalizing the submission, we discovered the manuscript of Abdi, Banihashem, Hajiaghayi, and Mittal, uploaded on September 16, 2026, which contains the same result via an essentially identical approach. We are sharing our manuscript nonetheless in case our exposition is of independent utility to the community, and we hope this experience stimulates broader discussion about concurrent discovery in the AI era. (From https://arxiv.org/abs/2609.20797 .)
The situation is pretty crazy, but it just means the bar for publishable results will go up. Theoretical computer science is a field built around conferences. The number of papers that can be published in reputable conferences ultimately depends on the size of the community and the willingness and ability of its members to attend those conferences.
I don't see what we lose if lead authors are forced to submit work themselves instead of going through the PI's account.
But yes, it'd be nice if arxiv also rate-limited co-author submissions with some higher number to discourage the most flagrant PI co-authorship abuses.
Insane. No. Mathematicians will not produce >2 "high SNR" papers per month as their rate calculated over any reasonable time scale; but their maximum per month? Very much the case that they will complete linked papers/dependencies/series of work/etc, spend a month in editing across multiple nearly-complete submissions, and so on.
And in mathematics, at least, priority on results is determined by the first time it appears out in the wild; for most, this is the arXiv.
Nobody is going to accept waiting a month on something that they fear may be scooped, costing them years of work and thought.
This affectively makes the arXiv obsolete for much of the purpose it has heretofore been put to.
Do you have a concrete example of a high SNR researcher who will be limited by this policy? Aka one who is submitting more than two high quality papers per month?
I hope something like this could also be adopted for some of our larger conferences - the absolute limits on co-authorship are what seem to cause the most grumbles.
> In September of 2016, arXiv received 9,869 submissions. In September of 2024, arXiv received 20,569 submissions. This September, arXiv received 40,363 submissions, which in turn generated almost 9,000 support tickets for arXiv staff and moderators.
> arXiv now limits submitters to up to two submissions per calendar month, with a limit of three total active submissions at any given time.
The hosting cost side I can understand, though they received quite some funding recently, but I assume it's not going towards actually running the site, but to who knows what broader impacts and so on.
Moderation for Arxiv is a silly idea. It already shouldn't be taken as a quality signal that something is able to be up on Arxiv. Obviously they should remove illegal stuff, but beyond that, requiring moderation is a misunderstanding of their role and reason for popularity.
If hosting is too expensive, I guess an alternative aggregator could also arise. With just metadata and a hash of the pdf that can be hosted anywhere, and as the sumbitter, if you move the file, you can change the URL.
The main reason for Arxiv's existence is the timestamping and the easy referencing. (Though I admit that the stable hosting is also a pretty important part, but they mention moderation effort as the reason, not the hosting costs.)
Academia is losing sight of the forest for the trees, can't see more than an arm's length ahead of their noses.
It sounds like you're advocating for arXiv to become viXra. On https://vixra.org/all/2610 the second paper is currently "Correct Interpretation of the Great Discoveries in Particle Physics: I. Reconsidering the Higgs Boson via Vedic Vortex Structure", followed by a paper arguing that classical electrodynamics is bogus, "Gauss’s Flux Theorem Does not Hold in Time-Varying Electric Fields", and then "On the Boundary Problem in the Origins of Matter, Life, and Consciousness".
I don't think arXiv will continue to receive "quite some funding" if this is what it contains.
Part of the reason that the value of Arxiv papers has gone up, is that it's pretty well known that the conference review system is already quite broken and random and lazy and superficial. So being rejected from there is not a good reason to ignore a paper, so conference-rejected Arxiv-only but good papers are a regular occurrence. If conference review had better signal, arxiv could have zero signal.
The other issue is the time delay. Conferences have stopped being an actual place to learn about new work, the way it had been 10+ years ago. Now it's all outdated stuff, and as a specialist you already know the important papers months before the actual conference, so the conference is a networking event basically. The real exchange of ideas moved to Arxiv and Github and social media, with its own problems, hype, algorithmic engagement optimization incentives etc.
But the pressure that shifted away from conferences to arxiv is continuing. The old guard was pearl clutching already by the erosion of conference rubber stamps as this big prestige. The newer ones are now trying to hold back the tide at the Arxiv line. But it will keep on moving faster and faster. And what is going to matter in the end is not how things used to be, but what delivers actual value. If the slop is slop, it will dwindle. If it starts to be actually good, all this will just turn into basically a dock-unions-opposing-automation story.
I suspect a logical conclusion Arxiv and elsewhere may be an identity management system with an aggressive filter, and shared blacklists. I suspect that classifying people as spam/slop-submitters, then banning them (or whatever identity they used; name, email, name + organization etc), applying incremental rate limits over the general one, may be required.
Then, if someone posts rejected papers regularly and is deprioritized, it naturally follows that they may try to submit three more... and get stuck naturally. There is no need to ban them. They will still get a response, and they will get a fair shake - when there is time.
The situation right now is pretty crazy, for example this established TCS professor [1] has at least 4 arXiv submissions co-authored by him in September [2]. This includes the recent breakthrough on the matroid secretary problem, which has a pretty interesting AI story of its own, if you haven't seen it yet [3].
Of course prolific professors can work around this by working with younger researchers who upload the work, but this just makes the rate limit a solo author bottleneck, which feels a bit weird.
[1]: https://en.wikipedia.org/wiki/Mohammad_Hajiaghayi
[2]: https://arxiv.org/search/cs?query=Hajiaghayi&searchtype=auth...
[3]: Concurrent Discovery Disclosure: The proof of the main result in this manuscript was obtained in a conversation with ChatGPT-6 Astra on Tuesday, September 15, 2026 at 1:02 AM PDT. We then prepared this manuscript for public release, with the intent of uploading it on the morning of Thursday, September 17, 2026. In the early morning hours of September 17, while finalizing the submission, we discovered the manuscript of Abdi, Banihashem, Hajiaghayi, and Mittal, uploaded on September 16, 2026, which contains the same result via an essentially identical approach. We are sharing our manuscript nonetheless in case our exposition is of independent utility to the community, and we hope this experience stimulates broader discussion about concurrent discovery in the AI era. (From https://arxiv.org/abs/2609.20797 .)
But yes, it'd be nice if arxiv also rate-limited co-author submissions with some higher number to discourage the most flagrant PI co-authorship abuses.
And in mathematics, at least, priority on results is determined by the first time it appears out in the wild; for most, this is the arXiv.
Nobody is going to accept waiting a month on something that they fear may be scooped, costing them years of work and thought.
This affectively makes the arXiv obsolete for much of the purpose it has heretofore been put to.