5 ms·
I might be missing something, but I still don't get the why. I don't see any "problem" that needs to be solved.
by psalminen 7mo ago
I might be missing something, but I still don't get the why. I don't see any "problem" that needs to be solved.
- kolinko 7mo agoThe article lists the reasons quite clearly.
- binsquare 7mo agoFor everyone else, The reason is because arxiv is growing significantly leading to 297,000 deficit in operating costs for 2025 alone. Corenell has helped with donation a long with other organizations that pay membership fees. As a result, donors + leaders of arxiv think it's best to spin off to increase funding.
- vl 7mo agoWhat is unclear why they need stuff of 27 and 6.7 million to operate essentially static hosting website in 2026.
- swiftcoder 7mo agoThe "essentially static hosting" isn't the cost centre (although with 5 million MAU, it's nothing to sneeze at). The real costs are on the input side - they have an ingestion pipeline that ensures standardised paper formatting and so on, plus at least some degree of human review.
- bonoboTP 7mo agoDo you mean that the CPU compute cost of turning latex into pdf/HTML is the main cost?
- swiftcoder 7mo agoNo, I mean that the pipeline requires software engineers to build/maintain, and salaries are (as in basically every tech organisation) the dominant cost
- bonoboTP 7mo agoThen drop it and make people upload a pdf and a zip of the latex sources. Most people I talk to hate that pipeline and spend a lot of debug hours on it when Arxiv can't compile what overleaf and your local latex install can.
- domoritz 7mo agoArxiv can recompile latex to support accessibility and html. Going to pdf submissions would be a major step backward.
- bonoboTP 7mo agoMake it an external service then, and leave the thing that's already working great to just be. The reason authors like and use arxiv is that it gives 1) a timestamp, 2) a standardized citable ID, and 3) stable hosting of the pdf. And readers like the no-nonsense single click download of the pdf and a barebones consistent website look. All else is a side show.
- OneDeuxTriSeiGo 7mo agoYou have to keep in mind that an increasing portion of their time and labor is going towards moderation and filtering due to a mass influx of nonsensical AI generated papers, non-academic numerology-tier hackery, and other useless drivel. Spinning the service off forces other the labor out onto other universities rather than leaving them to solely Cornell
- bonoboTP 7mo ago
- lou1306 7mo agoThe PDF formatting is all but standardised. They ingest LaTeX sources, which is formatted according to the authors' whims (most likely, according to whatever journal or conference they just submitted the manuscript to). I'll concede that the (relatively novel) HTML formatter gives paper a more uniform appearance. They also integrate a bunch of external services for e.g., citation metrics and cross-references. Still hard to justify such a high cost to operate, but eh. Also, the "human review" is a simple moderation process [1]. It usually does not dig into the submission's scientific merits. [1] https://info.arxiv.org/help/moderation/index.html https://info.arxiv.org/help/moderation/index.html
- OtherShrezzing 7mo agoI don't see it as an especially exuberant structure or budget. I've seen larger teams with bigger budgets struggle to maintain smaller applications. I've contracted into some consultancy teams which you could uncharitably describe as "15 people and $4mn/yr to create one PDF per month".
- planetoftofu 7mo agohttps://info.arxiv.org/about/reports/2024_arXiv_annual_report.pdf https://info.arxiv.org/about/reports/2024_arXiv_annual_repor... A critical component of the arXiv-CE project is moving our services entirely off of Cornell University’s infrastructure — this goal is also known as Milestone 1. Milestone 1 completion is projected for the end of fiscal year 2026. Assume if you are a library, and every day, half baked so-called books brought to the librarians where they have to make sure it is meaningful, readable and printable, 3000 of them, they accept and put them in the right bookshelf, and entire internet reads every one of them on the shelf multiple times by the AI bots, search engines and researchers. They are not only making a new library, they are also maintaining both and syncing two libraries because Cornell cannot handle the volume of access by bots. It is not static. It is essentially running two ships side-by-side, and two ships need to appear as one from the outside. And, the new ship is still only half built. The new ship is being designed, and being built. 27 seems small to me.
- sanex 7mo agoNow they're going to have a deficit of 600,000 in operating costs.
- pessimizer 7mo ago> The reason is because arxiv is growing significantly leading to 297,000 deficit in operating costs for 2025 alone. Dollars? So 300 people's cable bill? That's basically nothing. They're spending too much, and it's still nothing, and the solution is going to be to privatize it and eventually loot it. You can't hand out a collection plate and get $300K for Arxiv? Your local neighborhood church can. Civilization is obviously collapsing.
- u1hcw9nx 7mo agoI think the problem described in 6th paragraph needs to be solved.