8 ms·
I'm an academic. For years, my academic niche has tried to break free from the likes of Springer/Elsevier. Here are the bottlenecks: * There are wonderful "pr
by abhv 6y ago
I'm an academic.
For years, my academic niche has tried to break free from the likes of Springer/Elsevier. Here are the bottlenecks:
* There are wonderful "pre-print" servers like arxiv and eprint.iacr.org. However, these do not maintain the "archival quality" document storage that is needed for academic scientific literature. In day-to-day, all researchers use these to stay informed on recent results.
But how to guarantee that nobody hacks in and figures out how to change a few bytes in one paper that is 10years old? How to guarantee that these documents are available 75 years from now? I'm sure that many of you can devise solutions to this, but they will be costly, and they will need constant labor to implement. How do you pay for this?
It is OK when 20,000 researchers in a field are downloading papers every once in a while, but what happens when every student in the world wants to read these? The bandwidth charge becomes non-trivial.
It seems like it needs to be outsourced, and some commercial entity with experience handles it.
* The tenure process is slow to change. Many academics need publications in prestigious journals with "high impact factors" in order to get tenure because the upper-level tenure committees in older institutions use these metrics to evaluate cases. These people are not stupid: it is just hard to evaluate cases across a university when you are not an expert. Instead, you assume that certain journals represent "the highest quality work" and thus use the presence of those publications to judge researchers. This means that the top papers still end up in Elsevier/Springer journals.
When I was a grad student at MIT, it was easy to read papers; if your IP was from MIT, every paper was 1 click away. I wonder how it is going to work now that Elsevier's catalog won't work this way...
- huhtenberg 6y ago> But how to guarantee that nobody hacks in and figures out how to change a few bytes in one paper that is 10years old? Trusted timestamping. https://www.ietf.org/rfc/rfc3161.txt https://www.ietf.org/rfc/rfc3161.txt It also happens to be very widely deployed and supported by well-established companies, because it's an integral part of executable cross-signing. That is, this exists right now.
- mcv 6y ago> Many academics need publications in prestigious journals with "high impact factors" It wouldn't surprise me if this is a massive part of the problem. Any new system to replace Elsevier may be perfect in lots of ways, but it doesn't count as prestigious, which means everybody will still want/need to publish with Elsevier. How do you magically grant a new publishing platform this 'prestige'?
- bonoboTP 6y agoWhen the prestigious expert editorial board resigns at the same time and creates a new journal. It has happened several times, e. g. Glossa for linguistics. Also JMLR in machine learning is independent and still well regarded.
- Vinnl 6y ago"Flipping" journals is an option, but doesn't happen often because it's risk for the editors with little personal benefit. The answer the project I volunteer for [1] is that the prestige of a journal comes from the researchers who submit or review for it, so we can also employ their reputations without the middleman - by having them endorse works, thus having their names instead of the journal names attached to the works. [1] https://plaudit.pub/ https://plaudit.pub/
- vegetablepotpie 6y agoWhen they mess up is when it changes. When you look at old institutions and powerful people, sometimes their rule ends abruptly because of scandal, bad decisions, or corruption. Bear Sterns, Enron, and Nixon are examples of this. For a new publishing platform to succeed, the old one needs die. For an organization built on prestige to die, it needs to be mired in scandal wrapped up and packaged in the political zeitgeist at that moment that not only affects its small community but also develops the ire of the entire society. At that point a new platform will emerge, likely backed by, and inheriting its prestige from, another institution. Edit: I realize, unfortunately, this post doesn’t give anything actionable that anyone can enact. It at least offers hope that things can change.
- underdeserver 6y agoI'm a random industry software engineer. :) > In day-to-day, all researchers use these to stay informed on recent results. But how to guarantee that nobody hacks in and figures out how to change a few bytes in one paper that is 10years old? Printed versions + digitally signed and timestamped PDFs. This is a solved problem in the world, at least up to the level that Springer can solve it. > How to guarantee that these documents are available 75 years from now? I trust MIT and Harvard to keep PDFs and printed versions available much more than I trust Elsevier or Springer to be around in 75 years. > Many academics need publications in prestigious journals with "high impact factors" in order to get tenure because the upper-level tenure committees in older institutions use these metrics to evaluate cases. These people are not stupid: it is just hard to evaluate cases across a university when you are not an expert. Instead, you assume that certain journals represent "the highest quality work" and thus use the presence of those publications to judge researchers. This means that the top papers still end up in Elsevier/Springer journals. I don't disagree. This is why the change and the first wave of papers will likely come from already-tenured professors, who still publish high impact papers. > When I was a grad student at MIT, it was easy to read papers; if your IP was from MIT, every paper was 1 click away. I wonder how it is going to work now that Elsevier's catalog won't work this way... Now imagine the same situation, except you don't need your IP to be from MIT.
- olau 6y agoJust about the bandwidth costs: You can rent a server at Hetzner.de for 40 EUR with 1 Gbit/s. Let's say each PDF is around 100 kb, then you can serve 1000 PDFs per second. Say there are 50 million active research students in the world, then the single 40 EUR server can serve them about 10 PDFs/week on average.
- bufferoverflow 6y agoWith that requirements you can do much cheaper even. 1000 PDFs of 100KB is just 100MB. You can put them on a 1Gbit VPS for just $1.5 to $5/month. https://www.serverhunter.com/?search=PLJ-9TW-TWN https://www.serverhunter.com/?search=PLJ-9TW-TWN A dedicated box would cost just $7/month.
- oefrha 6y agoDigital archival of PDFs weighing a few hundred KBs to a few MBs is definitely a solved problem. And there are already arXiv overlay journals out there, and platforms supporting them. Tim Gowers' (Fields medalist) blog posts on this topic are quite informative: https://gowers.wordpress.com/2015/09/10/discrete-analysis-an-arxiv-overlay-journal/ https://gowers.wordpress.com/2015/09/10/discrete-analysis-an... https://gowers.wordpress.com/2019/10/30/advances-in-combinatorics-fully-launched/ https://gowers.wordpress.com/2019/10/30/advances-in-combinat... Highlights: $10 per submission, plus some fixed costs, including archival with CLOCKSS. No Elsevier extortion ring needed. Impact factors are of course kind of a chicken and egg problem. Need to have enough high profile journals move off Big Publishing, or have enough high profile ones started. > When I was a grad student at MIT, it was easy to read papers; if your IP was from MIT, every paper was 1 click away. When I was a grad student at <institution of similar caliber>, or an undergrad at <another institution of similar caliber>, accessing papers was rather painful off campus. One either has to use EZproxy, which might decide to block you if it doesn't like your IP range (say in a foreign country), or use some godawful proprietary VPN client that I would stay the hell away from unless necessary.
- aduitsis 6y agoToday it's much easier, practically all universities participate as Identity Providers in SAML Federations and digital libraries participate as Service Providers. So you can just use your institutional login credentials to the identity provider page of your university. The service provider receives a signed SAML assertion that, well, asserts that you belong to your university and you are, say, a student. Most popular software is Internet2 Shibboleth (IdP and SP) in the academic field. It all works very well and has been for some time. In the country where I live, you get access to office365, (physical) books, digital libraries (including Elsevier :)) and a wide variety of other services all via your institutional login.
- nihil75 6y agoFrom my brief time working at Springer, seeing how their business model shifted towards services and processes aimed at enabling as many publications as possible - I think basing tenure decisions on the fact papers were published there is based on archaic notions and misguided.
- hedora 6y agoI’m also an academic. Hash each paper, then hash the hashes. Publish the result with the proceedings. After year one, include the hash of the prior year(s). Problem solved. Recently I downloaded one of my old peer-reviewed papers. The “archival” service added a spammy logo to the bottom left corner of each page. I’ve been meaning to find the original and put it on my web page. Honestly, I might just add a list of all my papers with links to SciHub instead. I’m allowed to post them on my personal web page according to every copyright agreement I recall signing.
- ComodoHacker 6y ago>Problem solved Not at all. There are corrections and amendments. It isn't as complicated as in law, but still.
- sudosysgen 6y agoSigned declarations of amendment, then amended papers being added as new one with proof of amendment and link with the original. Kind of like how a keyserver deals with revoked keys.
- a1369209993 6y agoCorrections and amendments are separate (but related) documents. Preventing them from being retroactively applied to the original version of the source document is the specific thing that a archival-quality document storage is supposed to do (as opposed a non-archival-quality storage, which only needs to protect against data loss (as in turn opposed to a cache, which can rely on a backing storage)).
- ComodoHacker 6y agoLooks like you agree that an archival-quality document storage assumes a bit more than hashing papers. That was my point.
- a1369209993 6y ago
- code4tee 6y agoThese are valid concerns but in 2020 very easy problems to solve. There’s little reason why a small consortium of institutions couldn’t build a very robust system to accomplish all that. Use digital signatures and distributed mirrored storage and problem solved. Charge a very modest fee to members to cover fixed costs and make it free to the public. Heck a few we’ll organized S3 buckets with a search engine attached would be better than a lot of what’s out there today. Not to pick on academics but the commercial publishing houses basically prey on the stubbornness of the academic community here. In the pure private sector someone would come along tomorrow and make Elsevier and others obsolete and they would go bust quickly. MIT is making the moves that might just whip something into shape to remove Elsevier’s role in the market. On “high impact” if the top universities in the world en-mass unsubscribe from the commercial players that will change quick.
- dheera 6y ago> I wonder how it is going to work now that Elsevier's catalog won't work this way MIT alum here. In my experience you can always request a copy directly from the author by e-mail. There is ResearchGate which aims to make this easier, but doesn't because the fundamental problem is academics don't have time to respond to every e-mail, much less every ResearchGate notification. So yes sometimes you have to ping them by e-mail about 2 or 3 times. I think ResearchGate -- or even Google Scholar -- should add a feature to allow manuscript requests to be auto-replied with a copy of the document instead of waiting for the author to manually send a copy.
- ComodoHacker 6y ago>But how to guarantee that nobody hacks in and figures out how to change a few bytes in one paper that is 10years old? And how Springer/Elsevier guarantee that? Is it written into cintracts?
- pas 6y agoEven if it is... who enforces it? Who checks it? What does that guarantee worth? Who cares really about papers getting numbers changed... 90% of them are absolute crap and only max. 10% of the remaining 10% replicates just based on the text of the paper alone.
- x86_64Ubuntu 6y agoI don't mean to be rude, but I think you are greatly exaggerating the technical and cost considerations behind this effort.
- wolco 6y agoThis is the first blockchain example that makes sense.
- gwern 6y ago> However, these do not maintain the "archival quality" document storage that is needed for academic scientific literature. In day-to-day, all researchers use these to stay informed on recent results. But how to guarantee that nobody hacks in and figures out how to change a few bytes in one paper that is 10years old? How to guarantee that these documents are available 75 years from now? I have never had a link to Arxiv or Biorxiv break, and I have never had difficulty finding a copy of a paper on them either, going back to Arxiv's founding in 1991. On the other hand, on a daily basis, I struggle to get a copy of a paper published often just years or decades old from these 'archival quality' publishers like Elsevier, and they break my links so frequently that I spend some time every day fixing broken links on y website (and for new links, I have simply stopped link them entirely & host any PDF I need so I don't have to deal with their bullshit in the future). I guess "archival-quality publisher" is used in much the same way as the phrase "academic-quality source code"...
- MrGunn 6y agoHey Gwern, big fan of your GPT2 work. I notice I'm surprised to hear you say you struggle daily to fix broken links to the Elsevier catalog at ScienceDirect, because the links are used by libraries all over the world & they don't have the same feedback. Would you have a few examples available for me to send to the folks responsible?
- gwern 6y agoNature does it all the time. Here's one I fixed just this morning when I noticed it by accident: http://www.nature.com/mp/journal/vaop/ncurrent/full/mp2015225a.html http://www.nature.com/mp/journal/vaop/ncurrent/full/mp201522... (Note, by the way, how very helpfully Nature redirects it to the homepage without an error. That's what the reader wants, right? To go to the homepage and for Nature to deliberately conceal the error from the website maintainer? This is definitely what every 'archival quality' journal should do, IMO, just to show off their top-notch quality and helpful ways and why we pay them so much taxpayer money.) Oh, SpringerLink broke a whole bunch which I am still fixing, here's two from yesterday: http://www.springerlink.com/content/5mmg0gmtg69g6978/ http://www.springerlink.com/content/5mmg0gmtg69g6978/ http://www.springerlink.com/content/p26143p057591031/ http://www.springerlink.com/content/p26143p057591031/ And here's an amusing ScienceDirect example: https://www.sciencedirect.com/science/article/pii/S0006320719321032 https://www.sciencedirect.com/science/article/pii/S000632071... (I would have loads more specifically ScienceDirect examples except I learned many years ago to never link ScienceDirect PDFs because the links expire or otherwise break.)
- 1MoreThing 6y ago> But how to guarantee that nobody hacks in and figures out how to change a few bytes in one paper that is 10years old? I can't believe I'm going to say this, but this sounds like an actual problem well-suited to a blockchain solution.
- tialaramex 6y ago> But how to guarantee that nobody hacks in and figures out how to change a few bytes in one paper that is 10years old? Mostly you should stop worrying about this. Other people have explained various countermeasures that could be used, which are very cheap, but mostly nobody cares. And already today, without anybody altering anything, it is very common for papers to use misleading citations. You take a paper that found some clowns like cake, you write "Almost all clowns like cake" and you cite that paper. It's possible a reviewer will notice and push back, but very likely you will get published even though you've stretched that citation beyond breaking point. Why "hack in" to change the paper when you can just distort what it said and get away with it?
- jopsen 6y agoI heard the archiving argument before. But I can't imagine that this is an expensive role for an organization like library of Congress or similar to take on. Many counties have a national library of sorts. Bandwidth/storage costs are limited, we're talking about PDFs.
- bduerst 6y ago>When I was a grad student at MIT, it was easy to read papers; if your IP was from MIT, every paper was 1 click away. Isn't this predatory pricing?
- NoSorryCannot 6y agoRe: hacking, preservation - - does Elsevier make reasonable assurances in this space beyond being a strawman that can be torn apart in a lawsuit? Is there a reason an open access platform could not make technological assurances that are as good or better? This really seems like a good and fairly low-risk opportunity for universities to form a consortium of sorts to make their own publishing platform. And then maybe "impact" would be built in because they'd all be getting high on their own supply. But these institutions are strangely prone to silos despite the collaborative spirit of academia writ large so I don't see that happening.
- bufferoverflow 6y ago> how to guarantee that nobody hacks in and figures out how to change a few bytes in one paper that is 10years old? You publish hashes of all the documents. It's trivial to distribute lists of hashes. You can even put them into existing blockchains, which guarantee they won't change.
- dumb1224 6y agoI think there are field-specific differences. In biology / bioinformatics world the "high impact factors" journals are the norm for tenure or even confirmation of work. But highly influential computing related papers are rarely from those journals. Bioinformatics is an interesting exception because it's bridging these fields and you'll see references from both highly reputable journals and biorxiv.