9 ms·
65 out of the 100 most cited papers are paywalled
- StavrosK 9y agoFund SciHub! What I'd like to see is an IPFS feature that showed the "least shared" files in a set, so you could say "I want to help host the rarest 10 GB", for example.
- toomuchtodo 9y agoYou can do this, as all of SciHub is available as torrents. Edit: you should be able to find the URLs with some googling, not going to post them here
- m-p-3 9y agoNow I'm curious, how much data would that be with the latest dump?
- snowpanda 9y agoLast time I checked it was between 40TB - 60TB for just the articles, not including the libgen books. At that time I calculated that it would cost about $1000 in refurbished 4TB external hard drives. Which is probably not the best way to store it. Or about 2000 Empty Blu-Rays if you want to make searching for an article a nightmare.
- deleted 9y ago[deleted]
- dredmorbius 9y agoSearch indices != primary storage. Ask Google about that.
- snowpanda 9y agoDo you mind expanding on that? What are you suggesting I search for on Google? I genuinely am trying to understand.
- Skuzzzy 9y agoForgive my brevity, typos, and poor writing as I'm writing this on my phone. He is stating that you would never actually search on your primary storage mediums. You would build your search index off the primary storage medium such that when looking up an article you do not search over the primary medium and only search the index and then once the article is found you read from said primary medium.
- dredmorbius 9y agoThis, yes.
- Redoubts 9y agoWell sure, but then you have to crawl though thousands of pieces of physical medium to load the file you want...
- sli 9y agoDo search indexes typically not include an identifier for a given resource? Seems to me that an index would be almost useless without them, as the match itself is more often than not far less useful than the document that contains it. Unless the information you need happens to be in the surrounding blurb. So yeah, I'm not sure why you'd have to "crawl" through physical media when a search index ought to tell you right where your match is located. Is that not the entire point of indexing and searching?
- dredmorbius 9y agoAn index maps a query to a specific item. Say, a file reference within a storage hierarchy or record within a database. (The concepts are fundamentally identical.)
- userbinator 9y ago...or about 10 LTO-7 tapes (6TB each --- "15TB" compressed, but I assume the article collection is already compressed), at slightly less cost for the media, and lot more for the drive; but with better reliability and even worse access times...
- m-p-3 9y agoOh damn, this is way bigger than I imagined.
- StavrosK 9y agoThe problem with that is that you can't serve content directly over HTTP from a torrent, whereas with IPFS you can.
- shpx 9y agohttps://webtorrent.io https://webtorrent.io
- deleted 9y ago[deleted]
- StavrosK 9y agohttps://www.eternum.io/ipfs/QmNmrNkyLygYDt9t5ptXkukPdpGSabLByG9tYKZ5kyTTkq https://www.eternum.io/ipfs/QmNmrNkyLygYDt9t5ptXkukPdpGSabLB... Which UX is better?
- toomuchtodo 9y agoMore importantly, webtorrent is a client side hack when clients should be using webseeding [1], which is baked into torrent clients that use libtorrent. [1] https://en.wikipedia.org/wiki/BitTorrent#Web_seeding https://en.wikipedia.org/wiki/BitTorrent#Web_seeding
- deleted 9y ago[deleted]
- kanzure 9y agoDo you have a complete copy of scihub? The seeders seem to be very slow. Please let me know -- email is in my HN profile -- I'll even pay for complete copies.
- amelius 9y agoIf you get a hold of it, please seed! :)
- ric2b 9y agoIPFS could be such an amazing thing for the academic world, imagine every new paper including an IPFS hash alongside the citations, so you could instantly jump to them while reading the pdf. It would guarantee that you'd be seeing the same version that was cited and the link would never break as long as at least one person kept sharing it. It would achieve the original goal of the world wide web as envisioned by Tim Berners-Lee, except even better.
- StavrosK 9y agoYes! I had exactly the same idea. I also love how anyone can reshare the file transparently, whereas with HTTP, if the original server doesn't have it any more, you have to manually hunt around for alternate links.
- aalleavitch 9y agoThis is by far one of the most frustrating things to me about the current structure of the scientific community. Making all published research free to access for everyone would be a massive benefit to the general education of society and would allow anyone regardless of institutional affiliation to be involved in the process of science. Imagine how much better science reporting would be if every popular science article was expected to link directly to the full papers they were referencing. As someone who hasn't been involved with a university or a laboratory for many years, I find myself continually extremely frustrated by how difficult it can be for me to keep up with new developments in the fields I studied in college.
- jmnicholson 9y agoAgree! It's really ridiculous and quite sad. I am no longer at a University so basically everything I read is stolen via Scihub. Hopefully, more and more researcher will continue to preprint their work. Something we're trying to help with at https://www.authorea.com https://www.authorea.com (disclosure: I work there)
- mirimir 9y agoIt's a great site, which documents a horrible situation. It'd be cool if it also showed which papers have preprints online. And which (I presume virtually all) are available via SciHub. Even links, if possible :)
- knlje 9y agoAnd a funny thing is that, even though I work in a university, I still find myself using Scihub sometimes. Either we do not have the access (old papers are a problem) or I'm not present at work.
- SiVal 9y agoI agree, and I'll explicitly add "history" as a part of "science" in the research sense you meant. I'm having a heck of a time trying to access historical records of medieval England from here in Silicon Valley because, although the information I need (land deeds, court records, etc.) was uncovered & translated by Victorian historians, and much of it exists in both text and image form and is ALREADY ONLINE in Google Books, the various English universities & their presses, English historical associations, libraries, etc., all jealously guard the information instead of releasing it, even to the point of preventing Google from showing it. Example: Cambridge Univ Press takes books that were Victorian databases of medieval records, and in the name of "protecting our precious history", they just photocopy the Victorian pages, reprint them on newer paper, and put them under new copyright. If you want to look something up in the 10-volume database, then either blindly buy the $300 set of books and hope to find some items of interest scattered therein or go to the nearest library that has the full set, which turns out to be on the other side of the planet at Leeds Uni in Yorkshire, which will allow you to see these ancient texts printed in 2013 if you pay for library use and make an appointment at least 2 days in advance.... Meanwhile, the full text of the books is on Google Books at Google's expense and available in any browser, but Google is forced by the "protectors of history" at Cambridge to cloak many of the pages of 1000-yr-old data, because history is too precious to allow the unworthy to see their photocopied Victorian texts without first going on an old-fashioned quest, bribing the boatman, answering the troll's three questions, etc. My hope is that at some point there will be a cultural change among "historical preservation" organizations where they decide that the greatest thing they can do to promote their field is to find every original document in every collection, carefully photograph it in hi-rez using whatever optical frequencies bring out the most faded detail, and contribute it to a free online host. Next step is then to create and freely post transcriptions, translations, and indexes, so that ANYONE can use the data for research, not just those who have enough gold in their purse, time for the quest, and can "answer me these questions three".
- CapacitorSet 9y agoRelevant: [SciHub](https://scihub.org/ https://scihub.org/) is a project to "provide free access to research articles and latest research information without any barrier". It can also be used via Telegram at @scihubot.
- diggan 9y agoHuh, didn't recognize the URL for Sci-Hub but then realized this is not THE sci-hub but another one who stole it's name. Correct URL would be https://sci-hub.cc/ https://sci-hub.cc/
- CapacitorSet 9y agoMy apologies, I googled "scihub" and picked the first result in English.
- lunchladydoris 9y agoThose numbers seem a little disingenuous. If you work at a decent university you're not paying $20 to access every article. Plus, I'd be stunned if all the people making the citations actually read the full paper. Some papers are cited because everyone knows you need to cite them.
- Feniks 9y agoDaily reminder that universities are funded by society. As are a lot of those papers... So yes you ARE paying. Just not directly.
- lunchladydoris 9y agoI completely agree. I wasn't commenting on the thrust of the article, which I agree with, but rather the distorting figures that were used to promote it.
- SiVal 9y agoAnd if you're part of the 99% of us who AREN'T working at a university but still funding their research with taxes, we have to pay both the taxes AND the direct fees to access the papers.
- jefft255 9y agoYou rarely need to pay if you work at a university. When you're on the school network or VPN there's a way to bypass the paywall because the uni usually buys subscriptions for most journals. That does not work if you are not working in a university.
- deerpig 9y agoThe vast majority of universities outside of developed countries can't afford to pay. I work at a national university in Cambodia. I've been trying to get the ministry of education to pay for a heavily discounted account (because we're a developing country) from JSTOR for the last two years with no luck and that's just a few hundred a year. Without libgen and sci-hub we'd be really screwed.
- philipkglass 9y agoThe top 10 paywalled articles are all from the 20th century. The Open Access movement is great but it doesn't do anything to free up papers from the past. A large part of the problem is the ridiculous duration of copyright. "Adsorption of Gases in Multimolecular Layers" is from 1938 and still paywalled. In practice, almost all papers this popular will be available on random .edu sites and Google Scholar will find those technically-forbidden copies for you. But it is a significant problem if you don't have an institutional affiliation and you want to read articles that aren't among the top 5% cited. (Or at least it was a problem for me before sci-hub; I retained academic contacts who could email me any papers I wanted, but I had to cross a pretty high interest threshold before I'd bug someone to request that favor.)
- Houshalter 9y agoCopyright desperately needs reform. It's silly we treat completely different areas with the same set of rules. From code to movies to math papers. Fine, let the Mouse be protected indefinitely. But nonfiction works have objective value to society. It's insane that 100 year old scientific works are still copyrighted and paywalled. It's wrong that you can be sued and even go to prison for spreading and preserving humanities knowledge.
- aalleavitch 9y agoThe internet makes a lot of the concepts behind copyright fundamentally ridiculous. We honestly need to rewrite many of the rules from the ground up to take into account modern technology and what would best benefit society given the accessibility of information.
- Feniks 9y agoLibgen.
- seccess 9y agoSo, I tried to see if I could read some of the articles marked "paywall" and I had no trouble. My methodology: Google Scholar search the article title, and click the direct "PDF" link on the right side. Eg: https://scholar.google.com/scholar?q=Tissue+sulfhydryl+groups https://scholar.google.com/scholar?q=Tissue+sulfhydryl+group... EDIT: My point here is that the statement in the article "the world’s most important research is inaccessible from the majority of the world" isn't exactly true. This isn't supposed to be an endorsement of academic publishing practices: if anything the fact that these publishers are effectively trying to scam readers out of money is all the more evident.
- Sargos 9y ago> My point here is that the statement in the article "the world’s most important research is inaccessible from the majority of the world" isn't exactly true. This statement is still true even with your trick. The vast majority of people don't know about this and also even if they did know about this it would be technically difficult for many of them who aren't tech savvy. This is a huge barrier that shouldn't be discounted.
- seccess 9y agoI don't see what tech savvyness has to do with it, many people use Google search. In fact, Scholar isn't even necessary, doing a regular Google search has the direct PDF links at the top of the search results.
- aalleavitch 9y agoThe fact that we have to low-key pirate research papers is just silly, though. The academic publishing system is just goofy, and I hope one of the projects that are currently trying to establish something better ends up taking hold soon.
- seccess 9y agoIt isn't piracy, all the publishers (in my experience & field of study) allow for free third-party hosting. Still agree with your second sentence though.
- jwilk 9y agoArchived copy, which can be read with JS disabled: https://archive.is/hlFg1 https://archive.is/hlFg1
- zitterbewegung 9y agoThe best thing I learned in university was to figure out how to get paywalled articles for free. This involved looking through arxiv and looking for the authors website .
- coldcode 9y agoIf you can't read it it doesn't exist. Research results are meant to be available and visible to all, or they are someone's private science diary. Also I believe that Nobel prizes should not be given out to be people whose research is not available to the general public.
- beedogs 9y agoWe should all be supporting Sci-Hub. There is no reason for these papers to be locked away from the public.
- Simulacra 9y agoWithout SciHub and my access at MIT, most of the research I depend on would be out of bounds.
- stevespang 9y agoand that's why I love Sci Hub, it has most all the titles I ever searched for: http://sci-hub.cc/ http://sci-hub.cc/
- dredmorbius 9y agoAnswering the question "Why is Sci-Hub so popular?": Because it works. It delivers information and knowledge to those who need it. Because information and knowledge are public goods. As CUNY/GC says, an "increasingly unpopular idea",1,2,3 but an absolutely correct one. Because it democratises information. Because much the world cannot afford to pay US/EU/JP/AU prices for content. Including many of those in the US/EU/JP/AU. And most certainly virtually all outside. Billions and billions of people. Because the research is (often) publicly funded, conducted in public institutions, and meant for the public. Because information and markets simply don't work. https://redd.it/2vm2da https://redd.it/2vm2da Deadweight losses from restricted access and perverse incentives for publication both taint the system. Because much the content, EVERYTHING published before 1962, would have been public domain under the copyright law in force at the time, and much up through 1976 and the retrospective extensions of copyright it, and multiple subsequent copyright acts, have created. Because 30% profit margins are excessive by any measure. Greed, in this case, is not good. Because the interfaces to existing systems, a patchwork fragment of poorly administered, poorly designed, limited-access, and all partial systems are frankly far more tedious to navigate than Sci-Hub: Submit DOI or URL, get paper. Because unaffiliated independent research is a thing. Because the old regime is absolutely unsustainable. It will die. It is dying as we write this. Because the roles of financing research and publication need not parallel the activity of accessing content. Ronald Coase's "Theory of the Firm" (1937, ), a paper which should be public domain today under the law in which it was created and published, and should have been by 1991 at the latest, but isn't, tells us why: transactions themselves have costs. http://sci-hub.ac/http://papers.ssrn.com/sol3/papers.cfm?abstract%95id=2308556 http://sci-hub.ac/http://papers.ssrn.com/sol3/papers.cfm?abs... Because journals no longer serve a primary role as publishers of academic material, but as gatekeepers over academic professional advancement. This perpetrates multiple pathologies: papers don't advance knowledge, academics are blackmailed into the system, and access to knowledge is curtailed Because what the academic publishing industry calls "theft" the world calls "research". Notes See GC Presents, "At the Graduate Center, we believe knowledge is a public good. This idea inspires our research, teaching, and public events. We invite you to join us for timely discussions, diverse cultural perspectives, and thought-provoking ideas." https://www.gc.cuny.edu/Public-Programming/GC-Presents https://www.gc.cuny.edu/Public-Programming/GC-Presents See GC President Chase F. Robinson, introducing a conversation between Paul Krugman and Olivier Blanchard. A rare moment where the introduction itself contains some provocative thoughts. At about 50s into the video. (The remaining 72 minutes and 20 seconds aren't bad either if you're interested in discussions of global economics.) https://www.youtube.com/watch?v=zndOEQnMC44 https://www.youtube.com/watch?v=zndOEQnMC44 Joseph Stiglitz, "Knowledge as a Global Public Good," in Global Public Goods: International Cooperation in the 21st Century, Inge Kaul, Isabelle Grunberg, Marc A. Stern (eds.), United Nations Development Programme, New York: Oxford University Press, 1999, pp. 308-325. http://s1.downloadmienphi.net/file/downloadfile6/151/1384343.pdf http://s1.downloadmienphi.net/file/downloadfile6/151/1384343... https://www.reddit.com/r/dredmorbius/comments/4p2rwk/what_the_academic_publishing_industry_calls_theft/ https://www.reddit.com/r/dredmorbius/comments/4p2rwk/what_th... (This has proved to be among my more popular articles, including being picked up by the Open Access community.)
- chasedehan 9y agoAs a former professor I don't see much of an issue with this. Every research institution on the planet will have access to the articles. Even if you don't have an affiliation (or your school doesn't subscribe to a particular journal), if you use Google Scholar to search for an article you can easily find pre-prints which are essentially the same thing. Additionally, if that still fails then the next option is to just email the author - they actually want you to read their work and will just send it out. The real issue is that the societies are essentially extorting universities for hundreds of thousands of dollars per year when the writers have to pay to submit and readers have to pay to read. Many of the newer journals are becoming open access, but few of them have been able to make enough in roads to be considered "good journals." This is a completely separate topic than the one from the above article.
- forapurpose 9y ago> Even if you don't have an affiliation A very tiny number of people have affiliations; this isn't a realistic option. > if you use Google Scholar to search for an article you can easily find pre-prints which are essentially the same thing That hasn't been my experience. Lots of things I can't read.
- folli 9y agoI personally do see an issue if we make knowledge only accessible to a minority of people. In my field (microbiology/bioinformatics), it's almost impossible to find pre-prints of pay walled papers and emailing the authors and hoping that they will respond at all also doesn't seem an extremely efficient process for literature research.
- chasedehan 9y agoMaybe I am biased by my field - economics - in which even having "access" doesn't mean they are "accessible." There are very very few people without a PhD who are able to understand most of the good papers. There is an even smaller number who would be able to contribute to scientific discourse, which is the main reason these papers are made available. The other thing that may only be localised to my field is that every author recognizes this problem and makes a copy available on their website or university's working paper site. I would only rarely resort to going to an actual journal because Google Scholar was way easier to find a copy of the article. This is, of course, going to differ for a number of fields and there is a growing trend for authors and journals to open up their articles.
- itdpydypdpfyf 9y agonot with sci hub ^^
- laichzeit0 9y agoYeah and how many are still inaccessible with scihub or libgen? I have access through my university to most journals but I always use scihub because it’s the easiest and fastest way to access any paper.