12 ms·
A plan to mine the world’s research papers
- nafizh 7y agoIt would be interesting to see the server getting hacked with all the papers in the open, and then the researchers sending a sorry email to the publishers.
- killjoywashere 7y agoAnd use the same legal theory as Equifax, Target, etc? It would be beautiful.
- aaron695 7y agoSci-Hubs database weights 25kgs. This has long been solved https://mobile.twitter.com/sci_hub/status/697787778881490944 https://mobile.twitter.com/sci_hub/status/697787778881490944 Search is what needs doing now. The next step in furthering knowledge.
- TadaScientist 7y agoamen
- mirimir 7y agoRight, "hacked". More likely, a public-interest leak. Because, as Facebook (for example) is finding, leaks happen.
- modeless 7y agoSo, sci-hub, but less accessible?
- tuxguy 7y agoa, minable interface to sci-hub :)
- hyperbovine 7y agoAlso, marginally less likely to suddenly vanish in a poof of lawsuit one day. See Napster.
- jacquesm 7y agoBut storage costs have dropped by orders of magnitude since then and 'personal copies' of SciHub are a definite possibility.
- IronBacon 7y agoI've read somewhere a couple of years ago that SciHub was about 35TB of data, quite big but still manageable.
- snaky 7y ago$800, considering $200 for one 10TB HDD.
- jacquesm 7y agoIt's a bit more now, going on 80T or so. Still quite manageable. A large NAS will do.
- m1el 7y agoI just happen to have a 80T NAS that I plan to use as a sci-hub mirror, is there a way to download it? libgen torrents are dead.
- jacquesm 7y agoemail?
- IronBacon 7y agoSo if the 35TB figure was correct a few years ago, I didn't expect now to be more than doubled. Do you happen to know if it's only SciHub or it also includes Libgen? If it's only SciHub that's a lot of papers... ^__^;
- deleted 7y ago[deleted]
- dlkf 7y ago> No one will be allowed to read or download work from the repository, because that would breach publishers’ copyright. Instead, Malamud envisages, researchers could crawl over its text and data with computer software, scanning through the world’s scientific literature to pull out insights without actually reading the text. I find this totally unconvincing. The average scientific article isn't any good, and the NLP algorithms that do tasks like this are even worse. For scientific literature to be useful to those of us outside the academy, we need to be able to see the full document - what methodology was employed, what assumptions were made, how the data was gathered - just to be able to gauge whether the authors had any idea what they were doing. Ideally we would also be able to search over documents, explore the citation tree in some sort of UI, and access articles in multiple formats (sometimes you want a PDF, sometimes you want plaintext, and I imagine that the mathy-types might like LaTeX source). I applaud Malamud's efforts to overcome this problem, and I hope his results prove me wrong. I just think it's sad that we have to resort to hacks like this to overcome what is obviously enormous scam that is stealing our tax dollars and stifling academic and economic creativity.
- mirimir 7y agoYeah, this is silly. Elsevier and the like should be burned down for crimes against humanity. I mean, it generally is illegal to price gouge during famines.
- natechols 7y agoElsevier is a useless profiteer, but let's please remember that scientists are voluntarily submitting their articles to Elsevier journals, and the funding agencies (and universities) are doing nothing to stop them. I find it a little frustrating to read the wailing from institutions like UC (where I used to work) about subscription prices when they are part of the problem in the first place.
- deleted 7y ago[deleted]
- mirimir 7y ago
- aphextim 7y agoR.I.P. Aaron Swartz https://en.wikipedia.org/wiki/Aaron_Swartz https://en.wikipedia.org/wiki/Aaron_Swartz
- apo 7y ago> Over the past year, Malamud has — without asking publishers — teamed up with Indian researchers to build a gigantic store of text and images extracted from 73 million journal articles dating from 1847 up to the present day. Maybe I missed it, but the article doesn't seem to explain exactly how Malamud's group compiled its database. Throttling is a major problem with the naive approach of throwing wget on a publisher site. The publisher detects a bot on its network downloading everything in sight and either slows data transfer to a trickle or just shuts down access to it. The publishers may not win on copyright, but they may try to make a case based on criminality if Malamud's team actively took steps to circumvent throttling and defeat the defenses of the hosting sites. Especially if the publisher knows what to look for in its logs.
- toomuchtodo 7y agoScihub? Bonus points if your properly executed scraping project backfills Scihub where it is missing DOIs.
- Merrill 7y agoThis is good, more because it will make searching for information easier than because it will avoid copyright. Searching through journal articles to find information about a given topic is very hard, even at a university with more or less universal library access to the online literature. Authoring papers of uneven detail and quality and then publishing them in whatever prestigious journal that will have them is a terrible way to document the progress of science. Instead, new results should be added to an open science information base, fully linked to all previous results that bear upon the new results, either in support or contradiction. This is what some of the bioinformatic data bases attempt to achieve by scanning articles, but it would be better to omit the article step.
- kodz4 7y ago> A trigger for this mission came from a landmark Delhi High Court judgment in 2016. The case revolved around Rameshwari Photocopy Services, a shop on the campus of the University of Delhi. For years, the business had been preparing course packs for students by photocopying pages from expensive textbooks. With prices ranging between 500 and 19,000 rupees (US$7–277), these textbooks were out of reach for many students. In 2012, Oxford University Press, Cambridge University Press and Taylor and Francis filed a lawsuit against the university, demanding that it buy a license to reproduce a portion of each text. But the Delhi High Court dismissed the suit. In its judgment, the court cited section 52 of India’s 1957 Copyright Act, which allows the reproduction of copyrighted works for education. Another provision in the same section allows reproduction for research purposes. Good job India.
- jdjayded 7y agoI'm a little late to this party, but here's a mandatory research plug: My lab works in the area of evidence based medicine. My research focuses on Randomized Controlled Trials, and the overall goal is to automate (fully or partially) the meta-analysis of medical interventions. To that end, we collected a dataset of intervention pairs, statistical significance findings about these pairs, and a minimal rationale supporting the significance finding. Since these annotations were collected on full text PubMed Open Access articles, we can distribute both the articles and the annotations: https://github.com/jayded/evidence-inference https://github.com/jayded/evidence-inference ; paper: https://arxiv.org/abs/1904.01606 https://arxiv.org/abs/1904.01606 ; website: http://evidence-inference.ebm-nlp.com/ http://evidence-inference.ebm-nlp.com/ We're working on collecting more complete annotation. We hope to facilitate meta-analysis, evidence based medicine, and long document natural language inference. We might even succeed (somewhat) since this is a very targeted effort, as opposed to something more broad.