8 ms·
Torrent seeding effort: https://www.reddit.com/r/DataHoarder/comments/nc27fv/rescue_mission_for_scihub_and_open_science_we_are/ https://www.reddit.com/r/DataHoa
by sixtyfourbits 5y ago
Torrent seeding effort: https://www.reddit.com/r/DataHoarder/comments/nc27fv/rescue_mission_for_scihub_and_open_science_we_are/ https://www.reddit.com/r/DataHoarder/comments/nc27fv/rescue_...
All papers on sci-hub are available as torrents from library genesis. The full collection contains 85 million articles (before this announcement), and is about 80TB. If anything ever happens to sci-hub or library genesis, there's enough people out there with backups that a replacement can be set up fairly quickly, albeit without the proxy functionality to obtain new papers.
However, the more the merrier, so if you've got some spare hardware and bandwidth to share, I'd encourage you to contribute to the seeding effort if you're able. At current market prices of ~$30/TB, it costs ~$2400 to have a copy of the full collection sitting on your desk.
- animex 5y agoIs there any legal risk to users in North America that do this? Is this copyrighted material?
- sixtyfourbits 5y agoProbably; use a VPN. Yes they're copyrighted - albeit not by the authors who actually wrote them, but by the publishers who require copyright assignment for the privilege of having your work hosted on their website.
- blueblisters 5y agoSo when academics casually share papers with their collaborators, are they technically exposing themselves to lawsuits? Or is being part of a research organization/university with subscriptions to these services enough to mitigate that risk?
- goodcanadian 5y agoThat is an interesting question without a good answer. As I understand it (I am far from being an expert), there are sometimes specific exemptions to allow that sharing. However, generally speaking, I think it is a gray area. The journals would surely be shooting themselves in the foot, though, if they tried to sue their contributors. Academics would bring down hell on any journal that tried that. Moreover, it is not clear to me that they would win as it seems to be a customary practice even if it is not an explicitly allowed one. Finally, sharing a single paper might get research or academic exemptions as you aren't copying the whole journal issue. I doubt any publisher wants to go down that road. They are probably better off with the law remaining vague.
- CannoloBlahnik 5y agoSharing with collaborators may constitute fair use. I wouldn't put one of my own papers on my website, though, for example, for something that isn't published as open access.
- apdar 5y agoThey are usually shared under the title of “preprint” or “draft” often, but not always, before the time of actual publishing. Always seemed like a grey area to me. We didn’t really distribute the copy of the paper with the journal/conference’s name + copyright - though a perhaps a line under the title: “To be published in…”
- sixtyfourbits 5y agoIt depends on the circumstances and the publisher. In many cases publishers permit authors to host the accepted version of the paper (but not the final version that includes revisions based on reviewer feedback) on their personal/institutional website and to email copies of the paper on an individual basis to people who request them. For example, see https://www.elsevier.com/about/policies/copyright https://www.elsevier.com/about/policies/copyright (under "Author Rights") for what Elsevier permits you to do with your own work. On the other hand, publishers have sometimes filed lawsuits against sites where authors share their papers, e.g. ResearchGate: https://www.nature.com/articles/d41586-018-06945-6 https://www.nature.com/articles/d41586-018-06945-6 Elsevier have also sent takedown notices to universities where academics have made the final version available on their institutional websites: https://www.washingtonpost.com/news/the-switch/wp/2013/12/19/how-one-publisher-is-stopping-academics-from-sharing-their-research/ https://www.washingtonpost.com/news/the-switch/wp/2013/12/19... https://osc.universityofcalifornia.edu/2013/12/elsevier-takedown-notices/ https://osc.universityofcalifornia.edu/2013/12/elsevier-take... https://news.harvard.edu/gazette/story/newsplus/elsevier-takedown-notices-a-qa-with-peter-suber/ https://news.harvard.edu/gazette/story/newsplus/elsevier-tak... https://blogs.library.duke.edu/scholcomm/2014/01/28/setting-the-record-straight-about-elsevier/ https://blogs.library.duke.edu/scholcomm/2014/01/28/setting-...
- smarx007 5y agoYou are almost always allowed to self archive the final version, but you have to share the PDF you generated yourself, not the nicely formatted one from the publisher. And some publishers only allow self-archiving outside repositories like researchgate. But most importantly, Google Scholar and Semantic Scholar would pick up most links from blogs and Arxiv. Use https://v2.sherpa.ac.uk/romeo/ https://v2.sherpa.ac.uk/romeo/ to check. Also, EU projects in most cases now require open publishing and publishers make exceptions even when OA is forbidden (“self archiving is allowed if mandated by the funding agency”).
- lrem 5y agoWhat happens in practice: researchers are usually free to share the draft they had before the publishing process. Which means you can read the text before every sentence was wordsmithed in huge pain to make the paper a quarter page shorter.
- type0 5y ago> casually share papers with their collaborators, are they technically exposing themselves to lawsuits? For the final, not open access published papers they certainly do.
- michaelrpeskin 5y agoNo one talks about this, but for many (most?) of the papers they publish, there’s no copyright to assign. When I was in grad school, 100% of my research was federally funded (NSF, Navy, NASA) so they insisted that every product of that research be freely available with no restrictions. When we get an article ready for publication, the Journal would send us the standard “you assign your copyright to us” forms. We’d sign them but also send a copy of “the green form” stating that it was federally funded and we didn’t have any copyright to assign. They just ignore that.
- deleted 5y ago[deleted]
- ogma 5y agoLOL you're gonna get raided for sure. The feds are slaves to the RIAA and the rest of their corporate masters.
- baybal2 5y agoI encourage you to do this in the most audacious, and visible way possible. Let see how will they sue millions, upon millions of people. Make them face a fait accompli. They already lost.
- laurent92 5y agoThey never target millions. They target one, and ruin his life, preferably one with children so he also divorces. Every dystopian regime does that.
- deleted 5y ago[deleted]
- Griffinsauce 5y agoThey only need to create some examples.
- Griffinsauce 5y agoPlease look into the history of file sharing and how that went for individuals. These companies will pick examples and ruin their lives. It's hard to say how big the risk is here but this is reckless advice. Vote for copyright reform.
- zapataband1 5y agoFrom watching the Aaron Schwartz documentary, hopefully once scihub makes these publishers obsolete they will no longer have the power to press charges. But yeah they can and will get the Feds involved and will ruin your life like Schwartz
- tmp_anon_22 5y ago1. rent a seedbox for $10/mo hosted in a different country 2. seed from your seedbox, not your personal device 3. profit
- mishafb 5y agoIs it compressible or already compressed?
- sixtyfourbits 5y agoIt's all PDF files, which have their own compression, so it's unlikely there would be substantial gain from additional compression. Each torrent has 100 zip files, and each zip file has 1000 PDFs, but the files are stored uncompressed within the zips (i.e. using the STORE method).
- Thorentis 5y agoIs there some kind of searchable index included so that you can locate an article in a particular Zip? I'm assuming each article has some kind of ID numbers and the Zips are divided by ID range or something?
- sixtyfourbits 5y agoYep! See https://github.com/sci-hub-p2p/artifacts/releases/tag/0 https://github.com/sci-hub-p2p/artifacts/releases/tag/0 This project is in it's early stages and the documentation has quite some way to go, but the index that's part of the release contains all the necessary information. This tool also contains the code necessary to produce the index files if you have a local copy of the zips. Each torrent contains 100,000 files, comprised of 100 zip files with 1,000 PDFs each. They are named by DOI. There's a database dump at (http://libgen.rs/dbdumps/ http://libgen.rs/dbdumps/) (scimag.sql.gz) which has the id -> DOI mapping and other information. The specific torrent and zip file can be determined based on the id; torrent = id/100000 and zip = id/1000.
- andyxor 5y agoSci-Hub database/index is available here: http://libgen.rs/dbdumps/scimag.sql.gz http://libgen.rs/dbdumps/scimag.sql.gz and database documentation is available here: https://gitlab.com/lucidhack/knowl/-/wikis/References/Libgen-Articles-Tables https://gitlab.com/lucidhack/knowl/-/wikis/References/Libgen... also see introduction to Sci-Hub for developers: https://www.reddit.com/r/scihub/comments/nh5dbu/a_brief_introduction_to_scihub_for_developers/ https://www.reddit.com/r/scihub/comments/nh5dbu/a_brief_intr...
- deleted 5y ago[deleted]
- xoogler234 5y agoThis sounds very appealing. However, from a cursory search I can’t locate any NAS of that size anywhere close to that price point.
- yread 5y ago8-bay dock ~ 400$ (Icy Box 10-bay) 7x14TB HDD ~ 7*300$ (Toshiba MG07ACA14TE) total 2500$
- nicoburns 5y agoWhat we really need is an index of these torrents by DOI, and then ultimately by journal and issue. Are you aware of any work to make this happen?
- sixtyfourbits 5y agoThe only one I'm aware of currently is https://github.com/sci-hub-p2p/sci-hub-p2p https://github.com/sci-hub-p2p/sci-hub-p2p. Library genesis also hosts database dumps at https://libgen.rs/dbdumps/ https://libgen.rs/dbdumps/. There's really a need though for more developers to get involved with building tools for more easily searching and working with the collection, ideally with a nice UI and integration with things like crossref. This is a massively valuable data set and it would be great to see what people can come up with. Lots of awesome potential for data mining too. If it weren't for the legal issues (publishers using copyright law to restrict access to literature they got for free since they never pay authors for their work), there's no shortage of projects that could utilize this data and be enormously beneficial for the scientific community and humanity in general. Unfortunately such work can only be done in the shadows right now, which greatly limits the number of people/institutions likely to do so.
- Ericson2314 5y agoIsn't libgen already distributed via IPFS too?
- commoner 5y agoYes, LibGen is mirrored on IPFS: https://news.ycombinator.com/item?id=25209246 https://news.ycombinator.com/item?id=25209246
- moyix 5y agoI believe there is a database dump available with this information: http://libgen.rs/dbdumps/scimag.sql.gz http://libgen.rs/dbdumps/scimag.sql.gz