14 ms·
I really hope sci-hub survives this. Sci-hub and libgen are like an entirely different internet, one allowing you to dive as deep as you wish into any technical
by lcvw 5y ago
I really hope sci-hub survives this. Sci-hub and libgen are like an entirely different internet, one allowing you to dive as deep as you wish into any technical subject. There’s really no comparison I have found anywhere for the depth of material available. People always point to Wikipedia, but all of that is surface level. If you to build something, research something, or just really delve into it, there’s no substitute for having access to all of the latest textbooks, manuals, and papers. I’ve never found any source paid or otherwise that comes even close.
- londons_explore 5y agoI really wish someone could build a better UI for this research internet. Hyperlinks for all references would be a good start. Finding some way to make some automatic glossary of definitions of technical terms would make scientific papers substantially more accessible too.
- elcritch 5y agoDOI URIs work great and sci-hub understands them. Figuring out the DOI from the citations section is a bit more annoying still. A meta glossary would be fantastic. 80% of learning a new field is figuring out the jargon.
- goodmachine 5y ago> A meta glossary would be fantastic. 80% of learning a new field is figuring out the jargon. True!
- IshKebab 5y agoYeah it's called Web of Science. It's very good but it's not free unfortunately. And it doesn't go as far as hyperlinking references in PDFs or defining terms. I agree those would be great, but unfortunately there's not much incentive for paper authors to do that. https://en.wikipedia.org/wiki/Web_of_Science https://en.wikipedia.org/wiki/Web_of_Science
- Al-Khwarizmi 5y agoHyperlinking references in PDFs is trivial for authors if they use LaTeX and the journal/conference template supports it. It's just a matter of ensuring that the bibtex entry has an URL or a DOI, and most bibtex entries copied and pasted from curated sources already have them. If you are finding many papers without hyperlinked references, it's probably just because they're published in journals whose templates don't support it. In my particular research field, most publication venues' templates starting supporting those links around 3-4 years ago, so my papers from, say, 2015 have no hyperlinks in references, while those from 2019 do. This didn't require any significant extra effort on my part, in fact in general it requires less because well-curated bibtex entries are easier to come by now than some years ago.
- IshKebab 5y agoIt's not adding the links that is hard. It's choosing the destination URL. Where do you link to? I guess you could link to doi.org. Probably better than nothing but still not ideal because it doesn't actually take you to the PDF. Can you show an example paper with links?
- foobarbecue 5y agoDOI is the correct thing to link to because the author or publisher has chosen this as their canonical URL for the object. The DOI could and sometimes does point to the pdf, it's just conventional to point to an html version of the paper. It would make a lot of sense to me to standardize a field in the DOI metadata containing the PDF URL. (source: I manage the DataCite membership of a large organization)
- jltsiren 5y agoLinking directly to the PDF is usually the wrong choice. When you find a new paper, you often want to get the citation metadata, which the PDF document rarely contains in a convenient form. There are often multiple versions of the same paper, and you may want to determine which version you managed to find. Is it a preprint, the final authors' version, the published journal paper, an early version published in conference proceedings, or an unpublished extended version of the paper?
- geokon 5y agoOne thing I've recently discovered is this website: connectedpapers.com/ It builds a graph of referenced papers and makes it easier to narrow down which one are important/foundational for further research I don't think a glossary would help. What you need to find is a "review paper". These act as a primer to the field for new researchers. They're usually well written, with less jargon and have tons of references for you to dig into. That said, I don't have a good method for finding them.. I just stumble across them haphazardly..
- stevesimmons 5y agohttps://www.semanticscholar.org/ https://www.semanticscholar.org/ is very good too for reading paper abstracts, links to the references, citations and related papers. The full papers are included, if copyright allows. The references and citations are tagged and filterable, making it easy to see which are the most cited, or review papers, etc.
- beckman466 5y agoHave you heard of Alexandra Freeman's Octopus? https://www.science.org/careers/2018/11/meet-octopus-new-vision-scientific-publishing https://www.science.org/careers/2018/11/meet-octopus-new-vis... Edit for direct link to Octopus: https://science-octopus.org https://science-octopus.org
- sleepingsoul 5y agoHave you come across scite (https://scite.ai https://scite.ai) yet? We're also innovating in this space by extracting citation statements from full-text articles and classifying their intent. So let's say paper A cites paper B. If you look at paper B, we show you: - how many times it was cited - the direct paragraphs from paper A where it was cited - the sections from paper A where paper B was referenced - ... and a lot more You can also now search these citation statements directly to find evidence-based information pretty quickly. - Short video to showcase that search: https://www.youtube.com/watch?v=JYjCn-4uMJk https://www.youtube.com/watch?v=JYjCn-4uMJk - Website with a bit more details: https://citation.to/ https://citation.to/ You can also visualize citation networks similar to ConnectedPapers, set notifications for new citations on groups of papers you're interested in, and much more. (Disclaimer -- I work here!)
- KNrDajfZ 5y agoNo one cares. This thread is about free access to papers and not another paid service that forces you to pay monthly fees for something that could be a free service. In that sense you aren't any better than large online publishers. 8 bucks a month for a scientific paper search engine? Really?
- ComputerGuru 5y agoI vouched for your comment because you have a very valid point, but you could be more polite in making it. Welcome to HN.
- KNrDajfZ 5y agoThanks. I am really cranky today. I'll do better next time.
- deleted 5y ago[deleted]
- sleepingsoul 5y agoHiya, Well, I definitely agree with your sentiment in a normative sense that scientific papers should be free and readily accessible to all -- in part because a lot of it is funded through tax dollars! But given the current state of affairs, we're looking at making that information accessible to people without having to pay exorbitant fees to access individual research. We also offer steep discounts for students or anyone in academia. With that in mind I would push back a little that we're just a scientific paper search engine -- our system does a lot of work in extracting and classifying those citation statements, which makes it more powerful than traditional scientific search engines. And besides just using our search, a huge time-saving value of our service is the report pages which helps you quickly build a qualitative understanding of how something was cited. Even if all scientific papers were freely accessible, our report pages allow you to see the direct, relevant snippets from citing papers without having to manually read each and every single one. I think that is quite valuable! I know I've gone on a little tangent from the original discussion about scihub, and having free and open access to papers, but I did just want to throw that in because I think it's an important distinction. And as much as we all want that free and open world to exist, I think it's also interesting to think about how we can open up that information for people in the interim. Best, Ashish
- ZeroGravitas 5y agoSome of the background data for this is being collected by wikiCite (https://meta.wikimedia.org/wiki/WikiCite https://meta.wikimedia.org/wiki/WikiCite) and Scholia https://scholia.toolforge.org/ https://scholia.toolforge.org/ projects Part of this involves turning text data of authors into linked data that used to navigate between texts: https://author-disambiguator.toolforge.org/ https://author-disambiguator.toolforge.org/
- generalizations 5y agoSeems like a large part of what's needed is just being able to make the pdfs machine-readable, by making decent plain-text versions of the text content. Right now, IIRC, there's no hands-off way to get the text of a pdf. Especially if there's weirdness like multiple-columns (sometimes happens with this stuff).
- zapataband1 5y agoCurrently working in this field and this is actually the cutting-edge(!!) but it will be 100% possible/robust within the next year or so I believe. Really cool ML techniques being used for htis.
- sleepingsoul 5y agoHow well does it work with old OCR'd PDFs? :)
- generalizations 5y agoI had no idea. Is there anything that's open source, or is it all still being kept proprietary?
- zapataband1 5y agoLike you said the hard parts are the unstructured data/images/tables. There are pretty-good(80% of the way there) solutions tho. But nothing that could handle millions of paper without error
- jpeloquin 5y agoIt wouldn't take much to significantly improve things. We don't even have full-text search for paywalled articles. The paid search engines like Web of Science just do title, keywords, and abstract. Even considering the subset of open access articles, Google Scholar and Semantic Scholar do ok but don't offer much in the way of search refinement (e.g., "DTAF" NEAR "collagen"). They're good at finding _something_ related to your query but not good for systematic review.
- ehvatum 5y agoYes, it’s a new and better world. On the other side of the $275 per-paper Elsevier paywall, there are researchers who wish more people would read their papers. In my experience digging into robotics kinematics, authors are happy to answer questions and can point me to the right person when I want to send a check to support investigating specific research questions. The paywall deceives; science is neither an institution nor a copyright. It’s people. You can talk to them. You can learn from them and they can learn from you. When you apply their research, they often want to know about it! They might even discuss your application, in future papers. You don’t have to be Siemens or Big University Labs. The situation with for-profit journal publishers is diseased. Who the hell do they think they are? Elsevier should be dead, and Aaron Swartz should be alive.
- TaylorAlexander 5y agoI have been pleased that The Journal of Field Robotics is well represented on scihub. I have an open source off road robot I am designing and the journal is literally about robots out in fields and stuff. I am a "serious hobbyist" in that I believe my open source contributions to be at least somewhat helpful to others, but it's not the kind of thing that would justify paying for paywalled papers. I just want to glance over the material and keep track of what researchers are up to. Libgen is to me a vision of a world without copyright and intellectual property restrictions and I think it's a much better world than ours.
- tonyarkles 5y agoWow! The content in JFR is fantastic! Thank you for the pointer!
- TaylorAlexander 5y agoSo glad it was helpful!
- BeFlatXIII 5y agoMost authors of scientific papers will gladly send you a free PDF of their papers if you ask them (assuming they remember to check their e-mail and respond in the first place). The profitability of their publisher is of no concern to them.
- TheSpiceIsLife 5y agoWhat about a Netflix but for science information. Pay $15 a month for access to a rolling catalogue of science info.
- jules 5y agoWhy? The authors write the papers for free, the peer review is done by other scientists for free. Why should this "netflix for science" get to reap the profits by locking it behind a paywall? The reason why predatory publishers still exist is a coordination problem. The journals have prestige built up historically, and the scientists need to publish in prestigious journals for their career. It's a chicken and egg problem.
- qmmmur 5y agoAnd the papers are often funded by public money.
- GeckoEidechse 5y agoI'm glad initiatives like [Plan S](https://en.m.wikipedia.org/wiki/Plan_S https://en.m.wikipedia.org/wiki/Plan_S) exist ^^
- tchalla 5y agoHow much does it cost for an author to publish in a Nature or Science journal under Plan S?
- rsfern 5y agoI don’t know specifically about for Plan S, but most open access fees are in the 3-5k range per article
- mynameismon 5y agoDesktop version: https://en.wikipedia.org/wiki/Plan_S https://en.wikipedia.org/wiki/Plan_S
- KNrDajfZ 5y agoIt's the typical Eastern European (or non-US centric) Internet. The IP laws make sense only if people have the capital to buy stuff. When I was a kid in a post-Soviet country, no one ever bought anything original (like a CD with a game). Everything was bootleg or pirated from the Internet. After [the US lobby started to push for copyright enforcement in the EU under the threat of sanctions](https://falkvinge.net/2011/09/05/cable-reveals-extent-of-lapdoggery-from-swedish-govt-on-copyright-monopoly/ https://falkvinge.net/2011/09/05/cable-reveals-extent-of-lap...), things have changed. The copyright and IP laws that US lobbies push to the world are killing the idea of free Internet and only benefit large corporations that are untouchable. Elsevier is Dutch-based and they are fierce in suing everyone who tries to get away with getting free papers that were paid for by the taxpayers. The "free" Internet doesn't exist anymore, but it's good to have places like Russia where the IP law is not strictly enforced, because everything is broken so no one cares.
- Griffinsauce 5y ago> that were paid for by the taxpayers What do taxpayers have to do with it? I thought it was just a for-profit company?
- kanche 5y agoThe content i.e. research papers they publish are mostly from government funded research. Taxpayers fund those but don't get access: we have to buy the results from these for-profit corporations.
- zapataband1 5y agoThat's the best part about a good idea once it's out there! It's hard to kill. Really wish we had come up with an alternative to 20 streaming sites...
- codewithcheese 5y agoThere is. Torrents. The UX can be amazing if you know how, but we dont want spoil the party by sharing.
- rapnie 5y agoPeerTube at https://joinpeertube.org/ https://joinpeertube.org/ or Owncast at https://owncast.online/ https://owncast.online/ Both support live streams, and are (being) federated and ever more integrated with other Fediverse apps.