7 ms·
At the Internet Archive, we are working on this exact problem, and have been in communication with the pre-print's authors. We have built open infrastructure (o
by bnewbold 6y ago
At the Internet Archive, we are working on this exact problem, and have been in communication with the pre-print's authors. We have built open infrastructure (open source, open data) tracking "preservation coverage", for example:
https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3E1945+year%3A%3C%3D2019+%28type%3Aarticle-journal+OR+type%3Aarticle+OR+type%3Apaper-conference%29 https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3...
and are working to improve crawling. There is a "save paper now" feature, as well as an API for bots. Organizations like DOAJ, ISSN, DOI registrars (Crossref, Datacite, others) are crucial for this. In the broader ecosystem, we hope this can complement existing efforts that partner with large publishers (like LOCKSS, Portico, JSTOR) and institutional repositories. A natural niche for us is web-native (HTML) content, which we have crawled a lot of but are just getting started to index. For example, publications like d-lib, first monday, and distill.pub.
If folks want to help, it would be great to have a "youtube-dl for open access papers". There is a lot of content on large platforms and publishers which have anti-crawling measures (even for gold OA and hybrid content!), as well as a long tail of small publishers that don't use simple/common mechanisms like OAI-PMH and the `citation_pdf_url` HTML meta tag to identify fulltext content. The OAI-PMH ecosystem sadly is not very complete or helpful for the use case of mirroring.
- cxr 6y ago> If folks want to help, it would be great to have a "youtube-dl for open access papers". Zotero has an existing set of "translators". And somewhat related to this request: I scratched out some notes last year about how to get more out of "zero-obligation communities" (like the pool of prospective contributors in open source) <https://www.colbyrussell.com/2019/02/15/what-happened-in-january.html#underdeveloped> https://www.colbyrussell.com/2019/02/15/what-happened-in-jan.... The long and short of it is that instead of saying something like "if folks want to help, it would be great[...]", you should provide a place for people to sign up, take them at their word that they're willing to help, and then lay out a concrete set of tasks/deliverables. People get weird about trying to avoid being seen as not gentle enough with volunteers, but the end result is a lot of unharnessed human potential. You've got a pool of mechanical turks at your disposal. Take a break from polishing the arrangement of instructions you give to the computer and focus some energy on writing the "programs" that you want to be executed by meatbags.
- torgian 6y agoWhen does it become a prerequisite to eradicate meat bags and replace them with cold, calculating droi.... uh, code?
- garfieldnate 6y agoWintergatan's Martin recently transitioned a lot of the work for his new marble machine to volunteers, and he essentially followed this pattern, casting himself into the product owner role to harness the power of volunteers. It seems to be working pretty well!
- waheoo 6y agoHow does one get to work at the Internet Archive?
- cinquemb 6y ago> as well as a long tail of small publishers that don't use simple/common mechanisms like OAI-PMH and the `citation_pdf_url` HTML meta tag to identify fulltext content. The OAI-PMH ecosystem sadly is not very complete or helpful for the use case of mirroring. Most of orgs/journals/conferences in this group just don't have the resources to be able do this, nor maintain something like this. It was funny last year when a rep from google was doing a teleconference talk from Mountain View in Jakarta, covering things like this in front of a room full of like the top 200 journal managers in Indonesia (there are like tens of thousands of journals) and their eyes were glossing over when the rep was trying to address fixing some of the issues they encounter when crawling than hamper just indexing. Might as well be coming from a different planet…
- bnewbold 6y agoFrom what I have seen, the least technically resourced journals often use hosted platforms or free software like OJS (basically wordpress for journals), which comes with features like HTML meta tags and OAI-PMH by default. The trickier cases are when folks write their own platforms, or even write their own raw HTML with no templating, in which case adding tags to all landing pages or supporting an API would be a relatively large amount of work.
- cinquemb 6y ago> From what I have seen, the least technically resourced journals often use hosted platforms or free software like OJS (basically wordpress for journals), which comes with features like HTML meta tags and OAI-PMH by default. Yeah, and there are a lot of issues with OJS and the meta tagging and google being able to crawl a lot of these (not to mention site uptime where lots of sites go down for long stretches and google just assumes the site was taking offline permanently if kept down for a while). Esp when the meta tag locales don't map to the actual language used in the papers themselves (i.e. journal admins enter meta data tagged as EN but use indonesian for the text and the text of the paper) or like non iso locale codes in metadata fields, etc. Right now, when journals/orgs/conferences want help getting indexed in google and have oai endpoints, we pass all their content through locale detect stuff and don't trust their meta data by default.
- Vinnl 6y ago> If folks want to help, it would be great to have a "youtube-dl for open access papers". I think Unpaywall is already trying to do this? Or at least, for every DOI they index, they try to include a link to the direct article if known.
- bnewbold 6y agoUnpaywall is very helpful! However, even for direct PDF links, publishing platforms will often do things like check for a session cookie; if you don't have the correct cookie you get bounced back to the landing page, where you need find and follow another link. This isn't super complicated to work around (persist a cookie jar, use a headless browser, etc), but it doesn't work out-of-the-box with our crawlers, the same way crawling youtube doesn't work out-of-the-box so we rely on the youtube-dl community.
- Nemo_bis 6y agoYes, the Internet Archive already used those URLs, or tried to. Bryan was so kind as to share some statistics about it: https://groups.google.com/g/unpaywall/c/AbNwXdyWZfE/m/OVoVmxOTAQAJ https://groups.google.com/g/unpaywall/c/AbNwXdyWZfE/m/OVoVmx... Having an URL to something that is supposed to serve a PDF doesn't mean that you actually get it. Publishers like Elsevier or Wiley require a cookie/JavaScript dance and/or place heavy rate limits; whatever PDF hosted by them is never really accessible, even if it's ostensibly open access. The "youtube-dl for papers" would do what youtube-dl does, i.e. take the URL and execute whatever JavaScript or other trickery the target page/website requires before it serves the content.
- harlanji 6y agoI have a raspi1 that I tether to my phone with a script called download-video that wraps youtube-dl... I try to archive anything I consume before I consume it now. Maybe a little raspi image for researchers to behave like that could help. $100 archive setup, naturally it’s harder to use than most apps, I use Termius Premium for SSH + SFTP (share integration is nice) and run a screen session. IG same handle for photos and screenshots. The mirroring story is static gen + cqrs + cdn. I do this in my spare time, would like to do it in some sustainable capacity. New app studio for hire, basically. Patreon same handle for a little history. Email in profile, I don’t use other DMs.