5 ms·
Hi all, I’m one of the engineers at AI2 that helped make this happen. We’re excited about this for several reasons, which I’ll explain below. Most academic pa
by codeviking 5y ago
Hi all,
I’m one of the engineers at AI2 that helped make this happen. We’re excited about this for several reasons, which I’ll explain below.
Most academic papers are currently inaccessible. This means, for instance, that researchers who are vision impaired can’t access that research. Not only is this unfair, but it probably prevents breakthroughs from happening by limiting opportunities for collaboration.
We think this is partly due to the fact that the PDF format isn’t easy to work with, and thereby make accessible. HTML, on the other hand, has benefited from years of open contributions. There’s a lot of accessibility affordances, and they’re well documented and easy to add. In fact, our hope long-term is to use ML to make papers more accessible without (much) effort on the author’s part.
We’re also excited about distributing papers in their HTML form as we think it’ll allow us to greatly improve the UX of reading papers. We think papers should be easy to read regardless of the device you’re on, and want to provide interactive, ML provided enhancements to the reading experience like those provided via the Semantic Reader.
We’re eager to hear what you think, and happy to answer questions.
- isaacimagine 5y agoLooks great! Have you considered linking this up to something like arxiv or other preprint sites?
- codeviking 5y agoYup, we're definitely thinking about this. Our focus right now is on providing a tool folks can run it on whatever papers they have access to. For instance, some researchers might have access to documents that aren't available to the public. We want them to be able to run this against those. That said as we expand the effort I imagine we'll eventually pre-convert things that are publicly available, like those on ArXiv, etc.
- _delirium 5y agoThere's already this for arXiv: https://www.arxiv-vanity.com/ https://www.arxiv-vanity.com/ Their job is a little bit easier because arXiv papers have the .tex source available, so you can use one of the various tex2html variants, instead of having to extract the paper's contents from a rendered PDF.
- politelemon 5y agoI've never actually questioned the why, so maybe you could shine some light... why are they usually published as PDFs?
- codeviking 5y agoY'know, that's a good question. I'm not sure I know the answer. My guess is it's largely for historical reasons. At the time most venues were organized PDF was probably the best (or only) mechanism for sharing documents for print distribution. But we think it's time to change that :).
- kartoshechka 5y agoUnfortunately for my mental health my thesis was exactly about converting arxiv papers to modern looking html, and there's so much more broken, unjust and ugly things in academia then using pdfs... Regarding your question, I'd say that it is a natural continuation of centuries long tradition of writing on the actual paper. The invention of TeX actually made it easier to produce more papers, then came PDF, and you could produce virtual papers. Also science journals pretty much have monopoly on scientific knowledge distribution, and they are mostly paper too
- DoreenMichele 5y agoI have no idea at all but as a wild guess, I would assume it's because you can't edit PDFs. So you know it says the same thing forever and no one went and changed it in response to reading criticism of their paper or something.
- temp8964 5y agoWhat alternative do you have? Word file? PDF is the only widely supported format can guarantee accurate reprint.
- miohtama 5y agoAre papers printed anymore? HTML for text. SVGs for diagrams. Equations can be exported as images if needed.
- kahon65 5y agoDo you remove the pdf files we send to your servers? Edit https://allenai.org/terms https://allenai.org/terms point 5, you own all the uploads! So if by mistake we send a medical PDF for example or something else that is under gdpr, we can't ask you to delete it???? ? Wtfffff
- codeviking 5y agoWe don't retain the uploaded document. We cache the extracted content, as to make things more efficient. See https://papertohtml.org/about https://papertohtml.org/about: > What data do we keep? We cache a copy of the extracted content as well as the extracted images. This allows us to serve the results more quickly when a user uploads the same file again. We do not retain the uploaded files themselves. Cached content is never served to a user who has not provided the exact same document. Also, we can delete the extracted data on request. Just send a note to accessibility@semanticscholar.org. Sorry for the confusion!
- kahon65 5y agoAh okay, thank you. >Also, we can delete the extracted data on request. Just to be 100% clear, you are referring to the cached extracted data, right?
- codeviking 5y agoYup, that's right.
- kahon65 5y agoThank you very much!
- Telemakhos 5y agoIs there any thought about presenting the papers as TEI XML with XSLT to display the paper in a browser or screenreader? TEI provides pagination support (needed for citing page numbers, because most of academia still needs that) and extensive semantic markup for things like bibliographic information. It also serves as one data model that can be converted easily with existing tools (XSLT) to provide many representations for humans, while also serving as a machine-parsable text for datamining. Digital humanities has made heavy use of TEI for years, and this project seems like it could benefit from it.
- znpy 5y agoI'd love to see a way to re-export a paper into a digital-friendly format, say epub/mobi to use on my e-reader. Any plans on that?
- kwhitefoot 5y agoYou could give Calibre a try. The result will probably be a long way from perfect for complicated documents but it does work reasonably well for most things. Formulas don't translate well unfortunately.
- 1vuio0pswjnm7 5y ago"We're eager to hear what you think, ..." I think I will stick with pdftohtml, pdftotext, and pdfimages https://en.wikipedia.org/wiki/Poppler_(software) https://en.wikipedia.org/wiki/Poppler_(software). These take seconds not minutes. From user perspective I dont understand why not release the source code and let people compile a native application. (Did I miss the link to the source code.) Instead it looks like this is just a means of collecting free data (metadata, more training data, data from submitted papers by default) everytime someone submits a paper.