3 ms·
Does somebody know why we are still using pdfs for papers ? I know a lot of people that are trying to parse PDF files and it is an awful process. If somebody i
by tevlon 11y ago
Does somebody know why we are still using pdfs for papers ?
I know a lot of people that are trying to parse PDF files and it is an awful process.
If somebody is looking for an idea for a new venture, this is a problem, yet to be solved !
- stuxnet79 11y agohttps://www.readcube.com/ https://www.readcube.com/ http://www.jove.com/ http://www.jove.com/
- Blahah 11y agoreadcube is a spectacular somersault in the wrong direction. It is so much worse than a PDF - it actively fights me when I try to extract information.
- jessriedel 11y agoIsn't that the idea? My impression is that it's a format introduced to help publishers like Nature introduce frictions to copying their content.
- stuxnet79 11y agoWasn't very impressed by it either, but clearly there are a few startups in this area.
- jessriedel 11y agoWe are using PDFs because they have universal adoption and, importantly, they reliably produce the same document everywhere. Most alternatives you might think of will give variable results on different machines. There definitely needs to be something that makes it easier to parse and otherwise interact with a PDF. But, for network-effect reasons, it's probably easier to introduce a parseable overlay for PDFs than to replace the format wholesale.
- tevlon 11y agoi disagree. While i agree, that PDFs are there and they will stay for a long time. The problem can be easily solved by journals. They make you use of their own LaTeX templates. It would be very easy to just force the submitters to add the LaTeX together with the pdf. Sometimes the easy solution is too obvious, i guess
- jessriedel 11y agoWell, basically everything on the arXiv has the .tex available, but that doesn't seem to make the problem much better. The problem isn't getting raw access to the text. It's possible to copy-past from PDF's with labor too, or to examine the inside of it (which is its own typesetting system, like .tex). The problem is that this data is very difficult to parse.
- jandrese 11y agoThere is a tug of war taking place here. TeX is nice because the documents are formatted for whatever your reading situation happens to be. PDF is formatted for exactly one situation, the A4 sheet you targeted. But on the flipside, the TeX document will often be ugly no matter how you are reading it, and the author can't easily apply tweaks for aesthetics or readability--many try, hence the TeX markup horrorshow.
- jessriedel 11y agoHonestly, most of these issues are just a product of .tex's long and storied history. If some foundation plunked down $1-$10 million, it could definitely produce an open source successor to .tex (with extensive, maintained, and documented libraries like Mathematica) that avoids most of the badness.