4 ms·
My gripe with PDF is that it's the standard format for academic publishing, rendering a whole mass of scientific knowledge largely inaccessible for text process
by mytddu 10y ago
My gripe with PDF is that it's the standard format for academic publishing, rendering a whole mass of scientific knowledge largely inaccessible for text processing purposes. I've wanted to analyze the Libgen archive of journal articles for a long time but have never found an adequate solution for extracting text from PDFs. Any suggestions on this?
- nitrogen 10y agoPdf2txt wasn't helpful? http://manpages.ubuntu.com/manpages/precise/man1/pdf2txt.1.html http://manpages.ubuntu.com/manpages/precise/man1/pdf2txt.1.h...
- IngoBlechschmid 10y agoSure, the Linux tool "pdftotext" works just fine for this. Two small caveats: ligatures get converted to proper Unicode ligatures and not their ASCII fallback (as one might want or expect) and of course complex mathematical formulas are rendered badly.
- mytddu 10y agoI've tried both pdftotext and pdf2txt and I remember not being satisfied with either. Neither seem to handle non-ASCII characters very well, but I'll take another look soon though.
- based2 10y agohttps://pdfbox.apache.org/ https://pdfbox.apache.org/