38 ms·
Show HN: PDFLayoutTextStripper – Converts PDF to text while keeping the layout
- jlink 10y agowho would be interested by an online website doing the job?
- akouri 10y agoI bet most would, but privacy would be a big concern for me at least. script is optimal format for me
- rodw 10y agoFor what it's worth, here's a service that does that https://documentalchemy.com/demo/pdf2txt https://documentalchemy.com/demo/pdf2txt (and more: https://documentalchemy.com/demo https://documentalchemy.com/demo)
- jlink 10y agothanks for sharing this one, didn't know it.
- flexie 10y agoJust tried the demos on this website. I tried to extract text from a pdf that already has searchable text, which can be copy-pasted. This should be the easiest task of all but it made mistakes in every second word. Then I asked the website to make a pdf into a word-file. It just inserted the whole pdf as a picture in word.
- rodw 10y ago> Then I asked the website to make a pdf into a word-file. It just inserted the whole pdf as a picture in word. Really? I'm pretty sure that's not the way this works.
- logicallee 10y agoif you really want to rake it in, serve, at static speeds (meaning instantly, I swear, boot a ramdrive (Tmpfs) and serve static html from nginx all from RAM), text versions of the top 10,000 web sites. there is so much crap on most sites. re-crawl hourly. monetize via Google adwords. EDIT: I'm not sure why I'm being downvoted. I am not suggesting serving PDF's. I am suggesting serving tiny text renders of top sites, that otherwise are much too bloated. the hard part is getting the text and layout right. many people read many sites for the text IMO. So I am suggesting you make an all-text version. As an example, the front page of the New York Times right now, copied into Microsoft Word, is 2504 words. When I save from the word I copied into into .txt - I get a 16.4 KB file. By comparison, when I put the site into a Page Size Checker -- http://smallseotools.com/website-page-size-checker/ http://smallseotools.com/website-page-size-checker/ -- I get 214.23 KB. That is impressively small, and it's a fast page. If I try their competition, the Washington Post, I get 237 KB. If I try the Wall Street Journal, I get 938.15 KB -- nearly a full Megabyte. (This is actually more what I was expecting - I'm impressed by the Times.) Suppose someone desperately wants to glance at the Wall Street Journal from a poor connection where they barely get data. The difference between 12 KB and nearly a megabyte is huge. Its the difference between 4 seconds and 312 seconds: 4 seconds as compared with 5 full minutes. So there is a large need in my opinion for such a service in case someone desperately wants to see a text render. Preserving any formatting at all, helps hugely.
- 2_listerine_pls 10y agocheck docparser.com
- jlink 10y agointeresting service which was not present yet back in 2015 when I wrote my class.
- krakaukiosk 10y agoCorrect! We launched July 2016
- andreif 10y agoYeah, sure, a public one for not privacy-critical PDFs plus something like a Heroku button to build own secure app (with auth and no storage). See e.g. my file sharing app https://github.com/andreif/SecretFile https://github.com/andreif/SecretFile
- marak830 10y agoAhh this will be useful for my kitchen receipts. Thanks. Now I just need to roll that with an auto translator too.(I guess I have my day off project now :-) )
- tyingq 10y agoCurious if this works better than the pdftotext utility that comes in the Debian poppler-utils package. That has a --layout option that works really well sometimes and really terrible other times. Doesn't seem to be related to document complexity either.
- spcrngr 10y agoIt probably works reasonably well with the documents it has been tested with. It's a very hard problem to crack if you ask me. (edit: word choice)
- deleted 10y ago[deleted]
- dmoo 10y agoAlso available for windows and mac at http://www.foolabs.com/xpdf/download.html http://www.foolabs.com/xpdf/download.html
- krylon 10y agoLast year, my boss gave me a task that looked simple enough at first glance - get data on how many vacation days each employee has in total, how many they have used in the current year, and how many they have left, and put that data in our SharePoint server (so people can see when filling out a vacation request if they actually have enough days left). Most of that was fairly easy, except that the POS program that sits in the actual data only allows exporting data in one single format - PDF. Converting that PDF file to a CSV that I can feed into SharePoint was one of the nastiest things I did last year. I did manage to get it to work though, by toying around with pdftotext for a while and exploring its command line parameters. It was a pleasure to use! It took me a while to discover the correct set of command line parameters I needed, but I got it to work! Thanks, xpdf!
- tyingq 10y agoHad several somewhat similar experiences in my career. I think the general public would be surprised at the amount of duct tape and chewing gum that's behind things that appear to be important processes.
- rsync 10y agoThis is important for (al)pine users ... when reading email in a terminal it is very useful to be able to open a PDF attachment as text and view it in the (terminal) mailtool ... Yes, (al)pine is my mailtool in 2017.
- bsharitt 10y agoAlso mutt, which I've switched back to recently. I've got a little Atom powered Chromebook converted to Linux that just does not like modern heavy webmail clients(even GMail when it was still running ChromeOS, and this is one still on the market, Acer CB3-131) so a combination of mutt, mbsync, and msmtp is a much nicer combo. Mutt is a terrific mail reader but its internal SMTP and IMAP handling can be a bit iffy, hence mbsync and msmtp. Though I can generally open attachments just fine, this text rendering of PDFs would be useful for when I'm SSH'd into my home machine and reading stuff remotely(usually from work where I don't want to download my personal email).
- JetSpiegel 10y agoMutt + mbsync + msmtp is my setup too. I'm using Neomutt, since that's being actively maintained by a sizeable community of friendly people.
- zatkin 10y agoNo need to feel ashamed. I set up my own email server in 2016 and use mutt, squirrelmail, and iOS Mail very frequently.
- shakna 10y agoI use alpine in 2017. It's easy to use, pluggable, and faster than any GUI I've touched.
- phillc73 10y agoAs do I, because it's faster than browser based email and many GUI clients (like Thunderbird). I also like the fact that I can just copy across my .pinerc file to a new computer and my mail client is setup. I had not considered the PDF issue. I just open them with an external application. The potential of reading them within Alpine hadn't occurred to me, but now it has, I want it!
- kzrdude 10y agoBut does it keep both the layout and Sha-1 hash? Not sure it's HN worthy otherwise.
- nemild 10y agoFor those interested in converting PDF tables into CSV, there's also Tabula ( http://tabula.technology/ http://tabula.technology/ ) (Used by many journalists to analyze the data in PDFs)
- scrollaway 10y agoI find it absolutely ridiculous that we have to resort to these kinds of tools :/ We have digital formats, and we decided to standardize document distribution on the one that makes it as hard to extract data as if it were on physical paper.
- vog 10y agoThat's Adobe. Look at their other formats, and PDF seems to be one of their better ones. Compare to SWF, PSD, AI and so on. PDF is the successor of PostScript. PostScript is a stack-based programming language where anything can happen, while PDF enforces some document structure and metadata structure on top of it, so you can e.g. at least determine where pagebreaks are, without having to interpret ("run the code of") the whole document. Still, PDF is simpler than PostScript in the same sense that XML is a simplification of SGML. Jumping from PDF to a well-designed format would be like jumping from XML to JSON or S-Expr.
- halomru 10y agoPDF is a perfectly fine and rich digital format. It also allows you to do proper copy and paste, which is much saner than anything paper offers. Sure, PDF is a light on context clues for automation and is targeted purely at humans. But formats targeted at both computers and humans consistently fail (XML with accompanying XSLT comes to mind), and/or only have terrible tools for creating files (easily parsable, pretty HTML). Either there is very little real demand or we consistently fail at making alternatives viable.
- TeMPOraL 10y ago> It also allows you to do proper copy and paste, which is much saner than anything paper offers. Technically it does that, yes. Rarely do I see people taking advantage of it, though; most of the times I tried to copy some text out of PDF, the result had to undergo a significant cleanup before becoming usable.
- WalterGR 10y agoFairly frequently, OCR engines are posted here. But almost without exception, they lack layout analysis, which renders them largely useless. Is this something that could be combined with those OCR engines? (e.g. TesseractOCR...)
- 2_listerine_pls 10y agosome services allow you to set the layout manually: Docparser
- eumm 10y agoPDF.co offline tool (for Windows) supports OCR and partial OCR for pdf to text and pdf to csv with layout preserved. (disclaimer: i work on it)
- RandomBookmarks 10y agoI would not call these services useless ;) - but I wonder the same... Some apis like https://ocr.space https://ocr.space return the coordinates of each converted word. Can that be a used input? (I have not tried it yet)
- PretzelFisch 10y agoephesoft seems to use this for classifying and data extraction from documents.
- robinhowlett 10y agoNice. I recently got very familiar with PDFBox and parsing complex layouts - it is a great library.
- agumonkey 10y agoFun, I did the same thing as a clojure repl exploration to pipe PDF text to a bare Swing GUI (I know, a little absurd in a way). The deja vu made squint for a minute. ps: pdfbox is nice
- Animats 10y agoIs there a PDF to HTML converter which can consistently get line breaks right?
- curiousgal 10y agoAlthough I haven't tested this yet, these utilities tend to fail when fed a table with empty cells.
- gpvos 10y agoThe first example image in the linked article shows a conversion from a table with some empty cells. It looks fine.
- curiousgal 10y agoThose are at the end. I meant empty cells in the middle. The ones I tried don't account for them.
- jlink 10y agoIt works also with empty cells in the middle.
- ykaranfil 10y agoIn case it helps someone, for a data mining project at the research lab i work, i tried more than 10 different commercial and opensource libraries. The best one was commercial version of Foxit SDK, it always kept the layout perfectly. Function you need to use: FPDFText.FPDFText_PDFToText()