5 ms·
A timely link given there is a current discussion on the trouble with extracting meaningful text from pdfs in another thread on the front page! I look forward t
by GEBBL 6y ago
A timely link given there is a current discussion on the trouble with extracting meaningful text from pdfs in another thread on the front page! I look forward to reading the feedback on actual use of this
- nayuki 6y agoFor reference, it's https://news.ycombinator.com/item?id=24460142 https://news.ycombinator.com/item?id=24460142 ; https://filingdb.com/b/pdf-text-extraction https://filingdb.com/b/pdf-text-extraction
- mumblemumble 6y agoTika's PDF text extraction is fine if you're just trying to get searchable text. Which is what it's made for: Slurping doucments into Lucene. Fulltext search typically isn't terribly sensitive to getting the order of words right, and is even less sensitive to getting the formatting right. If you're trying to get something fit for consumption by human (including via a screen reader) or an NLP pipeline, though, all the problems discussed in that FilingDB article still apply.
- tonitosou 6y agoaspose can convert pdf to html
- chaps 6y agoFirefox's PDF reader does the same. A few years back, I wrote wrote a pdf to csv converter with selenium -- it worked surprisingly well! Though, after I finished I found tabula and the code became immediately useless, hah.
- rovr138 6y agoWhen that happens to me, it's due to a keyword I hadn't thought of looking for. Then, while building the project, it came to me.
- Ocelot20 6y agoI used Tika to build a search engine prototype, and it was fantastic for getting us up and running quickly. It's a really easy to use generic parser for a bunch of document types. The downside of being so generic and easy to use is that you end up lacking document-specific context that could be useful. For example: Do you consider the header/footer text to be important, or just noise (Page 1, Page 2, etc.)? Is the text contained in the Table of Contents or section headers important, or just the actual content? You won't find any ways to tweak the result, which could be a good or bad thing depending on your use case. We ended up using it as our "fallback" parser, writing more contextually aware ones for document types of greater importance to our use case (PDFs were high on the list).
- tonitosou 6y agoso how did u do to understand the "form" of the document such as table of contents and co.
- iav 6y agoI’ve found Tika’s PDF to HTML parser to be pretty good. My only complaint is that in a double spaced document where there is equal amount of space between paragraphs and normal lines, it labels every line as a separate paragraph.