5 ms·
Show HN: an API to extract text from a PDF
- ra 13y agoNice. Why no paid options? I'm guessing because this was a weekend project. If so, nice work!
- trez 13y agoThanks. The paying option would be coming if there is some interests in it. If there is people interested in it, please let us know in this thread or at info@stamplin.com.
- architgupta 13y agoDo you do OCR for text extraction?
- TillE 13y agoNeat, but practically who would want to do this with an API rather than installable software?
- kijin 13y agoI have some questions: 1. Why return an array of texts? Where do the texts get split up? At page boundaries? Column boundaries? At the end of each line? If a line is interrupted by a corner of an image and continues a couple of inches afterward, does it get treated as a separate text? (I once used a PDF->text extractor program that spit out every word sepearately, often in an incorrect order. That probably had to do with how the PDF was organized internally.) 2. "The PDF file should be smaller than 1 Mbit" -> You mean 1 megabyte, right? Because 1 megabit is only 125-128 kilobytes.
- trez 13y ago1. That's indeed dependant on the way the PDF has been created. Sometimes, you can even have letter splitted on different text token. 2. You're right, I mean Megabyte
- kijin 13y agoThanks for the clarifications. Since a lot of PDFs are badly organized (and I wonder if some programs deliberately do that to make text extraction difficult), perhaps you could try to analyze the location of each token on the page and merge the ones that seem to belong together. That would be already 100x better than most of the free PDF->text converters out there.
- alxbrun 13y agoOr even go the OCR approach !
- ygra 13y agoAll things considered that's pretty sad, though. A digital archive format that cannot reliably be read by machines, even if it contains just text.
- trez 13y agowe are already close to do that but with a really slow parser (this one can even replace some text on the pdf). Our problem now is to understand if developers would rather have better text extraction or some other features like image extractions, etc.. Let us know what you would prefer.
- rcfox 13y agoI've recently been working on extracting text from PDFs myself. I've found that `pdftohtml -xml` from the Poppler utils does a decent job of it, and includes a bounding box for each piece of text. I've submitted a few patches to their Bugzilla to also include the transformation matrix as well as some extra styling information.
- zdw 13y agoIf you're doing this local/cli `pdftext`, from http://www.foolabs.com/xpdf/ http://www.foolabs.com/xpdf/ For OCR, `pdfimages` (also from xpdf), combined with ImageMagick's `convert`, and `tesseract` (http://code.google.com/p/tesseract-ocr/ http://code.google.com/p/tesseract-ocr/) works passably well.
- surapaneni 13y agoThis is similar to what we do at http://searchtower.com http://searchtower.com , where you can store, view, index and search the data.
- midas 13y agoGoing from PDF to nicely formatted word doc would be huge for lawyers and people who do a lot of contract negotiations. It's hard to do well though.
- meomix 13y agoWord 2013 natively supports opening Word Docs, including some advanced features like tables, bookmarks, etc.
- zeckalpha 13y agoThis gets you some of the way, but if the two products were combined... http://www.docverter.com/api.html http://www.docverter.com/api.html
- adsr 13y agoIsn't a better solution to get software that supports PDF natively like Acrobat, then edit the documents there instead of doing a (poor) translation to Word.
- alkou 13y agodo you use pdftotext internally or something else?
- trez 13y agowe have our own parser for more complicated task and use xpdf when speed is key because it's much faster.
- aksx 13y agoI am writing something similar for a client. He needs data in tables extracted from the PDF. Which language are you using? I wrote two scripts, one using python and pdftotext and another using ruby pdf-reader, the ruby one gives each line of the PDF one by one which is good for extraction.
- deleted 13y ago[deleted]
- chenster 13y agoI googled "converting PDF to text" and "converting PDF to html". A tons of services already exist out there. Apparently, it's not something new. How do you plan to compete? Are you planning to focus on data extraction rather than conversion?
- trez 13y agothe initial plan wasn't about text extraction but text modification. We noticed we already have something to "give" and created this service. Following the lean methodology, we hope we are going to get some insights about the next step.
- hnriot 13y agoWhy not just system(pdf2html) - I don't see the point since this level of functionality is trivially achieved. If it did something over and above that it might be useful, like OCR, but even that's not hard to add.
- trez 13y agoindeed, at the moment, that's quite simple. The only benefit is the fact is quite easy to integrate and fast. More advance feature should be coming soon.
- ismaelc 13y agoHey I've documented this in Mashape - https://www.mashape.com/ismaelc/extract-text-from-pdfs#!documentation https://www.mashape.com/ismaelc/extract-text-from-pdfs#!docu...
- trez 13y agothanks!