3 ms·
I was really excited to try this until I saw that the only extraction methods are pdfminer, finereader, and tesseract. I was hoping there was something you roll
by staticautomatic 7y ago
I was really excited to try this until I saw that the only extraction methods are pdfminer, finereader, and tesseract. I was hoping there was something you rolled on your own. I've been trying for a long time to parse tables (and nested tables) but the available extractors seem to only work on really simple, idealized tables with virtually no skew or warping. The best I've found so far as Amazon's Textract, but it's not that great either. Alas, every attempt I've ever made at generalized table extraction has quickly regressed to templates.
- pierre 7y agoWe also support Google Document understanding API for OCR, with support for other cloud OCR vendor coming soon. We also support pdf.js as an alternative to pdfminer.
- staticautomatic 7y agoThanks. FYI the link to the Google Vision documentation in 2.1 Extractor Tools of your documentation is broken.
- chrisweekly 7y agoMaybe something like lnav (https://lnav.org https://lnav.org) would suit your needs? Edit: I mean as a part of a custom solution
- udayrddy 7y agoShameless plug: Would you like to join the club of happy customers at https://extracttable.com https://extracttable.com - API to extract tabular data from images and PDFs without worrying about co-ordinates. A comprehensive competitor comparison, along with outputs, is available at https://extracttable.com/compare.html https://extracttable.com/compare.html
- iudqnolq 7y agoSuggestion: use a simpler synonym for idempotent on your high level pricing overview. Something like "automatically makes sure you aren't charged for duplicates"