5 ms·
Make an API out of that. There is huge demand for it.
by dempseye 5y ago
Make an API out of that. There is huge demand for it.
- mvcatsifma 5y agoI believe such an API already exists: https://pdftables.com/ https://pdftables.com/ (no affiliation). Went to a presentation at a Golang meetup in Amsterdam by the guys behind this company. Seemed to know their stuff. But I have no real world experience using it.
- vic-traill 5y agoI had a look - from their FAQ [0]: However, some PDFs are scanned documents, or only contain images. PDFTables doesn't perform Optical Character Recognition (OCR) to turn these images into text. To process these kinds of documents, you will need to either enable OCR in your scanning software, or run the PDF through specialist OCR software before using PDFTables. [0] https://pdftables.com/faq https://pdftables.com/faq
- leesalminen 5y agoI’m in need of that right now! Would pay for it, especially if the proceeds benefitted that non-profit!
- mlboss 5y agoLooks this does it https://www.extracttable.com/ https://www.extracttable.com/
- mediascreen 5y agoI'm pretty sure AWS Textract does that already. https://aws.amazon.com/textract/ https://aws.amazon.com/textract/
- agustif 5y agoI needed to do this, so happy someone else made it! A nice FOSS library for it was on HN recently, sharing is caring I guess https://news.ycombinator.com/item?id=28680136 https://news.ycombinator.com/item?id=28680136 https://extract-table.com/ https://extract-table.com/ https://github.com/vegarsti/extract-table https://github.com/vegarsti/extract-table