4 ms·
Have you tried e.g. https://tabula.technology https://tabula.technology, https://pdftables.com https://pdftables.com, https://pypi.org/project/Camelot/ https://
by mjt58 8y ago
Have you tried e.g. https://tabula.technology https://tabula.technology, https://pdftables.com https://pdftables.com, https://pypi.org/project/Camelot/ https://pypi.org/project/Camelot/?
- BasHamer 8y agohttps://pdftables.com https://pdftables.com failed the test file, pretty good but inconsistent interpretation across rows, sometimes it split the cell, sometimes it did not. Tabula failed to detect multi-line rows, after manually changing the table it did do better than pdftables.com on splitting cells. Both failed the non-printable whitespace characters that created garbled outputs in the excel. The other one would take some time to rig up.
- ocrcustomserver 8y agoYou can also try https://docparser.com/ https://docparser.com/. If nothing works for you and you're comfortable with sharing an example file, you can send it to me and I could take a look.
- deleted 8y ago[deleted]
- cdolan 8y agoRather than the Camelot link you provided, I think you meant Excalibur? https://github.com/camelot-dev/excalibur https://github.com/camelot-dev/excalibur
- mjt58 8y agoOh yes, thanks :-)
- counciltime 8y agoI have a friend who has also developed a number of applications that use OCR specifically for PDF which uses Tesseract. The Report Miner application does a nice job of locating and extracting PDF tables. https://www.opait.com/tesseractstudio/ https://www.opait.com/tesseractstudio/ https://www.opait.com/Pdfreportminer/ https://www.opait.com/Pdfreportminer/
- minhtripham 8y agoWould love to learn more about the apps your friend developed--currently doing research into different OCR use cases + tech. can you shoot me an email at minh@docucharm.com?