4 ms·
If this can get me tables out of pdf's generated by crystal reports it would be a godsend for testing. This has been a nightmare to try and solve, the best opti
by BasHamer 8y ago
If this can get me tables out of pdf's generated by crystal reports it would be a godsend for testing. This has been a nightmare to try and solve, the best option so far has been adobe cloud but they don't offer an API for that. I'm excited to try it out.
- mjt58 8y agoHave you tried e.g. https://tabula.technology https://tabula.technology, https://pdftables.com https://pdftables.com, https://pypi.org/project/Camelot/ https://pypi.org/project/Camelot/?
- BasHamer 8y agohttps://pdftables.com https://pdftables.com failed the test file, pretty good but inconsistent interpretation across rows, sometimes it split the cell, sometimes it did not. Tabula failed to detect multi-line rows, after manually changing the table it did do better than pdftables.com on splitting cells. Both failed the non-printable whitespace characters that created garbled outputs in the excel. The other one would take some time to rig up.
- ocrcustomserver 8y agoYou can also try https://docparser.com/ https://docparser.com/. If nothing works for you and you're comfortable with sharing an example file, you can send it to me and I could take a look.
- deleted 8y ago[deleted]
- cdolan 8y agoRather than the Camelot link you provided, I think you meant Excalibur? https://github.com/camelot-dev/excalibur https://github.com/camelot-dev/excalibur
- mjt58 8y agoOh yes, thanks :-)
- counciltime 8y agoI have a friend who has also developed a number of applications that use OCR specifically for PDF which uses Tesseract. The Report Miner application does a nice job of locating and extracting PDF tables. https://www.opait.com/tesseractstudio/ https://www.opait.com/tesseractstudio/ https://www.opait.com/Pdfreportminer/ https://www.opait.com/Pdfreportminer/
- minhtripham 8y agoWould love to learn more about the apps your friend developed--currently doing research into different OCR use cases + tech. can you shoot me an email at minh@docucharm.com?
- RandomBookmarks 8y agoHow about https://ocr.space/tablerecognition https://ocr.space/tablerecognition It returns table data line by line.
- BasHamer 8y agohandled the non-printed whitespace but butchered the multi- line table headers, so re-building the headers is rough as it is line by line and you need to know what words go together and you have lost the structure.
- cdolan 8y agoCan you send me a copy of what you are trying to extract? We use proprietary stuff (we're in the business of extracting data and performing analysis on invoices for waste, recycling, cellular, etc... stuff that gets "lost" in the AP department. Happy to see if our tools can help. I've tried everything on the market - DocParser, MediusFlow, KOFAX, Ephesoft, etc... none work well enough in my opinion.