3 ms·
Tested it with the following documents: * Loan application form: It picks up checkboxes and handwriting. But it missed a lot of form fields. Not sure why? * E
by constantinum 2y ago
Tested it with the following documents:
* Loan application form: It picks up checkboxes and handwriting. But it missed a lot of form fields. Not sure why?
* Edsger W. Dijkstra’s handwritten notes(from Texas univ archive) - Parsing is good.*
* Badly(misaligned) scanned bill - Parsing is good. Observation: there is a name field, but it produced a synonymous name instead of the name in the bill — hallucination??
* Investment fund factsheet - It could parse the bar charts and tables, but it whimsically excluded many vital data points from the document.
* Investment fund factsheet, complex tables - Bad extraction, could not extract merged tables and again whimsical elimination of rows and columns.
Anyone curious, try LLMWhisperer[1] for OCR. It doesn't use LLMs, so no hallucination side effects. It also preserves the layout of the input document for more context and clarity.
There's also Docling[2], which is handy for converting tables from PDFs into markdown. While it uses Tesseract/EasyOCR under the hood, which can sometimes make the OCR results a bit less accurate
[1] - https://pg.llmwhisperer.unstract.com/ https://pg.llmwhisperer.unstract.com/
[2] - https://github.com/DS4SD/docling https://github.com/DS4SD/docling
- phren0logy 2y agoFYI, you can choose which OCR engine Docling uses (from a handful of predefined choices) - it doesn’t have to be Tesseract. https://ds4sd.github.io/docling/reference/pipeline_options/#docling.datamodel.pipeline_options.OcrEngine https://ds4sd.github.io/docling/reference/pipeline_options/#...