4 ms·
Show HN: BetterOCR combines and corrects multiple OCR engines with an LLM
- harwoodjp 3y agoCool!
- junhoyeo 3y agoMy pleasure
- tamimio 3y agoInteresting, will see how it performs
- junhoyeo 3y agoThanks!
- jamesnorden 3y agoIs the OpenAI key paid-only?
- raybb 3y agoSimilarly, is there any half decent general LLM that can run on consumer hardware?
- om154 3y agoI’ve been looking for an OCR engine and hoped that using LLMs would improve their output. Looks great! I’ll give it a go.
- junhoyeo 3y agoSuper -- I'm thrilled you enjoyed it!
- meetingthrower 3y agoAwesome - I'm a dabbler, but any thoughts on best engines for PDF tables? I've got tons of PDFs with similar tables embedded deep in them, but all formatted slightly differently. Seems like it should be easy....but nope!
- janderson215 3y agoAre you able to highlight the text on the PDF? If so, I highly recommend PDF2TXT to extract text from PDFs. Would require some parsing work on your part to convert it back to a table, but zero chance of error from inference since it’s using text extraction. If you can’t highlight the text, it won’t work.
- junhoyeo 3y agoThanks! PDF -> Markdown looks like a pretty great use case Just added box detection support -- maybe I'll start from here https://github.com/junhoyeo/BetterOCR#-box-detection https://github.com/junhoyeo/BetterOCR#-box-detection
- 2Gkashmiri 3y agoTabula. It does what you are looking for
- bomewish 3y agoI use popper pdftotext -layout flag, which preserves the shape of the tables. Gpt can then fix them up into whatever format.
- snac 3y agoWow, this is just up my alley. I started a side project recently using Tesseract to read book spines for inventory purposes and hooked it up to ChatGPT to clean up the text, having it "fill in the blanks" so to speak. I'll definitely give this a go, having using two OCR engines I should get better results. Any plans to add other OCR engines?
- junhoyeo 3y agoYup! But I'm still exploring options. (any recommendations would be welcomed!) Here are some candidates I'm considering: - https://github.com/mindee/doctr https://github.com/mindee/doctr - https://github.com/open-mmlab/mmocr https://github.com/open-mmlab/mmocr - https://github.com/PaddlePaddle/PaddleOCR https://github.com/PaddlePaddle/PaddleOCR (honestly I don't know Mandarin so I'm a bit stuck) - https://github.com/clovaai/donut https://github.com/clovaai/donut -- While it's primarily an "OCR-free document understanding transformer," I think it's worth experimenting with. Think I can sort this out by letting the LLM reason through it multiple times (although this will impact performance) - yesterday got a suggestion to consider https://github.com/kakaobrain/pororo https://github.com/kakaobrain/pororo -- don't think development is still active but the results are pretty great on Korean text
- qwerty456127 3y agoCool. Is there also a library which can be combined with this to automatically translate languages except some I understand (which I would specify as a list of ISO codes)?
- deleted 3y ago[deleted]