4 ms·
Wow, this is just up my alley. I started a side project recently using Tesseract to read book spines for inventory purposes and hooked it up to ChatGPT to clean
by snac 3y ago
Wow, this is just up my alley. I started a side project recently using Tesseract to read book spines for inventory purposes and hooked it up to ChatGPT to clean up the text, having it "fill in the blanks" so to speak. I'll definitely give this a go, having using two OCR engines I should get better results.
Any plans to add other OCR engines?
- junhoyeo 3y agoYup! But I'm still exploring options. (any recommendations would be welcomed!) Here are some candidates I'm considering: - https://github.com/mindee/doctr https://github.com/mindee/doctr - https://github.com/open-mmlab/mmocr https://github.com/open-mmlab/mmocr - https://github.com/PaddlePaddle/PaddleOCR https://github.com/PaddlePaddle/PaddleOCR (honestly I don't know Mandarin so I'm a bit stuck) - https://github.com/clovaai/donut https://github.com/clovaai/donut -- While it's primarily an "OCR-free document understanding transformer," I think it's worth experimenting with. Think I can sort this out by letting the LLM reason through it multiple times (although this will impact performance) - yesterday got a suggestion to consider https://github.com/kakaobrain/pororo https://github.com/kakaobrain/pororo -- don't think development is still active but the results are pretty great on Korean text