4 ms·
Is Tesseract still the best choice for local OCR in 2026? I was always underwhelmed with its real-world performance.
by rickcarlino 1mo ago
Is Tesseract still the best choice for local OCR in 2026? I was always underwhelmed with its real-world performance.
- _joel 1mo agoThere's EasyOCR and RapidOCR too, I guess benchmark and see what's best for your material? Oh and Multimodal LLMs :)
- polycancel 1mo agoBeing using granite model pretty small and good. Can't do JSON, but don't really care. Which models are EasyOCR and RapidOCR using?
- thiagolima 1mo agoI did a small benchmark for RapidOCR: https://thiagotigaz.github.io/ocr-it/bench/ https://thiagotigaz.github.io/ocr-it/bench/ For the input text rendered on screen, Tesseract did better on both accuracy and speed. We got about 0.1% character error vs 1–2% for RapidOCR, and Tesseract was roughly 2.5x faster. Blur was the biggest difference: 0.4% vs 14%. The big problem is that this is synthetic rendered text, which is basically the easy case and also the only kind of input this extension captures. I wouldn't assume the same results for scanned documents. I haven't tested EasyOCR yet.
- _joel 1mo agoLol, yea, not so rapid then! :)
- LeonardoTolstoy 1mo agoI, at this point, use Qwen2.5-VL-3B-Instruct for most of the small OCR I want to do. It is much much better than my experience with Tesseract in general. The nice thing about it is that if you give it, say, a movie poster you can ask for the "title of the movie" and it will, to the best of its ability, do just that, no need for regex or filtering after. For smallish images after loading the 3B model runs in <1 second. 7B takes longer but is obviously more accurate. I might be a bit behind, all of this is from early this year for the most part, but for something like "I have 3000 movie posters and I want to get the titles with like 90% accuracy" it is good (much better than Tesseract), and it'll do that in like an hour. EDIT: I guess one thing is Tesseract will kind of give gibberish back when it fails. The main issue with the LLMs are that instead they take a stab at it (like for a movie poster it'll give part of a quote, or a actor name) back. Makes knowing when it fails a little harder. As long as you have some way to verify when it is likely failing they are very good though.
- thiagolima 1mo agoLeo, i benchmarked Qwen2.5-VL-3B, 4-bit via MLX, against Tesseract on the same 24 samples: https://thiagotigaz.github.io/ocr-it/bench/ https://thiagotigaz.github.io/ocr-it/bench/ Its much lower and the error rate is much higher. It works, but for our usecase, "clean rendered text" (extract from kindle for example) tesseract is much better.
- zzleeper 1mo agoDefinitely not. Even Chrome has a built in OCR that performs amazingly. I got an LLM to write a quick python wrapper to it [1], so I'm sure you should be able to access it from an extension [1] https://github.com/sergiocorreia/clv-locro https://github.com/sergiocorreia/clv-locro
- deivid 1mo agoPaddlePaddle (v6) is fantastic and fast
- HackerThemAll 1mo agoIt works very well for me. It handles non-English characters and diacritics without issues.
- thiagolima 1mo agoWe did some benchmarking against other libraries like RapidOCR and EasyOCR results in https://thiagotigaz.github.io/ocr-it/bench/ https://thiagotigaz.github.io/ocr-it/bench/