39 ms·
There's EasyOCR and RapidOCR too, I guess benchmark and see what's best for your material? Oh and Multimodal LLMs :)
by _joel 1mo ago
There's EasyOCR and RapidOCR too, I guess benchmark and see what's best for your material? Oh and Multimodal LLMs :)
- polycancel 1mo agoBeing using granite model pretty small and good. Can't do JSON, but don't really care. Which models are EasyOCR and RapidOCR using?
- thiagolima 1mo agoI did a small benchmark for RapidOCR: https://thiagotigaz.github.io/ocr-it/bench/ https://thiagotigaz.github.io/ocr-it/bench/ For the input text rendered on screen, Tesseract did better on both accuracy and speed. We got about 0.1% character error vs 1–2% for RapidOCR, and Tesseract was roughly 2.5x faster. Blur was the biggest difference: 0.4% vs 14%. The big problem is that this is synthetic rendered text, which is basically the easy case and also the only kind of input this extension captures. I wouldn't assume the same results for scanned documents. I haven't tested EasyOCR yet.
- _joel 1mo agoLol, yea, not so rapid then! :)