4 ms·
"World's best OCR model" - that is quite a statement. Are there any well-known benchmarks for OCR software?
by ChemSpider 2y ago
"World's best OCR model" - that is quite a statement. Are there any well-known benchmarks for OCR software?
- xnx 2y agohttps://huggingface.co/spaces/echo840/ocrbench-leaderboard https://huggingface.co/spaces/echo840/ocrbench-leaderboard
- ChemSpider 2y agoInteresting. But no mistral on it yet?
- themanmaran 2y agoWe published this benchmark the other week. We'll can update and run with Mistral today! https://github.com/getomni-ai/benchmark https://github.com/getomni-ai/benchmark
- kergonath 2y agoExcellent. I am looking forward to it.
- cdolan 2y agoCame here to see if you all had run a benchmark on it yet :)
- themanmaran 2y agoUpdate: Just ran our benchmark on the Mistral model and results are.. surprisingly bad? Mistral OCR: - 72.2% accuracy - $1/1000 pages - 5.42s / page Which is pretty far cry from the 95% accuracy they were advertising from their private benchmark. The biggest thing I noticed is how it skips anything it classifies as an image/figure. So charts, infographics, some tables, etc. all get lifted out and returned as [image](image_002). Compared to the other VLMs that are able to interpret those images into a text representation. https://github.com/getomni-ai/benchmark https://github.com/getomni-ai/benchmark https://huggingface.co/datasets/getomni-ai/ocr-benchmark https://huggingface.co/datasets/getomni-ai/ocr-benchmark https://getomni.ai/ocr-benchmark https://getomni.ai/ocr-benchmark
- Thaxll 2y agoDo you benchmark the right thing though? It seems to focus a lot on image / charts etc... The 95% from their benchmark: "we evaluate them on our internal “text-only” test-set containing various publication papers, and PDFs from the web; below:" Text only.
- themanmaran 2y agoOur goal is to benchmark on real world data. Which is often more complex than plain text. If we have to make the benchmark data easier for the model to perform better, it's not an honest assessment of the reality.
- WhitneyLand 2y agoIt’s interesting that none of the existing models can decode a Scrabble board screen shot and give an accurate grid of characters. I realize it’s not a common business case, came across it testing how well LLMs can solve simple games. On a side note, if you bypass OCR and give models a text layout of a board standard LLMs cannot solve Scrabble boards but the thinking models usually can.
- resource_waste 2y agoIts Mistral, they are the only homegrown AI Europe has, so people pretend they are meaningful. I'll give it a try, but I'm not holding my breath. I'm a huge AI Enthusiast and I've yet to be impressed with anything they've put out.