5 ms·
How does an LLM approach to OCR compare to say Azure AI Document Intelligence (https://learn.microsoft.com/en-us/azure/ai-services/document-intelligence/overvie
by yoran 1y ago
How does an LLM approach to OCR compare to say Azure AI Document Intelligence (https://learn.microsoft.com/en-us/azure/ai-services/document-intelligence/overview?view=doc-intel-4.0.0 https://learn.microsoft.com/en-us/azure/ai-services/document...) or Google's Vision API (https://cloud.google.com/vision?hl=en https://cloud.google.com/vision?hl=en)?
- sandblast 1y agoNot sure why you're being downvoted, I'm also curious.
- ozgune 1y agoOmniAI has a benchmark that companies LLMs to cloud OCR services. https://getomni.ai/blog/ocr-benchmark https://getomni.ai/blog/ocr-benchmark (Feb 2025) Please note that LLMs progressed at a rapid pace since Feb. We see much better results with the Qwen3-VL family, particularly Qwen3-VL-235B-A22B-Instruct for our use-case.
- CaptainOfCoit 1y agoMagistral-Small-2509 is pretty neat as well for its size, has reasoning + multimodality, which helps in some cases where context isn't immediately clear, or there are few missing spots.
- cheema33 1y agoOmni OCR team says that according to their own benchmark, the best OCR is the Omni OCR. I am quite surprised.
- numpad0 1y agoClassical OCR still probably make undesirable su6stıtutìons in CJK from there being far too many of similar ones, even some absurd ones that are only distinguishable under microscope or by looking at binary representations. LLMs are better constrained to valid sequences of characters, and so they would be more accurate. Or at least that kind of thing would motivate them to re-implement OCR with LLM.
- fluoridation 1y agoHuh... Would it work to have some kind of error checking model that corrected common OCR errors? That seems like it should be relatively easy.
- colonCapitalDee 1y agoIt's harder then it first seems. The root problem is that for text like "hallo", correcting to "hello" may be fixing an error or introducing an error. In general, the more aggressive your error correction, the more errors you inadvertently introduce. You can try and make a judgement based on context ("hallo, how are you?"), which certainly helps, but it's only a mitigation. Light error correction is common and effective, but you can't push it to a full solution. The only way to fully solve this problem is to look at the entire document at once so you have maximum context available, and this is what non-traditional OCR attempts to do.
- fluoridation 1y agoOkay, but there way more common errors that should be easy to fix. "He11o", "Emest Herningway", incorrect diacritics like the other person mentioned, etc.
- make3 1y agoaren't all of these multimodal LLM approaches, just open vs closed ones
- daemonologist 1y agoMy base expectation is that the proprietary OCR models will continue to win on real-world documents, and my guess is that this is because they have access to a lot of good private training data. These public models are trained on arxiv and e-books and stuff, which doesn't necessarily translate to typical business documents. As mentioned though, the LLMs are usually better at avoiding character substitutions, but worse at consistency across the entire page. (Just like a non-OCR LLM, they can and will go completely off the rails.)
- stopyellingatme 1y agoNot sure about the others but we use Azure AI Document Intelligence and its working well for our resume parsing system. Took a good bit of tuning but we havent had to touch it for almost a year now.
- junto 11mo agoNot sure how it compares but we did some trials with Azure AI Document Intelligence and were very surprised at how good it was. We had a document example which was a poor photograph of a document that had quite a skew, and it (too our surprise), also detected the customer’s human legible signature and extracted their name from that signature.