10 ms·
Mistral OCR 4
- Ducki 4mo agoI was processing 55 year old paper files, most of them severely degraded, with its predecessor model. I was very impressed! I also tried Abbyy Finereader but it didn't even come close in my experience.
- philipkglass 4mo agoI used Abbyy Finereader for several years. I loved it. I completed some large projects with it. Modern VLMs put classic FineReader to shame for processing low-resolution/degraded/non-standard text. I'm personally using the small Qwen 3.5 models. If you have an OCR problem, Mistral OCR 4 is probably great. Open weights models that you can run on a laptop may also work great.
- jppope 4mo agoIs there something wrong with their certificate? Chromium is saying https isn't valid
- collabs 4mo agoLooks good to me on both brave (on android) and firefox (on windows 11). Lets see what ssl labs says (it is running now) https://www.ssllabs.com/ssltest/analyze.html?d=mistral.ai&latest https://www.ssllabs.com/ssltest/analyze.html?d=mistral.ai&la... Looks good so far, A+ on ipv4 as well as ipv6 Edit: I also asked Gemini 3.1 Pro to analyze the certificate and it looks good It looks like you have shared an `about:certificate` URL containing a chain of three Base64-encoded X.509 TLS/SSL certificates. This specific chain is used to secure connections to *mistral.ai*. Here is the decoded breakdown of the certificate chain you provided: ## Certificate Chain Overview This is a standard three-tier certificate chain issued by Google Trust Services for the Mistral AI domain. --- ### 1. Leaf Certificate (End-Entity) This is the specific certificate issued to the website to verify its identity and encrypt traffic. * *Subject (Common Name):* `mistral.ai` * *Subject Alternative Names (SANs):* `mistral.ai`, `workers.mistral.ai` * *Issuer:* WE1 (Google Trust Services) * *Valid From:* June 13, 2026 * *Valid To:* September 11, 2026 * *Key Type:* Elliptic Curve (ECDSA) ### 2. Intermediate Certificate This certificate acts as a bridge between the website's certificate and the trusted Root CA. * *Subject:* WE1 (Google Trust Services) * *Issuer:* GTS Root R4 (Google Trust Services LLC) * *Valid From:* December 13, 2023 * *Valid To:* February 20, 2029 * *Key Type:* Elliptic Curve (ECDSA) ### 3. Root Certificate This is the foundational trust anchor pre-installed in browsers and operating systems. * *Subject:* GTS Root R4 (Google Trust Services LLC) * *Issuer:* GTS Root R4 (Self-signed) * *Valid From:* June 22, 2016 * *Valid To:* June 22, 2036 * *Key Type:* Elliptic Curve (ECDSA)
- jppope 4mo agothanks I'm going to have to check whats going on with my setup then
- mdrzn 4mo agoIt'll be interesting to see how this ranks against https://github.com/baidu/Unlimited-OCR https://github.com/baidu/Unlimited-OCR
- cdnsteve 4mo agoRight, just announced https://x.com/BaiduAI_News/status/2069322806748410291 https://x.com/BaiduAI_News/status/2069322806748410291
- ge96 4mo ago1000 pages for $4? damn how does it compare to llama parse I wonder
- thenthenthen 4mo agoOr Apples local OCR/Vision models?
- aliljet 4mo agoI was just using infinity parser 2 (flash, to be fair) for pennies self-hosted to run through thousands of pages of documents with remarkable confidence. I decided to use https://huggingface.co/datasets/allenai/olmOCR-bench https://huggingface.co/datasets/allenai/olmOCR-bench to determine what was the best OCR tool, yesterday, but I've got no idea what the best is now. What is the dominant OCR eval right now? Between Baidu and Mistral this morning, I wonder if there's a new tool to switch to..
- freezed8 4mo ago(jerry from llamaindex here) we're gonna benchmark on ParseBench and report the results!
- tdubey 4mo agoAre there benchmarks for how this performs on charts, or maybe more accurately, plots? I've yet to find a model that can digitize a plot into X,Y points with some accuracy in my use case of digitizing old datasheets.
- utopiah 4mo ago"A note on out-of-scope use. OCR 4 is a document-understanding model, not a decision-maker. It is not intended for medical diagnosis, legal advice or judgment, high-stakes financial decisions, safety-critical systems, real-time/latency-sensitive processing, or non-document inputs (raw audio, video, etc.). " Can't wait for the "oh so innovative" manager who will suggest during the next meeting "Ok... but what if WE used it for high-stakes financial decisions on non-document inputs like a photo from my phone?" I guarantee you somebody on HN is going to comment about this "idea" next week.
- weird-eye-issue 4mo agoWhy would anybody do that you would simply get terrible results compared to dozens of other more capable models. It's for converting to text not answering questions. Just seems like you need some sort of weird angle to bring out an anti AI stance
- alex43578 4mo agoI think his comment is referring to a scenario where a decision is made on financial numbers that are misrecognized. E.g. 9.0% actual is OCR’d as 90%
- weird-eye-issue 4mo agoI don't think so But anyways just a side note one way to help reduce these errors is if you pass in both the original image and the OCR'd text to the models that make the decisions
- utopiah 4mo agoGuess you haven't met management yet. Clearly nobody should do that but that official warning is not going to stop them from trying.
- leoc 4mo ago“I delegated critical financial decisions to my OCR software, and you won’t believe what happened next.”
- gpm 4mo agoDo these models (this one or its competitors) do handwriting recognition?
- weird-eye-issue 4mo agoIf you mean handwriting to text then yes
- gpm 4mo agoYep that's what I mean, thanks :)
- 9cb14c1ec0 4mo agoYes, we have successfully used Mistral OCR for digitizing handwritten forms. You always have low percentage that need human review and adjustment, but overall Mistral has been highly accurate (their price is amazing, too).
- observationist 4mo agoIn the sense that you can get similarity scores for individual characters referenced against a known database of characters written by various individuals. You can get stylometry scores out of small LLMs that do demographic segmentation based on writing style using the same methods. They won't have the capacity to be fed an image of handwritten text and say "Ahh, this is a note written by Winston Churchill!". You could very easily use these models and your agent framework of choice, like Hermes, the Segment Anything models, and other foss tooling to build a dedicated, specialist handwriting recognition system. Or facial recognition, or fingerprint recognition, etc - these sorts of things can be done very procedurally, without a lot of interpretive AI.
- varenc 4mo agoI think OP meant converting handwriting to text, not identifying a person based on their handwriting style! (but that sounds quite interesting)
- 4mo ago
- Insanity 4mo agoRecently I tied OCR with Opus 4.8. (I know, not technically right tool for the job). All I needed to do was extract dates from receipts. It got about 20% of the dates wrong yet rated all as “high confidence”. Should have probably tried a more OCR specific model
- nik736 4mo agoOpus is very good at OCR. Way better than the small 1-4B VLMs. If Opus failed, most likely those smaller models will fail as well.
- MostlyStable 4mo agoHow long have you been testing this? Have you noted a large improvement? I tested Opus for this quite a while ago (maybe 4.5? Whatever was out about a year ago), and it performed quite poorly on my use case.
- nik736 4mo agoI have put together an internal benchmark on 1000s of business documents with weird tables, structure, etc. that I run on every relevant model release. Opus 4.8 performs very very well. But it is obviously overkill for the task (and expensive at doing so). I just wanted to respond to the OP.
- apawloski 4mo agoI’m curious what your findings are for the best model for your use case
- Insanity 4mo agoI'm assuming that the reason I didn't have good success rate is because it was not scanned documents, but photographs, and lighting conditions weren't always ideal. I think scanned business documents are a happy-case scenario in a way. (obv, you seem to run it against some complex documents, so that's impressive)
- stri8ted 4mo agoWay too expensive. Google vision OCR (which they failed to compare against), is $1.50 per 1k pages. Vs $4 from Mistral.
- kojoru 4mo agointeresting - an equivalent Azure Document Intelligence service (scanning with layout) is 10$/1k
- cvdub 4mo agoIt’s not the same service. Google’s vision OCR is pure text extraction, not layout. Pretty sure Google’s doc AI services that can identify headers vs body text is $10 per 1k pages.
- anon373839 4mo agoThat’s true, though worth beating a dead horse to say that traditional OCR won’t hallucinate sentences, perform unwanted translation, or change the meaning of whole paragraphs to something more “appropriate”.
- pmxi 4mo agoThis has been a niche where Mistral has actually been successful. Btw, Hindi and Japanese are bucketed in "Rare Languages," which is odd.
- ZiiS 4mo agoI read that as "languages under-represented in the training set".
- greenleafone7 4mo agoAfter paying for Mistral and using it for a while I genuinely hated it. It's a productivity black hole and can't realistically compete with anyone. I chose it only because it was European, but no. I'd rather let my one year subscription go to waste than use anything 'Mistral'.
- adlk 4mo agowhat did you use it for and when?
- amunozo 4mo agoSame, I got a refund 3 days later. It is unusable.
- maelito 4mo agoOpposite advice. It's very useful to me for dev and general tasks. Been using Claude in parallele, it's better not not that much, just 10x (or 100x ?) more expensive.
- greenleafone7 4mo agoSure, well for me it isn't. It has been awful for even toy tasks that opencode's free plan did without an issue. The general sentiment about it is that it is really bad. I wish I knew before paying.
- maelito 4mo agoWell, you lost 17 $. I'm not being sarcastic, that's not a lot for dev tasks.
- InsideOutSanta 4mo agoMistral's coding models aren't on par with current SOTA US and Chinese models if that's what you're referring to, but I rather like their OCR models.
- lxgr 4mo ago
- mcbetz 4mo agoLittle on differences other than bounding boxes and double the price compared to their previous OCR v3 model from December - https://mistral.ai/news/mistral-ocr-3/ https://mistral.ai/news/mistral-ocr-3/ - other benchmarks were used back then.
- MostlyStable 4mo agoDoes anyone know of OCR benchmarks that include hand-written documents? I'm currently using Gemini pro 3 for this, and error rates are quite good, but it's a little bit pricey, and I'd be interested in a cheaper model that could perform as well, but almost all the OCR benchmarks I'm aware of (and I believe all the ones included in this announcement) are about printed/typeset text.
- jimmypk 4mo ago[flagged]
- andrewmutz 4mo agoA tangential observation: the video on the linked page wasn't what I expected. I thought Mistral was a european AI company, so I didnt expect the video to be filmed in San Francisco featuring three people who don't seem to be european. I'm not against them being a global organization, that's wonderful. I was just surprised. I expected a parisian office and european accents.
- rjzzleep 4mo agoUnfortunately Europeans are terrible customers for making money. They ask a lot of questions and they're very stingy with their wallets. Americans on the other hand ...
- touwer 4mo agoYou're american?
- megous 4mo agoHe's a prince of whales.
- qingcharles 4mo agoIf he's a prince of whales, then I'm the king of eel-gland.
- throwa356262 4mo agoOh come on! Mistral has a successful business model and is actually making money. Not sure opening and anthropic are doing that yet.
- port11 4mo ago> Anthropic generated $4.8 billion in sales in the first quarter. Its quarterly revenue is now growing faster than Zoom did during the pandemic, and Google and Facebook in the run-up to their initial public offerings. It is set to turn an operating profit of $559 million in the June quarter.
- mrkn1 4mo agoThis runs for free on CPU https://github.com/kouhxp/textsnap https://github.com/kouhxp/textsnap
- coulix 4mo agoI wonder how it does compare to reducto, pulse, extendai.
- themanmaran 4mo agoIt's cheap at $4/1k, but I'm hesitant to even benchmark this one again since the previous versions were all "98% accurate based on internal benchmarks of 4 pdfs" and ended up falling short of almost everything else on the market [1]. Even in this one, they just report that OlmOCRBench and OmniDocBench have "known limitations" and that's why they report flagship numbers from their internal benchmark. https://getomni.ai/blog/benchmarking-open-source-models-for-ocr https://getomni.ai/blog/benchmarking-open-source-models-for-...
- v3ss0n 4mo agoNot opensource right?
- verdverm 4mo agoThe weights do not appear to be downloadable, "contact sales for self hosting"
- bastawhiz 4mo agoThe comparisons rank it against GPT and Gemini but not Claude. Is Claude's vision support simply not competitive when it comes to OCR tasks?
- abi 4mo agoI think until Fable, Claude's vision was significantly worse than GPT and Gemini in my personal experience. I eval almost every vision model since I work on screenshot to code conversion project: https://github.com/abi/screenshot-to-code https://github.com/abi/screenshot-to-code.
- Ninjinka 4mo agoIs there a complete list of the languages they support, and benchmarks by language, instead of just "Rare Languages"?
- dominotw 4mo agostarting y axis from 50 and 95 is a bit mileading
- sreekanth850 4mo agoTested with Malayalam, normal handwriting got accurate but a slight different style got detected as kannada. Have samples if required, which sarvam got done with 99% accuracy leaving one text error.
- civet_java 4mo agoI'm curious what's been your experience with Sarvam outside of Indic languages - Indian English (perhaps mixed with romanised indic verbiage) and also documents with complex layouts (figures, tables, etc). I've been quite curious but hesitant about Indian offerings, particularly because they seem to be priced a little higher than what I would think they should be (I could be wrong and simply be misrembering though).
- sreekanth850 4mo agoSarvam is exceptionally tuned for indic languages we have more than 20 languages and it perform well for all in ocr. Iam yet to test with other languages. No any models come close for indic languages like sarvam. I saw they recently dropped price per page to 0.5 inr which is much cheaper. The only downside is the zip file based delivery.
- deivid 4mo agoI am making (open) finetunes for malayalam and kannada (and bengali, gujarati, hebrew), and need someone to transcribe a few images for me. Could you contact me if you are interested in helping?
- sreekanth850 3mo agoI'm not an expert in fine tuning but i can obviously help if it falls under my expertise. You can shoot an email with my username and Gmail.
- JGB100 4mo agoNot well tested. It switched all U.S. (") double quotation marks to UK-style (') single quotation marks, ignoring the source document. Useless in the US.
- sscaryterry 4mo agoWhy the chart crimes?!
- beklein 4mo agoAll AI labs really need to stop using truncated y-axes for benchmark bar charts... https://mistral.ai/_astro/cm-engish_ZhlvoT.webp?dpl=6a3a94bd1f38530b2974c539 https://mistral.ai/_astro/cm-engish_ZhlvoT.webp?dpl=6a3a94bd...
- HDBaseT 4mo agoNo, they need to keep using truncated y-axes to increase the hype cycle.
- trilogic 4mo agoMistral keeps reminding us that doesn´t just brew great coffee, they can build great AI too. Hats off to the team. Mistral O.C.R. (Only Cool Results)
- ericyd 4mo agoI’ve always thought the US Postal Service is such a technological marvel. They somehow manage to identify and route billions of pieces of mail and I have to imagine their tech is significantly more primitive than this. Not only that but US addresses are absurdly non-standardized, you can often write the same address multiple ways and have it deliver to the same location. I’m sure there’s plenty of published knowledge in this area, but whenever I see announcements about OCR it feels like this should be a solved problem if it’s been accomplished at the scale of USPS for many years.
- alberth 4mo agoGreat video by Tom Scott on this subject: https://www.youtube.com/watch?v=XxCha4Kez9c https://www.youtube.com/watch?v=XxCha4Kez9c
- ericyd 4mo agohaha this was great!
- vel0city 4mo agoIIRC the USPS was one of the first big budget orgs behind early OCR systems all the way back in 1965. https://www.youtube.com/watch?v=V4LJs2ZoDR4 https://www.youtube.com/watch?v=V4LJs2ZoDR4
- thomasahle 4mo agoI used to part time for the (Danish) mail service. The only sorting that was done automatically was the post codes. That was enough to get the letter to the right post office. The rest was done by the mailmen/women early in the morning. It was a lot of fun trying to figure out what was meant by some of the addresses. The older people in particular often knew the story of why certain places were sometimes addressed in certain ways, or could guess the addresses based on the names of the people living there.
- adolph 4mo agoThe USPS Remote Encoding Center in Salt Lake City examined 841,260,847 images of poorly written addresses in fiscal year 2025. [0] Unfortunately the page does not have a base rate--the total number of mail pieces that were not prepared for automated processing. Total first class mail, which includes a lot of bills prepared for automation was 25.7 billion [1]. If 10% of that are non-automated, then .8 / 2.57 = .31 or a third of mail not prepared for automation is handled by "employees look at the image and type in address information" 0. https://facts.usps.com/remote-encoding-center-rec-deciphering-handwriting/ https://facts.usps.com/remote-encoding-center-rec-decipherin... 1. https://about.usps.com/what/financials/10k-reports/fy2025.pdf https://about.usps.com/what/financials/10k-reports/fy2025.pd...
- nickvec 4mo agoNaive question: is Claude no good at OCR? Was surprised to see that none of Anthropic's models were included in the benchmark comparisons.
- vasylvd 4mo ago[flagged]
- remus 4mo agoGiven this a test on some scans of magazines, generally pretty impressed with the results. Mags are generally pretty whacky layouts and it does a reasonable job working out what is where and pulling it together into a single coherent md file. The way it crops relevant pics and puts them into the doc is pretty nice. Haven't compared it with any other high tech OCR estups, but it's way better than the jank that comes as standard with my scanner.
- flakiness 4mo ago> On our internal multilingual evaluation, OCR 4 leads across all eight language groups — English, Western Europe, Eastern Europe, Middle Eastern, Chinese, East Asian, Southeast Asian, and specialized languages (Hindi, Japanese, Georgian, Bengali, Armenian, Hebrew, Greek, Gujarati, Tamil, Malayalam, Kannada, Telugu). The initial version of this page called these "minor languages" (vs specialized language), which is telling. If you're a speaker of one of these: This is why you need a sovereign set of models. (Japanese government: Are you listening?)
- deleted 4mo ago[deleted]
- zhivota 4mo agoAre there any open models focused on LPR (license plate recognition)? I have found some old ones but curious if there are new ones being developed like this OCR model. I may even try it for the purpose and see if it does well.
- dagni132 4mo ago[flagged]