13 ms·
Mistral OCR
- peterburkimsher 2y agoDoes it work for video subtitles? And in Chinese? I’m looking to transcribe subtitles of live music recordings from ANHOP and KHOP.
- Zopieux 2y agoSaving you a click: no, it cannot be self hosted (unless you have a few million dollars laying around)
- ChrisArchitect 2y ago[flagged]
- vessenes 2y agoNo comments there yet - this at the top of the home page, let’s use this one.
- hntiquated 2y ago[dead]
- ChrisArchitect 2y agoDon't need french /fr in the url. That is the one.
- vessenes 2y agoDang. Super fast and significantly more accurate than google, Claude and others. Pricing : $1/1000 pages, or per 2k pages if “batched”. I’m not sure what batching means in this case: multiple pdfs? Why not split them to halve the cost? Anyway this looks great at pdf to markdown.
- jacksnipe 2y agoI would assume this is 1 request containing 2k pages vs N requests whose total pages add up to 1000.
- abiraja 2y agoBatching likely means the response is not real-time. You set up a batch job and they send you the results later.
- sophiebits 2y agoBatched often means a higher latency option (minutes/hours instead of seconds), which providers can schedule more efficiently on their GPUs.
- Tostino 2y agoUsually (With OpenAI, I haven't checked Mistral yet) it means an async api rather than a sync api. e.g. you submit multiple requests (pdfs) in one call, and get back an id for the batch. You then can check on the status of that batch and get the results for everything when done. It lets them use their available hardware to it's full capacity much better.
- odiroot 2y agoMay I ask as a layperson, how would you about using this to OCR multiple hundreds of pages? I tried the chat but it pretty much stops after the 2nd page.
- 2y ago
- newfocogi 2y agoThey say: "releasing the API mistral-ocr-latest at 1000 pages / $" I had to reread that a few times. I assume this means 1000pg/$1 but I'm still not sure about it.
- bredren 2y agoYa, presumably it is missing the number `1.00`.
- svachalek 2y agoYeah you can read it as "pages per dollar" or as a unit "pages/$", it all comes out the same meaning.
- dgfl 2y agoGreat example of how information is sometimes compartmentalized arbitrarily in the brain: I imagine you have never been confused by sentences such as “I’m running at 10 km/h”.
- mkl 2y agoDollar signs go before the number, not after it like units. It needs to be 1000 pages/$1 to make sense, whereas 10km and 10h and 10/h all make sense so 10km/h does. I imagine you would be confused by km/h 10 but not $10.
- 2y ago
- pawelduda 2y agoIt outperforms the competition significantly AND can extract embedded images from the text. I really like LLMs for OCR more and more. Gemini was already pretty good at it
- sbarre 2y ago6 years ago I was working with a very large enterprise that was struggling to solve this problem, trying to scan millions of arbitrary forms and documents per month to clearly understand key points like account numbers, names and addresses, policy numbers, phone numbers, embedded images or scribbled notes, and also draw relationships between these values on a given form, or even across forms. I wasn't there to solve that specific problem but it was connected to what we were doing so it was fascinating to hear that team talk through all the things they'd tried, from brute-force training on templates (didn't scale as they had too many kinds of forms) to every vendor solution under the sun (none worked quite as advertised on their data).. I have to imagine this is a problem shared by so many companies.
- jcuenod 2y agoJust tested with a multilingual (bidi) English/Hebrew document. The Hebrew output had no correspondence to the text whatsoever (in context, there was an English translation, and the Hebrew produced was a back-translation of that). Their benchmark results are impressive, don't get me wrong. But I'm a little disappointed. I often read multilingual document scans in the humanities. Multilingual (and esp. bidi) OCR is challenging, and I'm always looking for a better solution for a side-project I'm working on (fixpdfs.com). Also, I thought OCR implied that you could get bounding boxes for text (and reconstruct a text layer on a scan, for example). Am I wrong, or is this term just overloaded, now?
- nicodjimenez 2y agoYou can get bounding boxes from our pdf api at Mathpix.com Disclaimer, I’m the founder
- kergonath 2y agoMathpix is ace. That’s the best results I got so far for scientific papers and reports. It understands the layout of complex documents very well, it’s quite impressive. Equations are perfect, figures extraction works well. There are a few annoying issues, but overall I am very happy with it.
- nicodjimenez 2y agoThanks for the kind words. What are some of the annoying issues?
- kergonath 2y agoI had a billing issue at the beginning. It was resolved very nicely but I try to be careful and I monitor the bill a bit more than I would like. Actually my main remaining technical issue is conversion to standard Markdown for use in a data processing pipeline that has issues with the Mathpix dialect. Ideally I’d do it on a computer that is airgaped for security reasons. But I haven’t found a very good way of doing it because the Python library wanted to check my API key. A problem I have and that is not really Mathpix’s fault is that I don’t really know how to store the figures pictures to keep them with the text in a convenient way. I haven’t found a very satisfying strategy. Anyway, keep up the good work!
- notepad0x90 2y agoI was just watching a science-related video containing math equations. I wondered how soon will I be able to ask the video player "What am I looking at here, describe the equations" and it will OCR the frames, analyze them and explain them to me. It's only a matter of time before "browsing" means navigating HTTP sites via LLM prompts. although, I think it is critical that LLM input should NOT be restricted to verbal cues. Not everyone is an extrovert that longs to hear the sound of their own voices. A lot of human communication is non-verbal. Once we get over the privacy implications (and I do believe this can only be done by worldwide legislative efforts), I can imagine looking at a "website" or video, and my expressions, mannerisms and gestures will be considered prompts. At least that is what I imagine the tech would evolve into in 5+ years.
- devmor 2y agoGood lord, I dearly hope not. That sounds like a coddled hellscape world, something you'd see made fun of in Disney's Wall-E.
- notepad0x90 2y agohence my comment about privacy and need for legislation :) It isn't the tech that's the problem but the people that will abuse it.
- devmor 2y agoWhile those are concerns, my point was that having everything on the internet navigated to, digested and explained to me sounds unpleasant and overall a drain on my ability to think and reason for myself. It is specifically how you describe using the tech that provokes a feeling of revulsion to me.
- notepad0x90 2y agoThen I think you misunderstand. The ML system would know when you want things digested to you or not. Right now companies are assuming this and forcing LLM interaction. But when properly done, the system would know based on your behavior or explicit prompts what you want and provide the service. If you're staring at a paragraph intently and confused, it might start highlighting common phrases or parts of the text/picture that might be hard to grasp and based on your reaction to that, it might start describing things via audio,tool tips,side pane,etc.. In other words, if you don't like how and when you're interacting with the LLM ecosystem, then that is an immature and failing ecosystem, in my vision this would be a largely solved problems, like how we interact with keyboards,mouse and touchscreens today.
- andoando 2y agoBit unrelated but is there anything that can help with really low resolution text? My neighbor got hit and run the other day for example, and I've been trying every tool I can to make out some of the letters/numbers on the plate https://ibb.co/mr8QSYnj https://ibb.co/mr8QSYnj
- busymom0 2y agoThere are photo enhancers online. But your picture is way too pixelated to get any useful info from it.
- tjoff 2y agoIf you know the font in advance (which you often do in these cases) you can do insane reconstructions. Also keep in mind that it doesn't have to be a perfect match, with the help of the color and other facts (such as likely location) about the car you can narrow it down significantly.
- zellyn 2y agoMaybe if you had multiple frames, and used something very clever?
- flutas 2y agoLooks like a paper temp tag. Other than that, I'm not sure much can be had from it.
- zinglersen 2y agoFinding the right subreddit and asking there is probably a better approach if you want to maximize the chances of getting the plate 'decrypted'.
- dewey 2y agoTo even get started on this you'd also need to share some contextual information like continent, country etc. I'd say.
- 2y ago
- TriangleEdge 2y agoOne of my hobby projects while in University was to do OCR on book scans. Doing character recognition was solved, but finding the relationship between characters was very difficult. I tried "primitive" neural nets, but edge cases would often break what I built. Super cool to me to see such an order of magnitude in improvement here. Does it do hand written notes and annotations? What about meta information like highlighting? I am also curious if LLMs will get better because more access to information if it can be effectively extracted from PDFs.
- jcuenod 2y ago* Character recognition on monolingual text in a narrow domain is solved
- jervant 2y agoI wonder how it compares to USPS workers at deciphering illegible handwriting.
- linklater12 2y agoDocument processing is where b2b SAAS is at.
- opwieurposiu 2y agoRelated, does anyone know of an app that can read gauges from an image and log the number to influx? I have a solar power meter in my crawlspace, it is inconvenient to go down there. I want to point an old phone at it and log it so I can check it easily. The gauge is digital and looks like this: https://www.pvh2o.com/solarShed/firstPower.jpg https://www.pvh2o.com/solarShed/firstPower.jpg
- dehrmann 2y agoYou'll be happier finding a replacement meter that has an interface to monitor it directly or a second meter. An old phone and OCR will be very brittle.
- haswell 2y agoNot OP, but it sounds like the kind of project I’d undertake. Happiness for me is about exploring the problem within constraints and the satisfaction of building the solution. Brittleness is often of less concern than the fun factor. And some kinds of brittleness can be managed/solved, which adds to the fun.
- arcfour 2y agoI would posit that learning how the device works, and how to integrate with a newer digital monitoring device would be just as interesting and less brittle.
- haswell 2y agoPossibly! But I’ve recently wanted to dabble with computer vision, so I’d be looking at a project like this as a way to scratch a specific itch. Again, not OP so I don’t know what their priorities are, but just offering one angle for why one might choose a less “optimal” approach.
- renewiltord 2y ago4o transcribes it perfectly. You can usually root an old Android and write this app in ~2h with LLMs if unfamiliar. The hard part will be maintaining camera lens cleanliness and alignment etc. The time cost is so low that you should give it a gander. You'll be surprised how fast you can do it. If you just take screenshots every minute it should suffice.
- ChemSpider 2y ago"World's best OCR model" - that is quite a statement. Are there any well-known benchmarks for OCR software?
- xnx 2y agohttps://huggingface.co/spaces/echo840/ocrbench-leaderboard https://huggingface.co/spaces/echo840/ocrbench-leaderboard
- ChemSpider 2y agoInteresting. But no mistral on it yet?
- themanmaran 2y agoWe published this benchmark the other week. We'll can update and run with Mistral today! https://github.com/getomni-ai/benchmark https://github.com/getomni-ai/benchmark
- kergonath 2y agoExcellent. I am looking forward to it.
- cdolan 2y agoCame here to see if you all had run a benchmark on it yet :)
- themanmaran 2y agoUpdate: Just ran our benchmark on the Mistral model and results are.. surprisingly bad? Mistral OCR: - 72.2% accuracy - $1/1000 pages - 5.42s / page Which is pretty far cry from the 95% accuracy they were advertising from their private benchmark. The biggest thing I noticed is how it skips anything it classifies as an image/figure. So charts, infographics, some tables, etc. all get lifted out and returned as [image](image_002). Compared to the other VLMs that are able to interpret those images into a text representation. https://github.com/getomni-ai/benchmark https://github.com/getomni-ai/benchmark https://huggingface.co/datasets/getomni-ai/ocr-benchmark https://huggingface.co/datasets/getomni-ai/ocr-benchmark https://getomni.ai/ocr-benchmark https://getomni.ai/ocr-benchmark
- deleted 2y ago[deleted]
- sashank_1509 2y agoReally cool, thanks Mistral!
- z2 2y agoIs there a reliable handwriting OCR benchmark out there (updated, not a blog post)? Despite the gains claimed for printed text, I found (anecdotally) that trying to use Mistral OCR on my messy cursive handwriting to be much less accurate than GPT-4o, in the ballpark of 30% wrong vs closer to 5% wrong for GPT-4o. Edit: answered in another post: https://huggingface.co/spaces/echo840/ocrbench-leaderboard https://huggingface.co/spaces/echo840/ocrbench-leaderboard
- dannyobrien 2y agoSimon Willison linked to an impressive demo of Qwen2-VL in this area: I haven't found a version of it that I could run locally yet to corroborate. https://simonwillison.net/2024/Sep/4/qwen2-vl/ https://simonwillison.net/2024/Sep/4/qwen2-vl/
- aperrien 2y agoIs this model open source?
- daemonologist 2y agoNo (nor is it open-weights).
- deleted 2y ago[deleted]
- cxie 2y agoThe new Mistral OCR release looks impressive - 94.89% overall accuracy and significantly better multilingual support than competitors. As someone who's built document processing systems at scale, I'm curious about the real-world implications. Has anyone tried this on specialized domains like medical or legal documents? The benchmarks are promising, but OCR has always faced challenges with domain-specific terminology and formatting. Also interesting to see the pricing model ($1/1000 pages) in a landscape where many expected this functionality to eventually be bundled into base LLM offerings. This feels like a trend where previously encapsulated capabilities are being unbundled into specialized APIs with separate pricing. I wonder if this is the beginning of the componentization of AI infrastructure - breaking monolithic models into specialized services that each do one thing extremely well.
- unboxingelf 2y agoWe’ll just stick LLM Gateway LLM in front of all the specialized LLMs. MicroLLMs Architecture.
- cxie 2y agoI actually think you're onto something there. The "MicroLLMs Architecture" could mirror how microservices revolutionized web architecture. Instead of one massive model trying to do everything, you'd have specialized models for OCR, code generation, image understanding, etc. Then a "router LLM" would direct queries to the appropriate specialized model and synthesize responses. The efficiency gains could be substantial - why run a 1T parameter model when your query just needs a lightweight OCR specialist? You could dynamically load only what you need. The challenge would be in the communication protocol between models and managing the complexity. We'd need something like a "prompt bus" for inter-model communication with standardized inputs/outputs. Has anyone here started building infrastructure for this kind of model orchestration yet? This feels like it could be the Kubernetes moment for AI systems.
- arcfour 2y agoThis is already done with agents. Some agents only have tools and the one model, some agents will orchestrate with other LLMs to handle more advanced use cases. It's pretty obvious solution when you think about how to get good performance out of a model on a complex task when useful context length is limited: just run multiple models with their own context and give them a supervisor model—just like how humans organize themselves in real life.
- alberth 2y agoCurious to see how this performance against more real world usage of someone taking a photo of text (which the text then becomes slightly blurred) and performing OCR on it. I can't exactly tell if the "Mistral 7B" image is an example of this exact scenario.
- roboben 2y agoLe chat doesn’t seem to know about this change despite the blog post stating it. Can anyone explain how to use it in Le Chat?
- kapitalx 2y agoLooks to be API only for now. Documentation here: https://docs.mistral.ai/capabilities/document/ https://docs.mistral.ai/capabilities/document/
- troyvit 2y agoI asked LeChat this question: If I upload a small PDF to you are you able to convert it to markdown? LeChat said yes and away we went.
- jacooper 2y agoPretty cool, would love to use this with paperless, but I just can't bring myself to send a photo of all my documents to a third party, especially legal and sensitive documents, which is what I use Paperless for. Because of that I'm stuck with crappy vision on Ollama (Thanks to AMDs crappy ROCm support for Vllm)
- jbverschoor 2y agoOhhh. Gonna test it out with some 100+ year old scribbles :)
- jbverschoor 2y agoIt did better than any other solution out there. However, I can only validate by the logic of the text. It's a recipe book.
- WhitneyLand 2y ago1. There’s no simple page / sandbox to upload images and try it. Fine, I’ll code it up. 2. “Explore the Mistral AI APIs” (https://docs.mistral.ai https://docs.mistral.ai) links to all apis except OCR. 3. The docs on the api params refer to document chunking and image chunking but no details on how their chunking works? So much unnecessary friction smh.
- cooperaustinj 2y agoThere is an OCR page on the link you provided. It includes a very, very simple curl command (like most of their docs). I think the friction here exists outside of Mistral's control.
- kergonath 2y ago> There is an OCR page on the link you provided. I don’t see it either. There might be some caching issue.
- WhitneyLand 2y agoHow is it out of their control to document what they mean by chunking in their parameters?
- deadbabe 2y agoLLM based OCR is a disaster, great potential for hallucinations and no estimate of confidence. Results might seem promising but you’ll always be wondering.
- menaerus 2y agoCNN-based OCR also have "hallucinations" and Transformers aren't that much different in that respect. This is a problem solved with domain specific post-processing.
- leumon 2y agowell already in 2013 ocr systems used in xerox scanners (turned on by default!) randomly altered numbers, so its not an issue only occuring in llms.
- utkarshphirke 2y agoAbsolutely right - we tried estimating LLM confidence and the results are not great. Any process that requires reliability will struggle with LLM OCR. https://news.ycombinator.com/item?id=43350816 https://news.ycombinator.com/item?id=43350816
- bob1029 2y ago> It takes images and PDFs as input If you are working with PDF, I would suggest a hybrid process. It is feasible to extract information with 100% accuracy from PDFs that were generated using the mappable acrofields approach. In many domains, you have a fixed set of forms you need to process and this can be leveraged to build a custom tool for extracting the data. Only if the PDFs are unknown or were created by way of a cellphone camera, multifunction office device, etc should you need to reach for OCR. The moment you need to use this kind of technology you are in a completely different regime of what the business will (should) tolerate.
- themanmaran 2y ago> Only if the PDFs are unknown or were created by way of a cellphone camera, multifunction office device, etc should you need to reach for OCR. It's always safer to OCR on every file. Sometimes you'll have a "clean" pdf that has a screenshot of an Excel table. Or a scanned image that has already been OCR'd by a lower quality tool (like the built in Adobe OCR). And if you rely on this you're going to get pretty unpredictable results. It's way easier (and more standardized) to run OCR on every file, rather than trying to guess at the contents based on the metadata.
- bob1029 2y agoIt's not guessing if the form is known and you can read the information directly. This is a common scenario at many banks. You can expect nearly perfect metadata for anything pushed into their document storage system within the last decade.
- themanmaran 2y agoOh yea if the form is known and standardized everything is a lot easier. But we work with banks on our side, and one of the most common scenarios is customers uploading financials/bills/statements from 1000's of different providers. In which case it's impossible to know every format in advance.
- SilentM68 2y agoI would like to see how it performs with massively warped and skewed scanned text images, basically a scanned image where the text lines are wavy as opposed as straight horizontal, where the letters are elongated. One where the line widths are different depending on the position on the scanned image. I once had to deal with such a task that somebody gave me with OCR software, Acrobat, and other tools could not decode the mess so I had to recreate the 30 pages myself, manually. Not a fun thing to do but that is a real use case.
- arcfour 2y agoGarbage in, garbage out?
- edude03 2y ago"Yes" but if a human could do it "AI" should be able to do it too.
- amelius 2y agoAre you trying to build a captcha solver?
- SilentM68 2y agoNo, not a captcha solver. When I worked in education, I was given a 90s paper document that a teacher needed OCRd but it was completely warped. It was my job to remediate those type of documents for Accessibility reasons. I had to scan and OCR it but the result was garbage. Mind you I had access to Windows, Linux and MacOS tools but still difficult to do. I had to guess what it said, which was not impossible but it was time-consuming, not doable in the time-frame I was given, so I had no option but to manually retype all the information into a new document and convert it that way. Document remediation and accessibility should be a good use case for A.I., in education.
- thegabriele 2y agoI use gemini to solve textual CAPTCHAS with those kind of distortions and more: 60% of the time it works every time.
- janalsncm 2y agoThe hard ones are things like contracts, leases, and financial documents which 1) don’t have a common format 2) are filled with numbers proper nouns and addresses which it’s really important not to mess up 3) cannot be inferred from context. Typical OCR pipeline would be to pass the doc through a character-level OCR system then correct errors with a statistical model like an LLM. An LLM can help correct “crodit card” to “credit card” but it cannot correct names or numbers. It’s really bad if it replaces a 7 with a 2.
- groby_b 2y agoPerusing the web site, it's depressing how much behind Mistral is on basic "how can I make this a compelling hook for customers" for the page. The notebook link? An ACL'd doc The examples don't even include a small text-to-markdown sample. The before/after slider is cute, but useless - SxS is a much better way to compare. Trying it in "Le Chat" requires a login. It's like an example of "how can we implement maximum loss across our entire funnel". (I have no doubt the underlying tech does well, but... damn, why do you make it so hard to actually see it, Mistral?) If anybody tried it and has shareable examples - can you post a link? Also, anybody tried it with handwriting yet?
- dehrmann 2y agoIs this burying the lede? OCR is a solved problem, but structuring document data from scans isn't.
- jslezak 2y agoHas anyone tried it for handwriting? So far Gemini is the only model I can get decent output from for a particular hard handwriting task
- pqdbr 2y agoI tried with both PDFs and PNGs in Le Chat and the results were the worst I've ever seen when compared to any other model (Claude, ChatGPT, Gemini). So bad that I think I need to enable the OCR function somehow, but couldn't find it.
- computergert 2y agoI'm experiencing the same. Maybe the sentence "Mistral OCR capabilities are free to try on le Chat." was a hallucination.
- troyvit 2y agoIt worked perfectly for me with a simple 2 page PDF that contained no graphics or formatting beyond headers and list items. Since it was so small I had the time to proof-read it and there were no errors. It added some formatting, such as bolding headers in list items and putting tics around file and function names. I won't complain.
- sunami-ai 2y agoMaking Transformers the same cost as CNN's (which are used in character-level ocr, as opposed to image-patch-level) is a good thing. The problem with CNN based character-level OCR is not the recognition models but the detection models. In a former life, I found a way to increase detection accuracy, and, therefore, overall OCR accuracy, and used that as an enhancement on top of Amazon and Google OCR. It worked really well. But the transformer approach is more powerful and if it can be done for $1 per 1000 pages, that is a game changer, IMO, at least of incumbents offering traditional character-level OCR.
- menaerus 2y agoIt certainly isn't the same cost if expressed as a non-subsidized $$$ one needs for the Transformers compute aka infra. CNNs trained specifically for OCR can run in real time on as small compute as a mobile device is.
- anon373839 2y agoA bit of a tangent, but aren’t CNNs still dominating over ViTs among computer vision competition winners?
- menaerus 2y agoI haven't watched that space very closely but IMO ViTs have a great potential to extract from since in comparison to CNNs they allow the model to learn and understand complex relations in the data. Where this matters, I expect it to matter a lot. OCR I think is not the greatest such example - while it matters to understand the surrounding context, I think it's not that critical for performance.
- srinathkrishna 2y agoGiven the fact that multi-modal LLMs are getting so good at OCR these days, is it a shame that we can't do local OCR with high accuracy in the near-term?
- coolspot 2y agoThis is $1 per 1000 pages. For comparison, Azure Document Intelligence is $1.5/1000 pages for general OCR and $30/1000 pages for “custom extraction”.
- kapitalx 2y agoCo-founder of doctly.ai here (OCR tool) I love mistral and what they do. I got really excited about this, but a little disappointed after my first few tests. I tried a complex table that we use as a first test of any new model, and Mistral OCR decided the entire table should just be extracted as an 'image' and returned this markdown: ```  ``` I'll keep testing, but so far, very disappointing :( This document I try is the entire reason we created Doctly to begin with. We needed an OCR tool for regulatory documents we use and nothing could really give us the right data. Doctly uses a judge, OCRs a document against multiple LLMs and decides which one to pick. It will continue to run the page until the judge scores above a certain score. I would have loved to add this into the judge list, but might have to skip it.
- infecto 2y agoWhy pay more for doctly than an AWS Textract?
- kapitalx 2y agoGreat question. The language models are definitely beating the old tools. Take a look at Gemini for example. Doctly runs a tournament style judge. It will run multiple generations across LLMs and pick the best one. Outperforming single generation and single model.
- nnurmanov 2y agoI did not try doctly, but AWS Textract does not support in my case Russian, so the output is completely useless
- the_mitsuhiko 2y agoWould love to see the test file.
- Starlord2048 2y agowould be glad to see benchmarking results
- owenpalmer 2y agoThis is incredibly exciting. I've been pondering/experimenting on a hobby project that makes reading papers and textbooks easier and more effective. Unfortunately the OCR and figure extraction technology just wasn't there yet. This is a game changer. Specifically, this allows you to associate figure references with the actual figure, which would allow me to build a UI that solves the annoying problem of looking for a referenced figure on another page, which breaks up the flow of reading. It also allows a clean conversion to HTML, so you can add cool functionality like clicking on unfamiliar words for definitions, or inserting LLM generated checkpoint questions to verify understanding. I would like to see if I can automatically integrate Andy Matuschak's Orbit[0] SRS into any PDF. Lots of potential here. [0] https://docs.withorbit.com/ https://docs.withorbit.com/
- generalizations 2y agoWait does this deal with images?
- ezfe 2y agoThe output includes images from the input. You can see that on one of the examples where a logo is cropped out of the source and included in the result.
- NalNezumi 2y ago>a UI that solves the annoying problem of looking for a referenced figure on another page, which breaks up the flow of reading. A tangent but this exact issue is what I was frustrated for a long time with pdf reader and reading science papers. Then I found sioyek that pops up a small window when you hover over links (references and equations and figures) and it solved it. Granted, the pdf file must be in right format, so OCR could make this experience better. Just saying the UI component of that already exist https://sioyek.info/ https://sioyek.info/
- PerryStyle 2y agoZotero's PDF viewer also does this now. Being able to annotate PDFs and having a reference manager has been a life saver.
- polytely 2y agoI don't need AGI just give me superhuman OCR so we can turn all existing pdfs into text* and cheaply host it. Feels like we are almost there. *: https://annas-archive.org/blog/critical-window.html https://annas-archive.org/blog/critical-window.html
- coolspot 2y agoThis is $1 per 1000 pages. For comparison, Azure Document Intelligence is $1.5/1000 pages for general OCR and $30/1000 pages for “custom extraction”.
- 0cf8612b2e1e 2y agoGiven the wide variety of pricing on all of these providers, I keep wondering how the economics work. Do they have fantastic margin on some of these products or is it a matter of subsidizing the costs, hoping to capture the market? Last I heard, OpenAI is still losing money.
- thegabriele 2y agoI'm using gemini to solve textual CAPTCHA with some good results (better than untrained OCR). I will give this a shot
- bugglebeetle 2y agoCongrats to Mistral for yet again releasing another closed source thing that costs more than running an open source equivalent: https://github.com/DS4SD/docling https://github.com/DS4SD/docling
- Squarex 2y agoI am all for open source, but where do you see benchmarks that conclude that it's just equivalent?
- bugglebeetle 2y agoWhere do you see open source benchmark results that confirm Mistral’s performance?
- anonymousd3vil 2y agoBack in my days Mistral used to torrent models.
- Asraelite 2y agoI never thought I'd see the day where technology finally advanced far enough that we can edit a PDF.
- randomNumber7 2y agoI never thought driving a car is harder than editing a pdf.
- pzo 2y agoIt's not about harder but about what error you can tolerate. Here if you have accuracy 99% for many applications it's enough. If you have 99% accuracy per trip of no crash during self driving then you gonna be dead within a year very likely. For cars we need accuracy at least 99.99% and that's very hard.
- rtsil 2y agoI doubt most people have 99% accuracy. The threshold of tolerance for error is just much lower for any self-driving system (and with good reason, because we're not familiar with them yet).
- KeplerBoy 2y agoHow do you define 99% accuracy? I guess something like success rate for a trip (or mile) would be a more reasonable metric. Most people have a success rate far higher than 99% for averages trips. Most people who commute daily are probably doing something like a 1000 car rides a year and have minor accidents every few years. 99% success rates would mean monthly accidents.
- lynx97 2y ago[dead]
- Apofis 2y agoFoxit PDF exists...
- thiago_fm 2y agoFor general use this will be good. But I bet that simple ML will lead to better OCRs when you are doing anything specialized, such as, medical documents, invoices etc.
- sureglymop 2y agoLooks good but in the first hover/slider demo one can see how it could lead to confusion when handling side by side content. Table 1 is referred to in section `2 Architectural details` but before `2.1 Multimodal Decoder`. In the generated markdown though it is below the latter section, as if it was in/part of that section. Of course I am nitpicking here but just the first thing I noticed.
- 0cf8612b2e1e 2y agoDoes anything handle dual columns well? Despite being the academic standard, it seemingly throws off every generic tool.
- serjester 2y agoThis is cool! With that said for anyone looking to use this in RAG, the downside to specialized models instead of general VLMs is you can't easily tune it to your use specific case. So for example, we use Gemini to add very specific alt text to images in the extracted Markdown. It's also 2 - 3X the cost of Gemini Flash - hopefully the increased performance is significant. Regardless excited to see more and more competition in the space. Wrote an article on it: https://www.sergey.fyi/articles/gemini-flash-2-tips https://www.sergey.fyi/articles/gemini-flash-2-tips
- deleted 2y ago[deleted]
- hyuuu 2y agogemini flash is notorious for hallucinating the output of the OCR, be careful with it. For straight forward, semi-structured, low page count (under 5) it should perform well, but the more the context window is stretched the more the output gets more unreliable
- oysterville 2y agoDupe of an hour previous post https://news.ycombinator.com/item?id=43282489 https://news.ycombinator.com/item?id=43282489
- beebaween 2y agoWonder how it does with table data in pdfs / page-long tabular data?
- blackeyeblitzar 2y agoA similar but different product that was discussed on HN is OlmOCR from AI2, which is open source: https://news.ycombinator.com/item?id=43174298 https://news.ycombinator.com/item?id=43174298
- hubraumhugo 2y agoIt will be interesting to see how all the companies in the document processing space adapt as OCR becomes a commodity. The best products will be defined by everything "non-AI", like UX, performance and reliability at scale, and human-in-the loop feedback for domain experts.
- trollied 2y agoThey will offer integrations into enterprise systems, just like they do today. Lots of big companies don't like change. The existing document processing companies will just silently start using this sort of service to up their game, and keep their existing relationships.
- hyuuu 2y agoI 100% agree with this, I think you can even extend this to any AI, in the end, IMO, as the llm is more commoditized, the surface of which the value is delivered will matter more
- lokl 2y agoTried with a few historical handwritten German documents, accuracy was abysmal.
- rvnx 2y agoProbably they are overfitting the benchmarks, since other users also complain of the low accuracy
- Thaxll 2y agoHTR ( Handwritten Text Recognition ) is a completely different space than OCR. What were you expecting exactly?
- riquito 2y agoIt fits the "use cases" mentioned in the article > Preserving historical and cultural heritage: Organizations and nonprofits that are custodians of heritage have been using Mistral OCR to digitize historical documents and artifacts, ensuring their preservation and making them accessible to a broader audience.
- anothermathbozo 2y agoOptical Character Recognition (OCR) and Handwritten Text Recognition (HTR) are different tasks
- 2y ago
- evmar 2y agoI noticed on the Arabic example they lost a space after the first letter on the third to last line, can any native speakers confirm? (I only know enough Arabic to ask dumb questions like this, curious to learn more.) Edit: it looks like they also added a vowel mark not present in the input on the line immediately after. Edit2: here's a picture of what I'm talking about, the before/after: https://ibb.co/v6xcPMHv https://ibb.co/v6xcPMHv
- resiros 2y agoArabic speaker here. No, it's perfect.
- evmar 2y agoI am pretty sure it added a kasrah not present in the input on the 2nd to last line. (Not saying it's not super impressive, and also that almost certainly is the right word, but I think that still means not quite "perfect"?)
- gl-prod 2y agoHe means the space between the wāw (و) and the word
- evmar 2y agoI added a pic to the original comment, sorry for not being clear!
- th0ma5 2y agoA great question for people wanting to use OCR in business is... Which digits in monetary amounts can you tolerate being incorrect?
- deleted 2y ago[deleted]
- kbyatnal 2y agoWe're approaching the point where OCR becomes "solved" — very exciting! Any legacy vendors providing pure OCR are going to get steamrolled by these VLMs. However IMO, there's still a large gap for businesses in going from raw OCR outputs —> document processing deployed in prod for mission-critical use cases. LLMs and VLMs aren't magic, and anyone who goes in expecting 100% automation is in for a surprise. You still need to build and label datasets, orchestrate pipelines (classify -> split -> extract), detect uncertainty and correct with human-in-the-loop, fine-tune, and a lot more. You can certainly get close to full automation over time, but it's going to take time and effort. But the future is on the horizon! Disclaimer: I started a LLM doc processing company to help companies solve problems in this space (https://extend.app/ https://extend.app/)
- risyachka 2y ago>> Any legacy vendors providing pure OCR are going to get steamrolled by these VLMs. -OR- they can just use these APIs, and considering that they have a client base - which would prefer to not rewrite integrations to get the same result - they can get rid of most code base, replace it with llm api and increase margins by 90% and enjoy good life.
- esafak 2y agoThey're going to become commoditized unless they add value elsewhere. Good news for customers.
- TeMPOraL 2y agoThey are (or at least could easily be) adding value in form of SLA - charging money for giving guarantees on accuracy. This is both better for customer, who gets concrete guarantees and someone to shift liability to, and for the vendor, that can focus on creating techniques and systems for getting that extra % of reliability out of the LLM OCR process. All of the above are things companies - particularly larger ones - are happy to pay for, because ORC is just a cog in the machine, and this makes it more reliable and predictable. On top of the above, there are auxiliary value-adds such a vendor could provide - such as, being fully compliant with every EU directive and regulation that's in power, or about to be. There's plenty of those, they overlap, and no one wants to deal with it if they can outsource it to someone who already figured it out. (And, again, will take the blame for fuckups. Being a liability sink is always a huge value-add, in any industry.)
- mvac 2y agoGreat progress, but unfortunately, for our use case (converting medical textbooks from PDF to MD), the results are not as good as those by MinerU/PDF-Extract-Kit [1]. Also the collab link in the article is broken, found a functional one [2] in the docs. [1] https://github.com/opendatalab/MinerU https://github.com/opendatalab/MinerU [2] https://colab.research.google.com/github/mistralai/cookbook/blob/main/mistral/ocr/structured_ocr.ipynb#scrollTo=svaJGBFlqm7_ https://colab.research.google.com/github/mistralai/cookbook/...
- owenpalmer 2y agoI've been searching relentlessly for something like this! I wonder why it's been so hard to find... is it the Chinese? In any case, thanks for sharing.
- thelittleone 2y agoHave you had a chance to compare results from MinerU vs LLM such a Gemini 2.0 or anthropic's native PDF tool?
- mvac 2y agoYes, i have. The problem with using just an LLM is that while it reads and understands text, but it cannot reproduce it accurately. Additionaly the textbooks I've mentioned have many diagrams and illustrations in them (e.g. books on anatomy or biochemistry). I don't really care about extracting text from them, I just need them extracted as images alongside the text, and no LLM does that.
- 101008 2y agoIs this free in LeChat? I uploaded a handwritten text and it stopped after the 4th word.
- bsnnkv 2y agoSomeone working there has good taste to include a Nizar Qabbani poem.
- rvz 2y ago> "Fastest in its category" Not one mention of the company that they have partnered with and that is Cerebras AI and that is the reason they have fast inference [0] Literally no-one here is talking about them and they are about to IPO. [0] https://cerebras.ai/blog/mistral-le-chat https://cerebras.ai/blog/mistral-le-chat
- pilooch 2y agoBut what's the need exactly for OCR when you have multimodal LLMs that can read the same info and directly answer any questions about it ? For a VLLM, my understanding is that OCR corresponds to a sub-field of questions, of the type 'read exactly what's written in this document'.
- daemonologist 2y agoIt's useful to have the plain text down the line for operations not involving a language model (e.g. search). Also if you have a bunch of prompts you want to run it's potentially cheaper, although perhaps less accurate, to run the OCR once and save yourself some tokens or even use a smaller model for subsequent prompts.
- ks2048 2y agoTons of uses: Storage (text instead of images), search (user typing in a text box and you want instant retrieval from a dataset), etc. And costs: run on images once - then the rest of your queries will only need to run on text.
- simonw 2y agoThe biggest risk of vision LLMs for OCR is that they might accidentally follow instructions is the text that they are meant to be processing. (I asked Mistral if their OCR system was vulnerable to this and they said "should be robust, but curious to see if you find any fun examples" - https://twitter.com/simonw/status/1897713755741368434 https://twitter.com/simonw/status/1897713755741368434 and https://twitter.com/sophiamyang/status/1897719199595720722 https://twitter.com/sophiamyang/status/1897719199595720722 )
- pilooch 2y agoFun, but LLMs would follow them post OCR anyways ;) I see OCR much like phonemes in speech, once you have end to end systems, they become latent constructs from the past. And that is actually good, more code going into models instead.
- troyvit 2y ago
- gatienboquet 2y agoI feel like i can't create an agent with their OCR model yet ? Is it something planned or it's only API?
- simonw 2y agoWhat do you mean by agent?
- gatienboquet 2y agoLa Plateforme agent builder - https://console.mistral.ai/build/agents/new https://console.mistral.ai/build/agents/new
- simonw 2y agoOh neat, thanks - I hadn't seen that. Looks like their version of an "agent" is a model with pre-baked system prompt and some examples.
- kiratp 2y agoIt's shocking how much our industry fails to see past its own nose. Not a single example on that page is a Purchase Order, Invoice etc. Not a single example shown is relevant to industry at scale.
- kashnote 2y agoFwiw, they have an example of a parking receipt in a cookbook: https://colab.research.google.com/github/mistralai/cookbook/blob/main/mistral/ocr/structured_ocr.ipynb https://colab.research.google.com/github/mistralai/cookbook/...
- guiomie 2y agoAgreed. In general I've had such bad performance for complex table based invoice parsing, that every few months I try the latest models to see if its better. It does say "96.12" on top-tier benchmark under the Table category.
- mtillman 2y agoWe find CV models to be better (higher midpoint on an ROC curve) for the types of docs you mention.
- simpaticoder 2y agoAnother good example would be contracts of any kind. Imagine photographing a contract (like a car loan) and on the spot getting an AI to read it, understand it, forecast scenarious, highlight red flags, and do some comparison shopping for you.
- JBiserkov 2y ago... imagining ... ... hallucinating during read ... ... hallucinating during understand ... ... hallucinating during forecast ... ... highlighting a hallucination as red flag ... ... missing an actual red flag ... ... consuming water to cool myself... Phew, being an AI is hard!
- simpaticoder 2y ago
- qwertox 2y agoWe developers seem to really dislike PDFs, to a degree that we'll build LLMs and have them translate it into Markdown. Jokes aside, PDFs really serve a good purpose, but getting data out of them is usually really hard. They should have something like an embedded Markdown version with a JSON structure describing the layout, so that machines can easily digest the data they contain.
- jgalt212 2y agoI think you might be looking for PDF/A. https://www.adobe.com/uk/acrobat/resources/document-files/pdf-types.html https://www.adobe.com/uk/acrobat/resources/document-files/pd... For example, if you print a word doc to PDF, you get the raw text in PDF form, not an image of the text.
- gpvos 2y agoPDF/A doesn't require preserving the document structure, only that any text is extractable.
- siva7 2y ago> We developers seem to really dislike PDFs, to a degree that we'll build LLMs and have them translate it into Markdown. Why Jokes aside? Markdown/html is better suited for the web than pdf
- d_llon 2y agoIt's disappointing to see that the benchmark results are so opaque. I hope we see reproducible results soon, and hopefully from Mistral themselves. 1. We don't know what the evaluation setup is. It's very possible that the ranking would be different with a bit of prompt engineering. 2. We don't know how large each dataset is (or even how the metrics are calculated/aggregated). The metrics are all reported as XY.ZW%, but it's very possible that the .ZW% -- or even Y.ZW% -- is just noise.[1] 3. We don't know how the datasets were mined or filtered. Mistral could have (even accidentally!) filtered out particularly data points that their model struggled with. (E.g., imagine good-meaning engineer testing a document with Mistral OCR first, finding it doesn't work, and deducing that it's probably bad data and removing it.) [1] https://medium.com/towards-data-science/digit-significance-in-machine-learning-dea05dd6b85b https://medium.com/towards-data-science/digit-significance-i...
- s4i 2y agoI wonder how good it would be to convert sheet music to MusicXML. All the current tools more or less suck with this task, or maybe I’m just ignorant and don’t know what lego bricks to put together.
- adrianh 2y agoTry our machine-learning powered sheet music scanning engine at Soundslice: https://www.soundslice.com/sheet-music-scanner/ https://www.soundslice.com/sheet-music-scanner/ Definitely doesn't suck.
- protonbob 2y agoWow this basically "solves" DRM for books as well as opening up the door for digitizing old texts more accurately.
- shmoogy 2y agoWhat's the general time for something like this to hit openrouter? I really hate having accounts everywhere when I'm trying to test new things.
- bondolo 2y agoSuch a shame that PDF doesn’t just, like, include the semantic structure of the document by default. It is brilliant that we standardized on an archival document format that doesn’t include direct access to the document text or structure as a core intrinsic default feature. I say this with great anger as someone who works in accessibility and has had PDF as a thorn in my side for 30 years.
- andai 2y agoTables? I regularly run into PDFs where even the body text is mangled!
- NeutralForest 2y agoI agree with this so much. I've tried to sometimes push friends and family to use text formats (at least I sent them something like Markdown), which is very easy to render in the browser anyways. But often you have to fall back to PDF, which I dislike very much. There's so much content like books and papers that are in PDF as well. Why did we pick a binary blob as shareable format again?
- meatmanek 2y ago> Why did we pick a binary blob as shareable format again? PDF was created to solve the problem of being able to render a document the same way on different computers, and it mostly achieved that goal. Editable formats like .doc, .html, .rtf were unreliable -- different software would produce different results, and even if two computers have the exact same version of Microsoft Word, they might render differently because they have different fonts available. PDFs embed the fonts needed for the document, and specify exactly where each character goes, so they're fully self-contained. After Acrobat Reader became free with version 2 in 1994, everybody with a computer ended up downloading it after running across a PDF they needed to view. As it became more common for people to be able to view PDFs, it became more convenient to produce PDFs when you needed everybody to be able to view your document consistently. Eventually, the ability to produce PDFs became free (with e.g. Office 2007 or Mac OS X's ability to print to PDF), which cemented PDF's popularity. Notably, the original goals of PDF had nothing to do with being able to copy text out of them -- the goal was simply to produce a perfect reproduction of the document on screen/paper. That wasn't enough of an inconvenience to prevent PDF from becoming popular. (Some people saw the inability for people to easily copy text from them as a benefit -- basically a weak form of text DRM.)
- OrvalWintermute 2y agoI'm happy to see this development after being underwhelmed with Chatgpt OCR!
- climb_stealth 2y agoDoes this support Japanese? They list a table of language comparisons againat other approaches but I can't tell if it is exhaustive. I'm hoping that something like this will be able to handle 3000-page Japanese car workshop manuals. Because traditional OCR really struggles with it. It has tables, graphics, text in graphics, the whole shebang.
- hyuuu 2y agoIt's a weird timing because I just launched https://dochq.io https://dochq.io - ai document extraction where you can define what you need to get out your documents in plain English, I legitimately thought that this was going to be such a niche product but hell, there has been a very rapid rise for AI-based OCR lately, an article/tweet even went viral 2 weeks ago I think? About using Gemini to do OCR, fun times.
- deleted 2y ago[deleted]
- sixhobbits 2y agoNice demos but I wonder how well it does on longer files. I've been experimenting with passing some fairly neat PDFs to various LLMs for data extraction. They're created from Excel exports and some of the data is cut off or badly laid out, but it's all digitally extractable. The challenge isn't so much the OCR part, but just the length. After one page the LLMs get "lazy" and just skip bits or stop entirely. And page by page isn't trivial as header rows are repeated or missing etc. So far my experience has definitely been that the last 2% of the content still takes the most time to accurately extract for large messy documents, and LLMs still don't seem to have a one-shot solve for that. Maybe this is it?
- hack_ml 2y agoYou will have to send one page at a time, most of this work has to be done via RAG. Adding a large context (like a whole PDF), still does not work that well in my experience.
- lysace 2y agoNit: Please change the URL from https://mistral.ai/fr/news/mistral-ocr https://mistral.ai/fr/news/mistral-ocr to https://mistral.ai/news/mistral-ocr https://mistral.ai/news/mistral-ocr The article is the same, but the site navigation is in English instead of French. Unless it's a silent statement, of course. =)
- lblume 2y agoFor me, the second page redirects to the first. (And I don't live in France.)
- anovick 2y agoHow does one use it to identify bounding rectangles of images/diagrams in the PDF?
- Oras 2y agoI feel this is created for RAG. I tried a document [0] that I tested with OCR; it got all the table values correctly, but the page's footer was missing. Headers and footers are a real pain with RAG applications, as they are not required, and most OCR or PDF parsers will return them, and there is extract work to do to remove them. [0] https://github.com/orasik/parsevision/blob/main/example/MultiPageInvoice.pdf https://github.com/orasik/parsevision/blob/main/example/Mult...
- deleted 2y ago[deleted]
- maCDzP 2y agoOh - on premise solution - awesome!
- cavisne 2y agoIts funny how Gemini consistently beats googles dedicated document API.
- jjice 2y agoI'm not surprised honestly - it's just the newer better things vs their older offering
- atemerev 2y agoSo, the only thing that stopped AI from learning from all our science and taking over the world was the difficulty of converting PDFs of academic papers to more computer readable formats. Not anymore.
- cytocync 2y ago[dead]
- noloz 2y agoAre there any open source projects with the same goal?
- submeta 2y agoIs this able to convert pdf flowcharts into yaml or json representations of them? I have been experimenting with Claude 3.5. It has been very good at readig / understanding/ converting into representations of flow charts. So I am wondering if this is more capable. Will try definitely, but maybe someone can chime in.
- bambax 2y agoIt's not bad! But it still hallucinates. Here's an example of an (admittedly difficult) image: https://i.imgur.com/jcwW5AG.jpeg https://i.imgur.com/jcwW5AG.jpeg For the blocks in the center, it outputs: > Claude, duc de Saint-Simon, pair et chevalier des ordres, gouverneur de Blaye, Senlis, etc., né le 16 août 1607 , 3 mai 1693 ; ép. 1○, le 26 septembre 1644, Diane - Henriette de Budos de Portes, morte le 2 décembre 1670; 2○, le 17 octobre 1672, Charlotte de l'Aubespine, morte le 6 octobre 1725. This is perfect! But then the next one: > Louis, commandeur de Malte, Louis de Fay Laurent bre 1644, Diane - Henriette de Budos de Portes, de Cressonsac. du Chastelet, mortilhomme aux gardes, 2 juin 1679. This is really bad because 1/ a portion of the text of the previous bloc is repeated 2/ a portion of the next bloc is imported here where it shouldn't be ("Cressonsac"), and of the right most bloc ("Chastelet") 3/ but worst of all, a whole word is invented, "mortilhomme" that appears nowhere in the original. (The word doesn't exist in French so in that case it would be easier to spot; but the risk is when words are invented, that do exist and "feel right" in the context.) (Correct text for the second bloc should be: > Louis, commandeur de Malte, capitaine aux gardes, 2 juin 1679.)
- bambax 2y agoAnother test with a text in English, which is maybe more fair (although Mistral is a French company ;-). This image is from Parliamentary debates of the parliament of New Zealand in 1854-55: https://i.imgur.com/1uVAWx9.png https://i.imgur.com/1uVAWx9.png Here's the output of the first paragraph, with mistakes in brackets: > drafts would be laid on the table, and a long discussion would ensue; whereas a Committee would be able to frame a document which, with perhaps a few verbal emundations [emendations], would be adopted; the time of the House would thus be saved, and its business expected [expedited]. With regard to the question of the comparative advantages of The-day [Tuesday]* and Friday, he should vote for the amendment, on the principle that the wishes of members from a distance should be considered on all sensations [occasions] where a principle would not be compromised or the convenience of the House interfered with. He hoped the honourable member for the Town of Christchurch would adopt the suggestion he (Mr. Forssith [Forsaith]) had thrown out and said [add] to his motion the names of a Committee.* Some mistakes are minor (emnundations/emendations or Forssith/Forsaith), but others are very bad, because they are unpredictable and don't correspond to any pattern, and therefore can be very hard to spot: sensations instead of occasions, or expected in lieu of expedited... That last one really changes the meaning of the sentence.
- neom 2y agoI gave it a bunch of my wifes 18th century English scans to transcribe, mostly couldn't do it, and it's been doing this for 15 minutes now, not sure why but i find quite amusing: https://share.zight.com/L1u2jZYl https://share.zight.com/L1u2jZYl
- dotnetkow 2y agoCongrats to the Mistral team for launching! A general-purpose OCR model is useful, of course. However, more purpose-built solutions are a must to convert business documents reliably. AI models pre-trained on specific document types perform better and are more accurate. Coming soon from the ABBYY team, we're shipping a new OCR API designed to be consistent, reliable, and hallucination-free. Check it out if you're looking for best-in-class DX: https://digital.abbyy.com/code-extract-automate-your-new-must-have-ocr-api-coming-soon https://digital.abbyy.com/code-extract-automate-your-new-mus...
- riffic 2y agoIt'd be great if this could be tested against genealogical documents written in cursive like oh most of the documents on microfilm stored by the LDS on familysearch, or eastern european archival projects etc.
- nyeah 2y agoIt's not fair to call it a "Mistrial" just because it hallucinates a little bit.
- vikp 2y agoI ran a partial benchmark against marker - https://github.com/VikParuchuri/marker https://github.com/VikParuchuri/marker . Across 375 samples with LLM as a judge, mistral scores 4.32, and marker 4.41 . Marker can inference between 20 and 120 pages per second on an H100. You can see the samples here - https://huggingface.co/datasets/datalab-to/marker_comparison_mistral_llm https://huggingface.co/datasets/datalab-to/marker_comparison... . The code for the benchmark is here - https://github.com/VikParuchuri/marker/tree/master/benchmarks https://github.com/VikParuchuri/marker/tree/master/benchmark... . Will run a full benchmark soon. Mistral OCR is an impressive model, but OCR is a hard problem, and there is a significant risk of hallucinations/missing text with LLMs.
- carlgreene 2y agoThank you for your work on Marker. It is the best OCR for PDFs I’ve found. The markdown conversion can get wonky with tables, but it still does better than anything else I’ve tried
- vikp 2y agoThanks for sharing! I'm training some models now that will hopefully improve this and more :)
- lolinder 2y ago> with LLM as a judge For anyone else interested, prompt is here [0]. The model used was gemini-2.0-flash-001. Benchmarks are hard, and I understand the appeal of having something that seems vaguely deterministic rather than having a human in the loop, but I have a very hard time accepting any LLM-judged benchmarks at face value. This is doubly true when we're talking about something like OCR which, as you say, is a very hard problem for computers of any sort. I'm assuming you've given this some thought—how did you arrive at using an LLM to benchmark OCR vs other LLMs? What limitations with your benchmark have you seen/are you aware of? [0] https://github.com/VikParuchuri/marker/blob/master/benchmarks/overall/scorers/llm.py https://github.com/VikParuchuri/marker/blob/master/benchmark...
- 2y ago
- rjurney 2y agoWhat about tables in PDFs?
- fsfsdfads 2y ago[dead]
- low_tech_punk 2y agoThis might be a contrarian take: the improvement against gpt-4o and gemini-1.5 flash, both of which are general purpose multi-modal models, seem to be underwhelming. I'm sensing another bitter lesson coming, where domain optimized AI will hold a short term advantage but will be outdated quickly as the frontier model advances.
- joeevans1000 2y agoCan someone give me a tl&dr on how to start using this? Is this available if one signs up for a regular Mistral account?
- ritvikpandey21 2y agoas builders in this space, we decided to put it to the test on complex nested tables, pie charts, etc. to see if the same VLM hallucination issues persist, and to what degree. while results were promising, we found several critical failure nodes across two document domains. check out our blog post here! https://www.runpulse.com/blog/beyond-the-hype-real-world-tests-of-mistrals-ocr https://www.runpulse.com/blog/beyond-the-hype-real-world-tes...
- Gnan 2y agoIs there an ocr with this kind of accuracy, but can run in a mobile device ? Looking for an ocr that can detect texts with high accuracy in realtime, so option of using cloud ocr is not viable.
- joeevans1000 2y agoI've found that the stunning OCR results so far were because the models were trained on the example file category. Is that the case here? Or can this recognize various documents?
- noosphr 2y agoAfter four years of "This VLM will solve OCR forever!" posts I've firmly put VLMs in the "useless until 100t parameters" category. Some day when there is enough internal state and training data that they can recognize tables, images and text we will get a gpt3 like moment that will make regular OCR obsolete. But that day is very far off and everyone who I've talked with and consulted over using VLMs in their pipeline is better served doing something else specific to their use case. An example I've used before: https://imgur.com/a/w2XKObh https://imgur.com/a/w2XKObh The reply: --- To extract the text contained within each box while ignoring the text in sub-boxes, we can follow these steps: 1. Identify the outermost box and extract its text. 2. Move to the next outermost box and extract its text, ignoring any text within its sub-boxes. 3. Continue this process for all boxes. Let's apply this to the image provided: 1. The outermost box contains the text: "This should be second". 2. The next outermost box (ignoring the sub-box within it) contains the text: "First". 3. The next box contains the text: "And also this". 4. The final box contains the text: "The quick brown fox". So, the extracted text from each box, ignoring sub-boxes, is: 1. "This should be second" 2. "First" 3. "And also this" 4. "The quick brown fox" --- As you can plainly see it is _wildly_ wrong and gives you no way to try and recover from those errors.
- applgo443 2y agoWhat's the simple explanation for why these VLM OCRs hallucinate but previous version of OCRs don't?
- prats226 2y agoTraditional OCR's usually have detection + recognition pipeline. So they will detect every word and then try to predict the text for every word. Errors obviously can happen in both parts, eg some words not detected which will get missed from output. Or word recognized incorrectly which is also common and more comparable to hallucination. However give its trained to work only on a small patch, accuracy is often higher. Comparing this to VLM's, they are looking at entire image/context and auto-regressively generating tokens/text which can also have lot of language bias, hence hallicinations.
- jwr 2y agoAlas, I can't run it locally. So it still doesn't solve the problem of OCR for my PDF archive containing my private data...
- porphyra 2y agoI uploaded a picture of my Chinese mouthwash [0] and it made a ton of mistakes and hallucinated a lot. Very disappointing. For example it says the usage instructions is to use 80 ml each time, even though the actual usage instruction on the bottle says use 5-20 mL each time, three times a day, and gargle for 1 minute. [0] https://i.imgur.com/JiX9joY.jpeg https://i.imgur.com/JiX9joY.jpeg [1] https://chat.mistral.ai/chat/8df2c9b9-ee72-414b-81c3-843ce74e1965 https://chat.mistral.ai/chat/8df2c9b9-ee72-414b-81c3-843ce74...
- yoelhacks 2y agoI was curious about Mistral so I made a few visualizations. A high level diagram w/ links to files: https://eraser.io/git-diagrammer?diagramId=uttKbhgCgmbmLp8OFf9R https://eraser.io/git-diagrammer?diagramId=uttKbhgCgmbmLp8OF... Specific flow of an OCR request: https://eraser.io/git-diagrammer?diagramId=CX46d1Jy5Gsg3QDzPOah https://eraser.io/git-diagrammer?diagramId=CX46d1Jy5Gsg3QDzP... (Disclaimer - uses a tool I've been working on)
- constantinum 2y agoI see a lot of comments on hallucination risk and the accumulation of non-traceable rotten data. If you are curious to try a better non-llm-based OCR, try LLMWhisperer.https://pg.llmwhisperer.unstract.com/ https://pg.llmwhisperer.unstract.com/
- shekhargulati 2y agoMistral OCR made multiple mistakes in extracting this [1] document. It is a two-page-long PDF in Arabic from the Saudi Central Bank. The following errors were observed: - Referenced Vision 2030 as Vision 2.0. - Failed to extract the table; instead, it hallucinated and extracted the text in a different format. - Failed to extract the number and date of the circular. I tested the same document with ChatGPT, Claude, Grok, and Gemini. Only Claude 3.7 extracted the complete document, while all others failed badly. You can read my analysis here [2]. 1. https://rulebook.sama.gov.sa/sites/default/files/en_net_file_store/SAMA_EN_10395_VER1.pdf https://rulebook.sama.gov.sa/sites/default/files/en_net_file... 2. https://shekhargulati.com/2025/03/05/claude-3-7-sonnet-is-good-at-pdf-processing/ https://shekhargulati.com/2025/03/05/claude-3-7-sonnet-is-go...
- michaelbuckbee 2y agoI'd mentioned this on HN last month, but I took a picture of a grocery list and then pasted it into ChatGPT to have it written out and it worked flawlessly...until I discovered that I'd messed up the picture when I took it at an angle and had accidentally cut off the first character or two of the bottom half of the list. ChatGPT just inferred that I wanted the actual full names of the items (aka "flour" instead of "our"). Depending on how you feel about it, this is either an absolute failure of OCR or wildly useful and much better.
- kccqzy 2y agoI have an actually hard OCR exercise for an AI model: I take this image of Chinese text on one of the memorial stones on the Washington Monument https://www.nps.gov/articles/american-mission-ningpo-china-220-level.htm https://www.nps.gov/articles/american-mission-ningpo-china-2... and ask the model to do OCR. Not a single model I've seen can OCR this correctly. Mistral is especially bad here: it gets stuck in an endless loop of nonsensical hallucinated text. Insofar as Mistral is design for "preserving historical and cultural heritage" it couldn't do that very well yet. A good model can recognize that the text is written top to bottom and then right to left and perform OCR in that direction. Apple's Live Text can do that, though it makes plenty of mistakes otherwise. Mistral is far from that.
- simonw 2y agoI built a CLI script for feeding PDFs into this API - notes on that and my explorations of Mistral OCR here: https://simonwillison.net/2025/Mar/7/mistral-ocr/ https://simonwillison.net/2025/Mar/7/mistral-ocr/
- raunakchowdhuri 2y agoWe ran some benchmarks comparing against Gemini Flash 2.0. You can find the full writeup here: https://reducto.ai/blog/lvm-ocr-accuracy-mistral-gemini https://reducto.ai/blog/lvm-ocr-accuracy-mistral-gemini A high level summary is that while this is an impressive model, it underperforms even current SOTA VLMs on document parsing and has a tendency to hallucinate with OCR, table structure, and drop content.
- hackernewds 2y agomeanwhile, you're comparing it to the output of almost a trillion dollar company
- HaZeust 2y ago... And? We're judging it for the merits of the technology it purports to be, not the pockets of the people that bankroll them. Probably not fair - sure, but when I pick my OCR, I want to pick SOTA. These comparisons and announcements help me find those.
- stann 2y agoThe tagline boasts that it is "introducing the world’s best document understanding API". So, holding them to their marketing seems fair
- raunakchowdhuri 2y agocomparisons to more outputs coming soon!
- zelcon 2y agoRelease the weights or buy an ad
- revskill 2y agoNextjs error is still uncauht correctly.
- hdjrudni 2y agoStill terrible at handwriting. I signed up for the API, cobbled together from their tutorial (https://docs.mistral.ai/capabilities/document/ https://docs.mistral.ai/capabilities/document/) -- why can't they give the full script instead of little bits? Tried uploading a tiff, they rejected it. Tried upload JPG, they rejected it (even though they supposed support images?). Tried resaving as PDF. It took that, but the output was just bad. Then tried ChatGPT on the original .tiff (not using API), and it got it perfectly. Honestly I could barely make out the handwriting with my eyes but now that I see ChatGPT's version I think it's right.
- InvidFlower 2y agoIt is confusing, but they have diff calls for pdfs vs images. In their example google colab: https://colab.research.google.com/drive/11NdqWVwC_TtJyKT6cmuap4l9SryAeeVt https://colab.research.google.com/drive/11NdqWVwC_TtJyKT6cmu... The first couple of sections are for pdfs and you need to skip all that (search for "And Image files...") to find the image extraction portion. Basically it needs ImageURLChunk instead of DocumentURLChunk.
- t_sea 2y agoThey really went for it with the hieroglyphs opening.
- monkeydust 2y agoSpent time working on OCR problem many years ago for a mobile app. We found at the time that the preprocessing was so critical to the outcome (quality of image, angle, colour/greyscale)
- sireat 2y agoIntriguing announcement, however the examples on the mistral.ai page seem rather "easy". What about rare glyphs in different languages using handwriting from previous centuries? I've been dealing with OCR issues and evaluating different approaches for past 5+ years at a national library that I work at. Usual consensus is that widely used open source Tesseract is subpar to commercial models. That might be so without fine tuning. However one can perform supplemental training and build your own Tesseract models that can outperform the base ones. Case study of Kant's letter's from 18th century: About 6 months ago, I tested OpenAi approach to OCR to some old 18th century letters that needed digitizing. The results were rather good (90+% accuracy) with the usual hallucination here and there. What was funny that OpenAI was using base Tesseract to generate the segmenting and initial OCR. The actual OCRed content before last inference step was rather horrid because the Tesseract model that OpenAi was using was not appropriate for the particular image. When I took OpenAi off the first step and moved to my own Tesseract models, I gained significantly in "raw" OCR accuracy at character level. Then I performed normal LLM inference at the last step. What was a bit shocking: My actual gains for the task (humanly readable text for general use) were not particularly significant. That is LLMs are fantastic at "untangling" complete mess of tokens into something humanly readable. For example: P!3goattie -> prerogative (that is given the surrounding text is similarly garbled)
- raffraffraff 2y agoForgive my absolute ignorance, I should probably run this through a chat bot before posting ... So I'm updating my post with answers now! Q: Do LLMs specialise in "document level" recognition based on headings, paragraphs, columns tables etc? Ie: ignore words and characters for now and attempt to recognise a known document format. A: Not most LLMs, but those with multimodal / vision capability could (eg DeepSeek Vision. ChatGPT 4). There are specialized models for this work like Tesseract, LayoutLM. Q: How did OCR work "back in the day" before we had these LLMs? Are any of these methods useful now? A: They used pattern recognition and feature extraction, rules and templates. Newer ML based OCR used SVM to isolate individual characters and HMM to predict the next character or word. Today's multimodal models process images and words, can handle context better than the older methods, and can recognise whole words or phrases instead of having to read each character perfectly. This is why they can produce better results but with hallucinations. Q: Can LLMs rate their own confidence in each section, maybe outputting text with annotations that say "only 10% certain of this word", and pass the surrounding block through more filters, different LLMs, different methods to try to improve that confidence? A: Short answer, "no". But you can try to estimate with post processing. Or am I super naive, and all of those methods are already used by the big commercial OCR services like Textract etc?
- utkarshphirke 2y agoLLMs are quite poor at rating their own confidence. Your best bet is to train a task specific LLM and ensure it is not overfit We benchmarked it here - https://news.ycombinator.com/item?id=43350816 https://news.ycombinator.com/item?id=43350816
- kinnth 2y agoThis looks like a massive win if you were the NHS and had to scan and process old case notes. Same is true if you were a solicitors/lawyers.
- lynx97 2y ago[dead]
- dwedge 2y agoBenchmarks look good. I tried this with a PDF that already has accurate PDF embedded just with new lines making pdftotext fail, and it was accurate for the text it found, but missed entire pages
- deleted 2y ago[deleted]
- mjnews 2y ago[dead]
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- mjnews 2y ago[dead]
- egorfine 2y agoI had a need to scan serial numbers from Apple's product boxes out of pictures taken by a random person on their phone. All OCR tools that I have tried have failed. Granted, I would get much better results if I used OpenCV to detect the label, rotate/correct it, normalize contrast, etc. But... I have tried the then new vision model from OpenAI and it did the trick so well it's wasn't feasible to consider anything else at that point. I have checked all S/N afterwards for being correct via third-party API - and all of theme were. Sure, sometimes I had to check versions with 0/o and i/l/1 substitutions but I believe these kind of mistakes are non-issues.
- lingjiekong 2y agoCurious that have people find more details regarding what is the architecture of this "mistral-ocr-latest". I have two question that 1. I was initially thinking this is VLM parsing model until I saw it can extract images. Then, I assume it is a pipeline of an image extraction and a VLM model while their result is combined to give the final result. 2. In this case, benchmark the pipeline result vs a end to end VLM such as gemini 2.0 flash might not be apple to apple comparison.
- yoeven 2y agoI ran Mistral AI OCR against JigsawStack OCR and beat their model in every category. Full breakdown here: https://jigsawstack.com/blog/mistral-ocr-vs-jigsawstack-vocr https://jigsawstack.com/blog/mistral-ocr-vs-jigsawstack-vocr
- 27theo 2y agoJust a small fyi, as viewed on an iPhone in Safari your tables don’t allow horizontal scrolling, cutting off the right column
- soyyo 2y agoI understand that is more juicy to get information from graphs, figures and so on, as every domain uses those, but i really hope to eventually see these models to be able to workout music notation, i have tried the best known apps and all of them fail to capture important details such as guitar performace symbols for bends or legato
- InvidFlower 2y agoWhile it is nice to have more options, it still definitely isn't at a human level yet for hard to read text. Still haven't seen anything that can deal with something like this very well: https://i.imgur.com/n2sBFdJ.jpeg https://i.imgur.com/n2sBFdJ.jpeg If I remember right, Gemini actually was the closest as far as accuracy of the parts where it "behaved", but it'd start to go off the rails and reword things at the end of larger paragraphs. Maybe if the image was broken up into smaller chunks. In comparison, Mistral for the most part (besides on one particular line for some reason) sticks to the same number of words, but gets a lot wrong on the specifics.
- thomasahle 2y agoI'm surprised they didn't benchmark it against Pixtral. They test it against a bunch of different Multimodal LLMs, so why not their own? I don't really see the purpose of the OCR form factor, when you have multimodal LLMs. Unless it's significantly cheaper.
- strangescript 2y agoI think its interesting they left out Gemini 2.0 Pro in the benchmarks which I find to be markedly better than flash if you don't mind the spend.
- jhatemyjob 2y agoAs far as open source OCRs go, Tesseract is still the best, right?
- ein0p 2y agoCould anyone suggest a tool which would take a bunch of PDFs (already OCR-d with Finereader), and replace the OCR overlay on all of them, maintaining the positions? I would like to have more accurate search over my document archive.
- f_k 2y agohttps://getsearchablepdf.com https://getsearchablepdf.com (I'm the founder)
- meepmeepinator 2y ago[flagged]
- Zufriedenheit 2y agoHow can I use these new OCR tools to make PDF files searchable by embedding the text layer?
- f_k 2y agoShameless plug: https://getsearchablepdf.com https://getsearchablepdf.com
- deleted 2y ago[deleted]
- Creator23 2y ago[dead]
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- Creator56 2y ago[dead]
- jojogh 2y agoHigh accuracy is the goal! But the multimodal approach introduces some complexities that can impact real-world performance. We break it down in our review: https://undatas.io/blog/posts/in-depth-review-of-mistral-ocr-a-pdf-parsing-powerhouse-tailored-for-the-ai-era/ https://undatas.io/blog/posts/in-depth-review-of-mistral-ocr... As for use cases, it really depends on how well it handles edge cases…
- ronald263 2y ago[dead]