8 ms·
Mistral OCR 4.1
- mainecoder 2mo agoThe chinese did it better, mistral is alive thanks to regulations.
- rtaylorgarlock 2mo agoI've been a bit more careful about complaining about regulations broadly due to competitive advantage, e.g. ITAR
- mangecoeur 2mo agoi.e. it's one AI company that's basically guaranteed to never fail since it has a market niche guaranteed by European companies and governments.
- petcat 2mo agoWhich is also why their most recent model "Shieldstral" does nothing except monitor and moderate internet content. After stuff like Chat Control I think they're obviously seeing a big demand for this kind of "internet safety" technology in Europe.
- gregorygoc 2mo agoI don’t think any European AI model can reach the experience and productivity of average NSA/CIA analyst.
- maelito 2mo agoNo it's the other way round : the Chinese do better thanks to regulations : massive amounts of money from Big tech and public money.
- Bombthecat 2mo agoYeah, I'm not sending personal bills etc to china. No thanks
- gkbrk 2mo agoIt's still open-weight models that you can download and run locally. You don't need to send anything to China.
- sajithdilshan 2mo agoand you trust the french?
- MrDrMcCoy 2mo agoI trust them more than the Americans and Chinese. Genuinely rooting for them to close the gap.
- sajithdilshan 2mo agoGood luck with that. Mistral is never going to produce a frontier model. Their business model is to rely on the EU regulations and cater to the bureaucracy. I won't be surprised if their next model is specialized on generating laws/regulations for EU lawmakers so they can sell it to EU. That way their business would survive for the next 5-10 years till EU implodes.
- MrDrMcCoy 2mo agoYour disdain for the EU is genuinely puzzling to me.
- rdm_blackhole 2mo agoWhy is someone critiquing the EU such an affront to you? Is the EU above reproach for some reason?
- t3hTao 2mo ago[dead]
- eloisant 2mo agoIt's not regulations, it's European companies being (legitimately) worried about their data sovereignty if they use US or Chinese AI providers.
- king_crimson 2mo agoAt this point I lost all hope for Europe playing any significant role in the AI race. If that’s a good or a bad thing I don’t know, but it seems to me like that’s the reality.
- kubb 2mo agoIt's not a race. You don't get anything for winning.
- bpodgursky 2mo agoIt's red queen. You stay alive by winning, you lose everything by losing.
- ben_w 2mo agoNot sure you get either outcome in either case. Race dynamics increases p(doom) for everyone. The non-doom scenarios include "utopia for all", and "power flows to investors, not citizens of whichever nation the winning model's corp. was registered in". Independently, "oh look all the investors went bankrupt" can happen in both "doom" and "normal technology" timelines.
- procgen 2mo agoThe only prize is control of the light cone.
- ChrisClark 2mo agoUnless you manage to build a god, and keep it under control... okay, we're all going to lose
- Palpatineli 2mo agoHow about "being able to align ASI somewhat to your values"?
- ben_w 2mo ago
- merb 2mo ago1000 Pages / 3.5€ this is expensive as hell. If this is not fastly superior than something like tesseract it is not worth it.
- beernet 2mo agoAgreed. Does the GTM team there really sit together like "oh yeah, that sounds reasonable" while being totally beyond typical market prices?
- Oras 2mo agoEven comparing to AWS Textract or Azure Document Intelligence, this is very expensive (more than double)
- ttyyzz 2mo agoHow does it compare to those services? Are there any benchmarks yet?
- coredog64 2mo agoAnecdotal as I haven't had time for a full benchmark, but I have a dense tabular handwritten form that is pretty challenging: LLMs with vision don't do well because the content isn't English words. Textract did poorly because it was just too dense for their model (although I haven't tried recently) Whereas I tried this model and it was a champ: Reasonable Markdown format with accurate content.
- Oras 2mo agofaced something similar with Azure, try upscaling the page (2x or more) before sending to the OCR, it increased the accuracy a lot for tiny tables on landscape pages.
- ComputerPerson 2mo agoI've got a scan from a book that I OCR with new releases. Ligatures, critical sigla, Fraktur letterforms, subscripts, superscripts, etc. Nothing special about this model for overly-detailed work like mine. It's been a while since I last tested (and discontinued my subscription), but the "pro" models from OpenAI dominate. Not surprising, given the price difference, but it would be nice if an OCR-specific model could perform better. It's worth mentioning that even the highest-end models do a pretty poor job with intricate text like mine.
- petcat 2mo ago> the "pro" models from OpenAI dominate. Not surprising considering the price difference, but it would ne nice if an OCR-specific model could do better. I haven't been impressed with any of Mistral's models. They obviously realized that they couldn't compete at the frontier so they decided to go for smaller focused models but even those have not been that good.
- booi 2mo agoThis is what I found as well. We moved away from Cursor but I was looking for a model that would help with FIM (fill-in-middle) multiline autocompletion and people were recommending Mistral's Codestral. We gave it a shot and it was lackluster at best.. Even Google's Gemini did a significantly better job than Codestral. Ultimately Opus-class models got good enough and I don't do much manual coding anymore.
- rtaylorgarlock 2mo agoYet: how is pricing? Evaluating contents and routing appropriately isn't a new challenge in OCR, one of the oldest fields of applications in ML. Thus, how do the smaller open models perform in tandem with relatively pricy $/pg models & APIs? Your use case is remarkably rare relative to the volume and price sensitivity of enterprise data warehouse ops.
- deleted 2mo ago[deleted]
- ianhawes 2mo agoI won't comment on accuracy, but in internal benchmarks, Mistral OCR is significantly faster than comparable APIs.
- Johnny_Bonk 2mo agoHow does this compare to Baidu Unlimited OCR. I've been very impressed with Baidu and it's essentially free to run on a decent computer, other than electricity costs.
- spiderfarmer 2mo agoWhere do your documents go?
- rescbr 2mo agoThey go to the decent computer hosting the model, which can be yours if you pay the electricity costs
- maz1b 2mo agoHow does this compare to 4?
- piterrro 2mo agoFor anyone interested, I have an ocr pipeline running on rented GPUs, doing around 1000pages for 0.05-01 usd with around 0.8 seconds per page with full bounding boxes support for grounding. If you’re interested you can find contact to me via this profile. 3.5 usd/1000 pages is just too expensive…
- x3ro 2mo agoYou should put contact details in your profile :)
- aliljet 2mo agoAccuracy is truly what people die for in the OCR game. Price isn't the primary function here.. it's an equation of price, accuracy, speed, and in mayn cases regulation.
- merb 2mo agoTbf even with tesseract you already get shit ton of accuracy and you can probably do these 1000 pages for way less than 3.5€. For 3.5€ you can spin up a cloud instance with 8vCPU+32gb on gcloud for 11 hours (or 11 instances for an hour) which can do way more than 1000 pages per hour on tesseract. It takes you around 6 second per page +-4 seconds start/stop depending on what you are doing on that instance size without too much optimization (you can probably even run multiple processes on a single node) Google documentai costs 1.5$ per 1000 which is probably better in quality and speed.
- kergonath 2mo agoTesseract is not a substitute for these models, which understand complex layouts and also extract bounding boxes for things like tables and pictures. They are also much better at making sense of cursive scripts. I’ve been there, implementing a way to linearise text from a document with pages with 1, 2 or 3 columns, some of them in landscape is a nightmare. And that’s not even considering equations. In the end it’s way easier to use a specialised model, trained by other people to do exactly what I need.
- ad_fontes 2mo agoI've been experimenting with using NuExtract this week on locally OCRing bank statements that don't have a predefined document structure. It's way better than Tesseract or a generic vision-enabled model. It runs great on a single RTX 4090 at the modest throughput I need. Their hosted, API-based service is something like a third of the cost of this model.
- hmokiguess 2mo agoWhoever is paying all that for OCR is being scammed.
- raverbashing 2mo ago* scanned
- parhamn 2mo agoMistral is bumping the price of this thing every release. I think we're at 2x now?
- maelito 2mo agoGiven the latest vibe release's new "follow default" model option, we should see a new coding / general Mistral model, mistral 4, soon.
- ks2048 2mo agoDoes anyone know a site that lets you browse examples of input / output pairs?, particularly with layout analysis (bounding boxes of figures, tables, etc).
- jdrhyne 2mo ago[flagged]
- t3hTao 2mo ago[dead]
- oliveralbertini 2mo agoI'm wondering if this model performs better on french (and other european languages) documents than others
- tethys 2mo agoWho the hell at Mistral thinks it is a good idea to register CMD + T as a shortcut for switching theme!?
- sajithdilshan 2mo agoit's control + T right?
- waldrews 2mo agoThe VLM's are so good at complex document understanding now. But you just can't trust them not to invisibly censor sensitive clinical/legal docs, even at the maximally permissive settings. And the deep learning OCR-only models won't censor, but can and do hallucinate. I've yet to see a 'scan with different approaches and reconcile and say you're not sure if they don't agree' system just work for generic complex documents.
- kolinko 2mo agoThe way my harness set it up is going through 2 or 3 providers, and cross-checking through them, also with plain text extracted if available. I think we also had a layer that for any quote extracted tested it back if it exists within the original. If you wanted 100% accuracy, I think it wouldn't be too difficult nowadays to guess the font&size&other text settings, and re render the crucial parts.
- dangoodmanUT 2mo ago> But you just can't trust them not to invisibly censor sensitive clinical/legal docs What's an example of this?
- nc55g3g 2mo ago[flagged]
- fumeux_fume 2mo agoI think people misunderstand the utility of Mistral's OCR. It's not going to beat SOTA models for extraction on edge-case docs, but it's MUCH cheaper and faster and does an excellent job on simple ones. I've been working on converting PDFs to EPUBs and Mistral has been making steady improvements. On a chapter of Bleak House it was able to extract and tag the header, titles, and references at the bottom every time. The only thing it struggled on was line numbers in the right margin which it correctly tagged as "aside text" 3/5 times, but always separated from the core text each time. The important thing to keep in mind is that there's no prompting needed, just upload the PDF and voila!. There's even a batch mode with a 50% discount.
- mkbkn 2mo ago> The important thing to keep in mind is that there's no prompting needed, just upload the PDF and voila! I'm sorry, noob here. I have a special book that I bought which I can open only inside the Kindle app (Windows/mobile). I have been meaning to screenshot the pages and convert them into a document/PDF. What do I have to do to make it fast? Just upload all the screenshots one by one and tell Mistral "Chat" to OCR them?
- einpoklum 2mo ago> it's MUCH cheaper and faster and does an excellent job on simple ones. OCR should be: 1. Privacy-respecting, i.e. running on your own machine without network communications. 2. Fully open-source. 3. Gratis. The first one is a must, the second is very important for the public interest, and the third one is a nice-to-have. Mistral does not appear to satisfy even the first-, let alone all three.
- einpoklum 2mo agoHow does this perform with non-Latin and non-LTR scripts? Say, Chinese, Arabic, Devangari, Adlam, etc.?
- yunusislegel 2mo agoI didn't quite understand.
- Utkarsh736 2mo agoMistral has been very good with handwritten ocr, expecting the new models to get better with that across languages.
- felooboolooomba 2mo agoBenchmarks? I'm currently using Tesseract (via OcrMyPDF) and would like to compare the difference.