6 ms·
GLM-OCR – A multimodal OCR model for complex document understanding
- aliljet 8mo agoThis is actually the thing I really desperately need. I'm routinely analyzing contracts that were faxed to me, scanned with monstrously poor resolution, wet signed, all kinds of shit. The big LLM providers choke on this raw input and I burn up the entire context window for 30 pages of text. Understandable evals of the quality of these OCR systems (which are moving wicked fast) would be helpful... And here's the kicker. I can't afford mistakes. Missing a single character or misinterpreting it could be catastrophic. 4 units vacant? 10 days to respond? Signature missing? Incredibly critical things. I can't find an eval that gives me confidence around this.
- cinntaile 8mo agoDeciphering fax messages? What is this, the 90s?
- xyproto 8mo agoFax is still hard to hack, so some organizations have kept it alive for security.
- meatmanek 8mo agoI think the most useful thing about faxes, security-wise, is that in their basic form they require zero digital storage of the image being sent. The only record on either side of the transmission is a piece of paper.* Contrast that with email, which is store-and-forward by design, and now you have to put in effort to ensure both the sending and receiving email providers delete the message in a timely manner. * obviously you can add store-and-forward behavior to either fax machine, but it's not the default.
- kergonath 8mo agoWe have decades of internal reports on film that we’d like to make accessible and searchable. We don’t do it with new documents, but we have a huge backlog.
- daveguy 8mo agoIf your needs are that sensitive, I doubt you'll find anything anytime soon that doesn't require a human in the loop. Even SOTA models only average 95% accuracy on messy inputs. If that's a per character accuracy (which OCR is generally measured by), that's going to be 5+ errors per page of 100+ words. If you really can't afford mistakes you have to consider the OCR inaccurate. If you have key components like "days to respond" and "units vacant" you need to identify the presence of those specifically with bias in favor of false positives (over false negatives), and human confirmation of the source-> OCR.
- kergonath 8mo ago> If you really can't afford mistakes you have to consider the OCR inaccurate. Isn’t this close to the error rate of human transcription for messy input, though? I seem to remember a figure in that ballpark. I think if your use case is this sensitive, then any transcription is suspicious.
- aliljet 8mo agoThis is precisely the real question. If you're exceeding human transcription, you may be generally pretty good. The question is what happens when you tell a human to become surgical about some part of the document, how then does the comparison change..
- coder543 8mo agoIf you want OCR with the big LLM providers, you should probably be passing one page per request. Having the model focus on OCR for only a single page at a time seemed to help a lot in my anecdotal testing a few months ago. You can even pass all the pages in parallel in separate requests, and get the better quality response much faster too. But, as others said, if you can't afford mistakes, then you're going to need a human in the loop to take responsibility.
- HPsquared 8mo agoYou could maybe then do a second pass on the whole text (as plain text not OCR) to look for likely mistakes.
- kergonath 8mo agoThis is not always easy. The models I tried were too helpful and rewrote too much instead of fixing simple typos. When I tried I ended up with huge prompts and I still found sentences where the LLM was too enthusiastic. I ended up applying regexes with common typos and accepted some residual errors. It might be better now, though. But since then I’ve moved to all-in-one solutions like Mathpix and Mistral-OCR which are quite good for my purpose.
- staticman2 8mo agoGemini Pro 3 seems to be built for handling multiple page PDFs. I can feed it a multiple page PDF and tell it to convert it to markdown and it does this well. I don't need to load the pages one at a time as long as I use the PDF format. (This was tested on A.i. studio but I think the API works the same way).
- coder543 8mo agoIt's not that they can't do multiple pages... but did you compare against doing one page at a time? How many pages did you try in a single request? 5? 50? 500? I fully believe that 5 pages of input works just fine, but this does not scale up to larger documents, and the goal of OCR is usually to know what is actually written on the page... not what "should" have been written on the page. I think a larger number of pages makes it more likely for the LLM to hallucinate as it tries to "correct" errors that it sees, which is not the task. If that is a desirable task, I think it would be better to post-process the document with an LLM after it is converted to text, rather than asking the LLM to both read a large number of images and correct things at the same time, which is asking a lot. Once the document gets long enough, current LLMs will get lazy and stop providing complete OCR for every page in their response. One page at a time keeps the LLM focused on the task, and it's easy to parallelize so entire documents can be OCR'd quickly.
- chrsw 8mo agoI'm keeping my eye on progress in this area as well. I need to free engineering design data from tens of thousands of PDF pages and make them easily and quickly accessible to LLMs.
- aliljet 8mo agoAll of healthcare is crying. Trust me.
- Imustaskforhelp 8mo agoI suppose tears of joy?
- fragmede 8mo agoOf sadness because they're not allowed to use it yet.
- saidinesh5 8mo agoDo you have more details about this?
- deleted 8mo ago[deleted]
- deleted 8mo ago[deleted]
- renewiltord 8mo agoI’m sure you’ve tried all this but you’ve tried inter-rater agreement via multiple attempts on same LLM vs different LLM? Perhaps your system would work better if you ran it through 5 models 3 times and then highlighted diffs for human chooser.
- coder543 8mo agoThere are a bunch of new OCR models. I’ve also heard very good things about these two in particular: - LightOnOCR-2-1B: https://huggingface.co/lightonai/LightOnOCR-2-1B https://huggingface.co/lightonai/LightOnOCR-2-1B - PaddleOCR-VL-1.5: https://huggingface.co/PaddlePaddle/PaddleOCR-VL-1.5 https://huggingface.co/PaddlePaddle/PaddleOCR-VL-1.5 The OCR leaderboards I’ve seen leave a lot to be desired. With the rapid release of so many of these models, I wish there were a better way to know which ones are actually the best. I also feel like most/all of these models don’t handle charts, other than to maybe include a link to a cropped image. It would be nice for the OCR model to also convert charts into markdown tables, but this is obviously challenging.
- StableAlkyne 8mo agoHow do these compare to something like Tesseract? I remember that one clearing the scoreboard for many years, and usually it's the one I grab for OCR needs due to its reputation.
- deleted 8mo ago[deleted]
- kergonath 8mo agoTesseract does not understand layout. It’s fine for character recognition, but if I still have to pipe the output to a LLM to make sense of the layout and fix common transcription errors, I might as well use a single model. It’s also easier for a visual LLM to extract figures and tables in one pass.
- rdos 8mo agoIs it possible for such a small model to outperform gemini 3 or is this a case of benchmarks not showing the reality? I would love to be hopeful, but so far an open source model was never better than a closed one even when benchmarks were showing that.
- amluto 8mo agoOff the top of my head: for a lot of OCR tasks, it’s kind of worse for the model to be smart. I don’t want my OCR to make stuff up or answer questions — I want to to recognize what is actually on the page.
- rdos 8mo agoInteresting. Won't stuff like entity extraction suffer? Especially in multilingual use cases. My worry is that a smaller model might not realize some text is actually a persons name because it is very unusual.
- kergonath 8mo agoThe model does not need to be that smart to understand that a name it does not know that starts with a capital letter is a the name of a place or a person. It does not need to be aware of whom this refers to, it just needs to transcribe it. Also, there are generalist models that have enough of a grasp of a dozen or so languages that fit comfortably in 7B parameters. Like the older Mistral, which had the best multi-lingual support at the time, but newer models around that size are probably good candidates. I am not surprised that a multilingual specialised model can fit in 8B or so.
- retrac 8mo agoSometimes what is on the page is ambiguous. Imagine a scan where the dot over the i is missing in a word like "this". What's on the page is "thls" but to transcribe it that way would be an error outside of forensic contexts. I am reminded it's basically impossible to read cursive writing in a language you don't know even if it's the same alphabet.
- 8mo ago
- alaanor 8mo agoThere was so many OCR models released in the past few months, all VLM models and yet none of them handle Korean well. Every time I try with a random screenshot (not a A4 document) they just fail at a "simple" task. And funnily enough Qwen3 8B VL is the best model that usually get it right (although I couldn't get the bbox quite well). Even more funny, whatever is running on an iphone locally on cpu is insanely good, same with google's OCR api. I don't know why we don't get more of the traditional OCR stuff. Paddlepaddle v5 is the closest I could find. At this point, I feel like I might be doing something wrong with those VLMs.
- ghrl 8mo agoI remember someone building a meme search engine for millions of images using a cluster of used iPhone SE's because of Apple's very good and fast OCR capabilities. Quite an interesting read as well: https://news.ycombinator.com/item?id=34315782 https://news.ycombinator.com/item?id=34315782
- fzysingularity 8mo agoApple OCR even on the Mac is insanely good, in fact way better than AWS textract/GCP cloud vision OCR. Any idea what model is being used?
- AlphaSite 8mo agoProbably some custom model built for their hardware.
- Stagnant 8mo agoChrome ships a local OCR model for text extraction from PDFs which is better than any of the VLM or open source OCR models i've tried. I had a few hundred gigs of old newspaper scans and after trying all the other options I ended up building a wrapper around the DLL it uses to get the text and bboxes. Performance and accuracy on another level compared to tesseract, and while VLM models sometimes produced good results they just seemed unreliable. I've thought of open sourcing the wrapper but havent gotten around to it yet. I bet claude code can build a functioning prototype if you just point it to "screen_ai" dir under chrome's user data.
- deleted 8mo ago[deleted]
- bugglebeetle 8mo agoI tested this pretty extensively and it has a common failure mode that prevents me from using: extracting footnotes and similar from the full text of academic works. For some reason, many of these models are trained in a way that results in these being excluded, despite these document sections often containing import details and context. Both versions of DeepseekOCR have the same problem. Of the others I’ve tested, dot-ocr in layout mode works best (but is slow) and then datalab’s chandra model (which is larger and has bad license constraints).
- droidjj 8mo agoI have been looking for an OCR model that can accurately handle footnotes. It’s essential for processing legal texts in particular, which often have footnotes that break across pages. Sadly I’ve yet to encounter a good solution.
- kergonath 8mo agoI found Mathpix to be quite good with this type of documents, including footnotes but to be fair my documents did not have that many. It’s also proprietary.
- sgc 8mo agoI can get multiple sets of footnotes (critical + content notes) reliably recognized and categorized using gemini-3-flash-preview. I took 15-20 hours to iterate on my prompt for a specific format. Otherwise it would not produce good enough results. It was a slow process because results from batch did not mirror what I was getting from the chat mode, and you have to wait for batch results while analyzing the last set. There was also a bit of debugging of the batch protocol going on at the same time. Flash is also surprisingly affordable for the results I am getting, 4-5x less than I had anticipated. I gave up on gemini-3-pro pretty quickly because it overthinks and messes things up.
- ks2048 8mo agoI've been trying different OCR models on what should be very simple - subtitles (these are simple machine-rendered text). While all models do very well (95+% accuracy), I haven't seen a model not occasionally make very obvious mistakes. Maybe it will take a different approach to get the last 1%...
- raphaelmolly8 8mo ago[dead]
- sinandrei 8mo agoHas anyone experiment with using VLM to detect "marks"? Thinking of pen/pencil based markings like underlines, circles,checkmarks.. Can these models do it?
- leetharris 8mo agoNone of them do it well from our experience. We had to write our own custom pipeline with a mixture of legacy CV approaches to handle this (AI contract analysis). We constantly benchmark every new multimodal and VLM model that comes out and are consistently disappointed.
- coder543 8mo agoIf someone releases a benchmark/dataset, I'm sure that significantly increases the chances of one of these AI labs training on the task.
- mikae1 8mo agoText me back when there's a working PDF to EPUB conversion tool. I've been waiting (and searching for one) long enough. :D EDIT: https://github.com/overcuriousity/pdf2epub https://github.com/overcuriousity/pdf2epub looks interesting.
- simpleusername 8mo ago[dead]
- surfacedamage 8mo agoThis might be a niche question, but does glm-ocr (or other libraries) have the ability to extract/interpret QR code data?
- simpleusername 8mo agozbar, pyzbar, OpenCV + QRCodeDetector, Gemini + Vision AI, Claude + Vision
- ThrowawayTestr 8mo agoWhat's the current SOTA for Japanese and Korean OCR? BalloonsTranslator has a great workflow but the models are pretty old.
- bartread 8mo ago> Option 1: Zhipu MaaS API (Recommended for Quick Start) > Use the hosted cloud API – no GPU needed. ... > Option 2: Self-host with vLLM / SGLang So, first off, this looks really cool and, given I'm looking for OCR at the moment, I'm pretty interested in this and other OCR models. With that said, the README implies that option 2 requires a GPU. That's fine but it would be incredibly helpful if the README were explicit about requirements, and especially the amount of memory it needs. EDIT: Looking at the links under option 3, the docs for macOS setup suggest 8GB of unified memory is enough to run the model, which is pretty modest, so I'd imagine Option 2 is similar. Ollama also offers a CPU only option (no idea how that will perform - not amazingly, I'm guessing), but that would suggest to me that if your volume requirements are low and you can't shell out for or source a beefy enough GPU and don't want to pay the sometimes exhorbitant hire costs, you should be able to punt it on to a machine with enough memory to run the model without too much difficulty.
- TZubiri 8mo agoWas it trained by distilling some other model?