6 ms·
Benchmarking vision-language models on OCR in dynamic video environments
- echelon 2y agoAre there any benchmarks (speed, accuracy, etc.) for non-OCR use cases? I want to label images and videos, but don't really care about text.
- nolok 2y agoI have lots of customer files and I've looked around with all these AI tools for something, paid or self hosted or whatever, where I point it to a folder with xlsx and pdf and then I can query "Whats the end date or M Smith contract" or "How much does M Smith still owe" and I've been very disappointed by that, it's either very complicated, or they break down with non text based pdf, or... It feels to me that if you need to provide schema and preprocess the data and this and that at the end all AI provide is a way to do some SQL in natural language, meaning yes it's better but it doesn't remove the actual pain point if you're a tech user. Then again maybe I'm wrong, didn't find the right tool or didn't understand it. Is what I'm looking for something that actually exists (and works, not just on simple cases)?
- fhd2 2y agoI worked on this a bit 1-2 years ago. Back then, LLMs weren't really up to the task, but I found them OK for suggestions that a human double checks. Brings us to the Ironies of Automation though (human oversight of automation with a review process doesn't really work, it's a paper worth reading). We tried several dedicated services for extracting structured data and factoids like that from documents: First Google Document AI, then a dedicated provider focusing solely on our niche. Back then, that gave the best results. There wasn't enough budget to go deeper into this and we just reverted to doing it manually. But I think a really cool way to do this would be to make a user friendly UI where they can see suggestions and the text snippets they were extracted from as they skim through the document, with a simple way to modify and accept these. I think that'd work to scale the process quite a bit. Focusing the attention of the human at the relevant parts of the document basically. Haven't worked on this space since then, but I'm pretty bearish on fully automated fact extraction. Getting stuff in contracts and invoices wrong is typically not acceptable. I think a solid human in the loop approach is probably still the way to go.
- tpm 2y agoI'm not completely up to date but a few months ago Qwen2-VL (runnable locally) was able to perfectly read text from images. So I'd say you would still need to preprocess that folder to texts to get any reasonable speed for queries but after that if you feed the data to a LLM with long enough context it should just work. If on the other hand it's too much data and the LLM is required to use tools then it is indeed still too soon. But it is coming.
- silveraxe93 2y agoPosted 4 days ago: > Three state of the art VLMs - Claude-3, Gemini-1.5, and GPT-4o Literally none of those are state of the art. Academia is completely unprepared to deal with the speed Ai develops. This is extremely common in research papers. That's literally in the abstract. If I can see a completely wrong sentence 5 seconds into reading the paper, why should I read the rest?
- lisnake 2y agoThey may have been SotA at the moment of writing
- silveraxe93 2y agoSure, but they posted this 4 days ago. The minimum I'd expect for quality research is for them to skim the abstract before posting and change that line to: "Models from leading AI labs" or similar. Leaving it like now signals either sloppiness or dishonesty
- michaelt 2y agoWhat models would you recommend instead, for sophisticated OCR applications? Honestly I thought Claude-3 and GPT-4o were some of the newest major models with vision support, and that models like o1 and deepseek were more reasoning-oriented than OCR-oriented.
- silveraxe93 2y agoFor Google, definitely flash-2.0; It's a way better model. GPT-4o is kinda dated now. o1 is the one I'd pick for OpenAI. It's basically their "main" model now. I'm not that familiar with Claude for vision. I don't think Anthropic focusses on that. But the 3.5 family of models is way better. If 3.5 Sonnet supports vision that's what I'd use
- thelittleone 2y agoAnthropic has a beta endpoint for PDFs which has produced impressive results for me with long and complex PDFs (tables, charts etc).
- _stillmind 2y agoThe paper says, "GPT-4o achieves the highest overall accuracy, while Gemini-1.5 Pro demonstrates the lowest word error rate." Saying Gemini "beats everyone" in this benchmark is misleading.
- nolist_policy 2y agoNotably, they tested Gemini-1.5 Pro while Gemini 2.0 is another step up.[1] Between this and their 1M token context it is getting hard to ignore Google's models. [1] https://news.ycombinator.com/item?id=42952605 https://news.ycombinator.com/item?id=42952605
- spwa4 2y agoThis looks a lot like "compared to a bunch of people who are 10 years behind (non-transformer, vision-only models), and people who aren't trying (aren't optimizing for OCR) Google is doing real well" EasyOCR is LSTM-CTC from 2007, RapidOCR is a ConvNet approach from 2021, both focused on speed. Both will vastly outperform almost any transformer model, and certainly a big one, on speed and memory usage, but they aren't state of the art on accuracy. This is well known, for a decade at this point. 2 decades for LSTM-CTC. Plus, I must say the GPT-4o results look a lot saner. "COCONUT" (GPT-4o) vs "CONU CNBC" (Gemini) vs Ground Truth "C CONU CNBC". And, obviously the ground truth should be "COCONUT MILK" (the word milk is almost entirely out of the picture, but is still the right answer that a human would give). The "C CONU" comes from the first O of COCONUT being somewhat obscured by a drawing of ... I don't know what the hell that is. It's still very obvious it's meant to be "COCONUT MILK", so the GPT-4o answer is still not quite perfect, but heaps better than all the others. Now this looks very much like it might be temperature related, and I can find nothing in the paper about changing the temperature, which is imho a very big gap (temperature gives transformer models more freedom to choose more creative answers. The better performance of GPT-4o might well be the result of such a more creative choice, and might also explain why Gemini is trying so hard to stay so very close to the ground truth. It's still quite the accomplishment to succeed, but GPT-4o is still better)
- yorwba 2y agoWhat would you say is currently the most accurate OCR solution if you're not concerned about speed and memory usage?
- OJFord 2y agoNot GP but it depends what you mean by accuracy. If you want inference like the 'coconut milk' described then obviously an LLM. If you want accurate as-written transcription, then I don't know the state of the art, but it'll be something purpose built for CV & handwriting recognition. It'll also depend if you care about tabular data, whether a 'minor' numerical error (like 0 & 8 mismatched sometimes) is significantly worse than a 'typo' as it were in recognising a word, etc.
- breadislove 2y agoThe systems they tested against the LLMs are mostly used as a part of a larger system. A more fair comparison would be to use something like MinerU [1] and proper benchmark like the OHR Bench [2] and Reductos table bench [3]. This paper is really bad... [1]: https://github.com/opendatalab/MinerU https://github.com/opendatalab/MinerU [2]: https://github.com/opendatalab/OHR-Bench https://github.com/opendatalab/OHR-Bench [3]: https://github.com/reductoai/rd-tablebench https://github.com/reductoai/rd-tablebench
- croes 2y agoDoes everyone also need huge data centers at lots of energy?
- retskrad 2y agoPeople say CPU benchmarks are meaningless (what does even 10-15% better mean in practice?) but LLM benchmarks are even more of a mystery. The same LLM will produce a novel output everytime you given it the exakt same prompt.
- stavros 2y agoNo it won't.
- casey2 2y agoIt's not surprising that google has such a huge mote with their highly illegal and unethical activity of scanning and digitizing billions of pages of copyrighted work to train their models. Oh wait, google books search was fair use. I got it confused with LLMs.
- Terretta 2y ago> It's not surprising that google has such a huge mote with their highly illegal and unethical activity of scanning and digitizing billions of pages of copyrighted work to train their models. Excellent Freudian slip (proverb allusion suggesting Google has a blind spot, while discussing OCR).
- alberto-m 2y agoIt seems to me that the software is occasionally doing better than the supposed “ground truth” (who annotated that?), and I don't understand why the authors are blindly following the latter, and the reviewers apparently approved that. In Figure 1 the authors complain that Gemini “misreads 'ss ety!' as 'ness ety!'”, but even a casual look at the image reveals that Gemini's reading is correct. In Figure 11, they state that Claude is “altering the natural sequence of ideas in the ground truth”, except that the sequence in the ground truth makes no sense, while Claude's order does (only the initial “the” is misplaced).
- virgilp 2y agoI think the goal here was to convince the AI to actually read chars ("OCR") rather than speculate what might be written on paper/in the image. Hence why the ground truth is explicitly removing the letters & word parts that are obscured, even when they can be guessed. TBH, I'm not sure it's a good test. I can somewhat see the argument against "BASELINE" for ground truth - the underlying text might have been BASE(IAKS), for all we know. But, IMO the ground truth should have been "Direction & ess" at the very least. And, more significantly than that - it's a fake scenario, that we don't care for in practice. Why use that? Use invoices with IDs that sound like words but are not. Use license plates and stuff like that. Heck, use large prints of random characters, mixed with handwritten gibberish. For at least some of images that they used, the expectation from a good text reader is actually to understand context and not blindly OCR. Take "Trader Joe's": we *know* that's an 's', but only from outside context; from OCR, it might've been an 8, there's really no way to tell. Why accept the "s" in ground truth, but reject the full world "Coconut" (which is obviously what is written on the can, even if partially obscured)? Furthermore, a human would know what kind of products are sold by Trader Joe's, and coupling that with the top of the letters "M I L" that are visible, would deduce that's Coconut Milk. So really, Claude nailed that one.
- 8organicbits 2y agoI think there are multiple possible goals we could imagine in text recognition tasks. Should the AI guess the occluded text? That could be really helpful in some instances. But if the goal is OCR, then it should only recognize characters optically, and any guessing at occluded characters is undesired.
- belter 2y agoGemini is so bad I gladly cancelled my paid account. But hey, maybe AI and 50B dollars is what was needed to get a better OCR...
- bobjordan 2y agoReally? This surprises me because I use open-AI pro for $200 per month and I still fall back to using my Gemini $20 per month account a lot these days, I like the new 2.0 experimental speed and how it defaults to diving into producing usable code, immediately. Whereas, my open ai pro mode will spend a few minutes to give me an initial answer that beats around the bush at a much higher level to start. So, my workflow has evolved to using Gemini to iterate my initial thinking and frame out requirements and first draft code. Then, when I get about 2,000 - 3,000 lines for a detailed initial pro mode prompt, I send that to open ai pro mode and then it shines. But, I really like starting with the Gemini 2.0 model first. The main thing I dislike about Gemini is I often need to tell it “please continue” when it reaches its output limit. But it nearly always just picks up right where it left off and continues its output. This is critical in using Gemini.
- malanj 2y agoIf you're wondering how they prompt the models: "Perform OCR on this image. Return only the text found in the image as a single continuous string without any newlines, additional text, or commentary. Separate words with single spaces. For any truncated, partially visible, or occluded text, include only the visible portions without attempting to complete or guess the full text. If no text is present, return empty double quotes." Found in: https://github.com/video-db/ocr-benchmark/blob/main/prompts.yaml https://github.com/video-db/ocr-benchmark/blob/main/prompts....
- Terretta 2y agoTL;DR: For original object truth rather than image truth, this paper shows VLMS are superior, even though prompt shows the authors are "holding it wrong". Yet another paper where the authors don't address what tokens are. It's like publishing Rolling pin fails at math or Calculator fails to turn dough ball into round pizza. While I can understand where they're coming from in a desire to avoid hallucination when doing some letter for letter transcription from an image, certainly most times you reach for OCR you want the original copy, despite damage to its representation (paper tears, coffee stains, hands in front of it). Turns out token conjunction probability conjectures come in handy here! Whether the image of an object, or the object, is "Ground Truth" is an exercise left to the user's goal. Almost all use cases would want what was originally written on the object, not its present occlulded [sic] representation.
- hubraumhugo 2y agoAs someone building in this space, we've found that raw OCR accuracy is just one piece (and it's becomming a commodity). The real challenge is building reliable and accurate ETL pipelines (document ingestion from web, OCR, classification, validation, etc.) that work at scale in production. The best products will be defined by everything "non-AI", like UX, performance, and human-in-the loop feedback loop for non-techies. Avoiding over-reliance on specific models also helps. With good internal eval data and benchmarks, you can easily switch or fine-tune models.
- mtrovo 2y agoThat’s the point of using AI in the first place. If your product is just a polished interface on top of a prompt, then your moat isn’t that strong, and chances are your product will be commoditized soon. By building a good UX and integrating it with other processes that require traditional collaboration, you increase the chances that replicating your secret sauce is either infeasible or too difficult for newcomers to bother.
- HannesWes 2y agoThis looks very interesting. I conducted some explorations of whether LLMs can be used to extract information from hand-written forms [0][1]. Such a system could allow users to snap pictures of forms and other legal documents, automatically extract structured information, and use this information to e.g. automatically fill out new forms or determine whether the user has the right to a government benefit. The initial results were quite promising, as GPT-4o could reliably identify the correct place in the form for the information, and moderately reliably extract the values, even if the image was blurry or the text was sloppily written. Excited to see how Gemini 2.0 would do on this task! [0] https://arxiv.org/abs/2412.15260 https://arxiv.org/abs/2412.15260 [1] https://github.com/hwestermann/AI4A2J_analyzing_images_of_legal_documents https://github.com/hwestermann/AI4A2J_analyzing_images_of_le... (code and data)
- deivid 2y agoAre there any "good" OCR models that run in restricted/small environments? I'm thinking about local models for phone-sized CPUs Obviously these models would have lower accuracy, but running at all would be nice.
- Terretta 2y agoSee comment above: https://news.ycombinator.com/item?id=43048326 https://news.ycombinator.com/item?id=43048326 Throw those into Google search along with term iOS or Android.