3 ms·
We’ve been doing exactly this by doubling-down on VLMs (https://vlm.run https://vlm.run) - VLMs are way better at handling layout and context where OCR systems
by fzysingularity 2y ago
We’ve been doing exactly this by doubling-down on VLMs (https://vlm.run https://vlm.run)
- VLMs are way better at handling layout and context where OCR systems fail miserably
- VLMs read documents like humans do, which makes dealing with special layouts like bullets, tables, charts, footnotes much more tractable with a singular approach rather than have to special case a whole bunch of OCR + post-processing
- VLMs are definitely more expensive, but can be specialized and distilled for accurate and cost effective inference
In general, I think vision + LLMs can be trained to explicitly to “extract” information and avoid reasoning/hallucinating about the text. The reasoning can be another module altogether.
- yigitkonur35 2y agoI did a ton of Googling before writing this code, but I couldn't find you guys anywhere. If I had, I'd have definitely used your stuff. You might want to think about running some small-scale Google Ads campaigns. They could be especially effective if you target people searching for both LLM and OCR together. Great product, congratz!
- fzysingularity 2y agoHey, thanks! DM me if you want to test it out (sudeep@vlm.run). Agreed on SEO - we’re redoing our landing page and searchability. We recently rebranded, hence the lack of direct search hits for LLM / OCR.