3 ms·
It's still 40. Why not use Ollama-OCR?
by vaxman 1y ago
It's still 40.
Why not use Ollama-OCR?
- krapht 1y agoBecause I benchmarked both on my dataset and found that Tesseract was better for my use-case?
- rafram 1y agoI’ve tested a bunch of vision models on particularly difficult documents (handwritten in a German script that’s no longer used), and I have yet to be impressed. They’re good at BSing to the point that you almost think they nailed it, until you realize that it’s mostly/all made-up text that doesn’t appear in the document.
- yjftsjthsd-h 1y ago> It's still 40. Is it, though? If the important parts of the code are new, does it matter that other parts are older or derived from older code? (Of course, I think this whole line of thought is pointless; what matters is not age, but how well it works, and tesseract generally does seem to work.)
- vaxman 1y agoYeah it is, it does (especially with OOP) and "ABBYY" kicked Tesseract's arse a long time ago anyway. Maybe try OpenAI GPT-4o or Google's Document AI https://cloud.google.com/document-ai https://cloud.google.com/document-ai