3 ms·
Anything that mentions tesseract is about 10 years out of date at this point.
by abc-1 1y ago
Anything that mentions tesseract is about 10 years out of date at this point.
- amelius 1y agoWell, at least I can apt-get install tesseract. That doesn't hold for any of the GPU-based solutions, last time I checked.
- booder1 1y ago5.5.0 released November last year. Still a very active project as far as I can tell and runs on CPU. Even compared to best open source GPU option it is still pretty good. VLMs work very differently and don't work as well for everything. Why is it out of date?
- cbsmith 1y agoI don't know that that is true: https://researchify.io/blog/comparing-pytesseract-paddleocr-and-surya-ocr-performance-on-invoices https://researchify.io/blog/comparing-pytesseract-paddleocr-... Using Surya gets you significantly better results and makes almost all the work detailed in the article largely unnecessary.
- booder1 1y agoSurya weights for the models are licensed cc-by-nc-sa-4.0 so not free for commercial usage. Also, as far as I know, the training data is 100% unavailable. Given they use well trained, but standard models, it isn't really open source and barely, maybe, open weight. I kinda hate how their repo says gpl cause that is only true for the inference code. The training code is closed source.
- cbsmith 1y agoI did not know that the training code is closed source. That is troubling.
- krapht 1y agoI just built a pipeline with tesseract last year. What's better that is open source and runnable locally? VLLM hallucination is a blocker for my use case.
- criddell 1y agoIf you are stuck with open source, then your options are limited. Otherwise I'd say just use your operating system's OCR API. Both Windows and MacOS have excellent APIs for this.
- stavros 1y agoHow is a hallucination worse than a Tesseract error?
- jgalt212 1y agoHallucinations are hard to detect unless you are a subject-matter expert. I don't have direct experience with Tesseract error detection.
- krapht 1y agoBecause the VLM doesn't know it hallucinated. When you get a Tesseract error you can flag the OCR job for manual review.
- amelius 1y agoIt could hallucinate obscene language, something which is less likely with classic OCR.
- gessha 1y agoLatter is more likely to get debugged.
- fxtentacle 1y agoQuite simply, you’re completely wrong. Modern tesseract versions include a modern LSTM AI. It can very affordably be deployed on CPU, yet its performance is competitive with much more expensive large GPU-based models. Especially if you handle a high volume of scans, chances are that tesseract will have the best bang per buck.
- nicman23 1y agoi remember that you could not train it your self in a font like you could in older versions, it that still the case?
- ianhawes 1y agoMy company probably spent close to 6 figures overall creating Tesseract 5 custom models for various languages. Surya beats them all and is open source (and quite faster).
- booder1 1y agoSurya weights for the models are licensed cc-by-nc-sa-4.0. They have an exception for small companies. If you're company is not small you either need to pay them or use them illegally. Their training code and data is closed source. They are barely open weight and only inference is open source.