4 ms·
> Honestly I was kind of surprised that good basic OCR isn't a totally solved issue with an ecosystem of fully open-source solutions by now. Yes! Can anyone co
by vortex_ape 6y ago
> Honestly I was kind of surprised that good basic OCR isn't a totally solved issue with an ecosystem of fully open-source solutions by now.
Yes! Can anyone comment on why this is the case, since OCR is proclaimed to be a solved problem?
I've always wondered why Google Lens works "out of the box" and shows great accuracy on extracting text from images taken using a phone camera, but open-source OCR software (Tesseract, Ocropy etc.) needs a lot of tweaking to extract text from standard documents with standard fonts, even after heavily pre-processing the images.
PS: Has Google released any paper on Google Lens?
- craftinator 6y agoI've been wondering this ever since I used Lens. My hobby applications doing OCR always fall way short of Len's magic.
- vortex_ape 6y agoYeah! And Lens is not the only closed-source OCR solution that works. I've gotten great accuracy using ABBYY and docparser.com in the past. But one needs to pay per page after the free trial ends :(
- claudeganon 6y agoI’ve found that none of the open source stuff works well for Japanese language documents. Most of the time, I’ve just ran them through Adobe Acrobat’s OCR and dumped the results into a text file. There are still mistakes, but it at least returns a passable result compared to others.
- shiredude95 6y agoI was building an image search engine[0] a while back and faced the same issues you mentioned with OCR. What i realized is tesseract[1](one of the more popular ocr framework) works so long as you are able to provide it data similar to the one it was trained on. We were basically trying to transcribe message screenshots which should have been relatively straightforward given the homogeneity of the font. But this was not the case as tesseract was not trained in the layout of msg screenshots. The accuracy of raw tesseract on our test dataset was somehwere about 0.5-0.6 BLEU. Once we were able to isolate individual parts of the image and feed it to tesseract, we were able to get around 0.9 BLEU on the same dataset. TLDR;Some nifty image processing is required to make tesseract perform as expected. [0] (https://www.askgoose.com https://www.askgoose.com) [1] (https://github.com/tesseract-ocr/tesseract https://github.com/tesseract-ocr/tesseract)