5 ms·
From what I can tell (without having read the research papers) it looks like this is just an easy to use package for sparse scene text extraction. It seems to d
by dclusin 6y ago
From what I can tell (without having read the research papers) it looks like this is just an easy to use package for sparse scene text extraction. It seems to do okay if the scene has sparse text but it
falls down for dense text detection. The results are going to be pretty bad if you try and do a task like "extract transactions from a picture of a receipt." Here's an example of input you might get for a production app: https://www.clusin.com/walmart-receipt.jpg https://www.clusin.com/walmart-receipt.jpg
Notice the faded text from the printer running out of ink and the slanted text. From limited experience each of these are thorny problems and the state of the art CV algorithms won't help you escape from having to learn how to algorithmicly pre-process images and clean them up prior to feeding them into a CV algorithm. You might be able to use Google's Cloud OCR but that charges per image, although it is pretty good. Even if you use that you've graduated to the next super difficult problem which is Natural Language Processing.
Once you have the text you need to determine if it has meaning to your application. That's basically what NLP is about. For the receipts example, how do you know you're looking at a receipt? What if its a receipt on top of a pile of other receipts? How do you extract transactions from the receipt? Does a transaction span multiple lines? How can you tell? etc etc etc.
- jonathanstrange 6y agoSo your point is that this library is not a magic unicorn that solves all problems related to OCR and natural language processing?
- nacs 6y agoTry reading the post. There’s a lot more there but the gist is that this is optimized for a different set of OCR uses and not the more typical scan a book/receipt cases.
- dclusin 6y agoThis is a fair point. I think my criticism more generally is that they position it as easy to use but its still just another library for a subset of OCR problems: sparse text extraction from a scene. As I said in a sibling post there doesn't seem to be a library that stitches together OCR approaches for all the different use cases and chooses an approach based on analyzing the image itself. That would be truly easy to use.
- D13Fd 6y agoI'm just happy to see some advancement in open source OCR for Python. Last time I had a Python project that needed OCR, I found that the open-source options were surprisingly limited, and it required some effort to achieve consistently good results even with relatively clean inputs. Honestly I was kind of surprised that good basic OCR isn't a totally solved issue with an ecosystem of fully open-source solutions by now.
- dclusin 6y agoFrom my experience the algorithms & implementations seem to be pretty good but the caveat is that you the developer need to be aware of all the different approaches and when it is appropriate to apply them. There just doesn't seem to be a good general purpose library that stitches them all together and knows when to use which approach based analyzing the image.
- catalogia 6y agoI've found that often for tools related to natural language (ORC, text-to-speech, and speech-to-text) it feels like you need a PhD in the subject just to figure out how to anything done at all. I heartily welcome efforts to package these sort of things up in ready-to-use ways.
- dclusin 6y agoThis is good news if you have one of these PhD's. Your career probably isn't going anywhere any time soon :)
- vortex_ape 6y ago> Honestly I was kind of surprised that good basic OCR isn't a totally solved issue with an ecosystem of fully open-source solutions by now. Yes! Can anyone comment on why this is the case, since OCR is proclaimed to be a solved problem? I've always wondered why Google Lens works "out of the box" and shows great accuracy on extracting text from images taken using a phone camera, but open-source OCR software (Tesseract, Ocropy etc.) needs a lot of tweaking to extract text from standard documents with standard fonts, even after heavily pre-processing the images. PS: Has Google released any paper on Google Lens?
- sinuhe69 6y agoYes, you're right. I tried with some scanned pages from a Vietnamese book but the result was very bad (say <5% accurate). The scans was pretty OK, though. Probably the model was not trained much for the Vietnamese language but I think it's more likely that it does not do the necessary per-processing steps.
- tasogare 6y agoI had very bad results on Vietnamese using Tesseract and their trained model. French output was mostly fine. I guess less attention is given to some language, and the huge number of diacritics used in Vietnamese make it harder to process too.
- impostervt 6y agoI've been very impressed with the OCR on an app called Fetch, which you use to scan your grocery receipts and get points you can use to redeem for gift cards. Even if I pull a receipt out of my pocket and it's wrinkly, it still seems to read it very well.
- pjc50 6y agoCan you get the data from them yourself, or is it purely for them? I've just tried easyocr on a receipt, and it's pretty bad. I've also just noticed that ASDA have a "mojibake" problem and print ú instead of £ on the entire receipt ...
- impostervt 6y agoI haven't looked into it, I believe it's purely for them. It's sort of like a reverse-coupon app. You buy stuff, and get extra points for say, Lipton iced tea. That's supposed to encourage you to buy more of that stuff next time.
- maire 6y agoAbout a year ago I surveyed the available OCR packages for receipts. This was for pristine scans (not the crumpled scan you have in your image). In my survey all OCRs failed except google cloud OCR! If there is another OCR that works I would love to know.
- memexy 6y agoI use TesseractOCR for general screenshot text extraction. Granted they're not receipts but Tesseract works well enough. What packages did you survey? Do you still have the data and code?