4 ms·
Lots of reasons they haven't been wiped off the map. 1. Full page OCR is trivial these days. Anyone who's just doing full-page OCR has no business using one of
by staticautomatic 9y ago
Lots of reasons they haven't been wiped off the map.
1. Full page OCR is trivial these days. Anyone who's just doing full-page OCR has no business using one of these commercial offerings. The commercial stuff is really for extraction of structured info from unstructured and semi-structured documents. ABBYY and Nuance are the only ones with products that can handle it. There are alternatives for simple capture tasks (like DocParser), but not for complex ones.
2. ABBYY and Nuance have a lot of IP in the space on lock.
3. The market for complex data extraction (at least the kind of stuff I do) may not actually be big enough for smaller players to bother pursuing (most new players are going after tasks of intermediate difficulty, like invoice and receipt capture).
4. There's still a need for on-prem style solutions for people doing huge volumes (hundreds of thousands of pages a year), whereas most of the new market entrants are cloud only.
5. I haven't used Nuance's stuff much but ABBYY's products are actually incredibly robust under the hood-- stuff you wouldn't want to build yourself unless you absolutely had to. FlexiCapture will do 95%+ accuracy out of the box (on a character by character basis) with no human verification.
- derefr 9y ago> Anyone who's just doing full-page OCR has no business using one of these commercial offerings. So what are you supposed to use for full-page OCR?
- ocrcustomserver 9y agoFor an opensource solution that uses Tesseract, check out ocrmypdf: https://github.com/jbarlow83/OCRmyPDF https://github.com/jbarlow83/OCRmyPDF
- rb2018 9y agoAs far as hosted solutions go, the best are Google Cloud Vision, Azure OCR and OCr.space You can compare all three here: https://ocr.space/compare-ocr-software https://ocr.space/compare-ocr-software
- ocrcustomserver 9y agoI wouldn't say that full page OCR is trivial. Using an opensource solution (99% based on Tesseract) is going to get you ok-ish results if your input is relatively clean (no complex layout, scanned documents from a flatbed scanner, standard fonts) and you don't care about speed. If you care about recognition accuracy then Tesseract isn't going to cut it (at least not without some serious effort). Replying to points 1 and 3: For smaller players and/or complex tasks you can always implement your own custom parser. I'm doing work as a contractor in this space.
- staticautomatic 9y agoI agree with you that Tesseract isn't great out of the box, but if you aren't doing huge volumes, there are plenty of cloud options available. Respectfully, I disagree about this being a parsing issue. The whole reason so-called "zonal ocr" exists is because of the challenges of reliably inferring the structure of a document at parsing time. Yes, there are some kinds of documents where parsing logic alone will suffice, but for more complex tasks you need what ABBYY and Nuance are selling.
- ocrcustomserver 9y agoJust to be sure that we're talking about the same thing, by "custom parser" I meant implementing your own barebones "zonal OCR" functionality with just the features that are needed for the specific problem. I think it boils down to the needs of each individual application. Some cases have a lot of templates and need the "automatic fuzzy matching" functionality and the extra bells and whistles. But smaller players often deal with just a handful of relatively simple templates where FlexiCapture would just be overkill (not to mention a couple of other problems that I'm covering at the end of the post). This is of course not an easy task because you need someone who can design and implement an end to end system that possibly involves image processing, "zonal OCR", include an OCR engine and also perform reliable text extraction from images/PDFs (extracting text from PDFs is tricky). It's way easier for a non developer to think about what rulesets/logic to apply and not having to think about the image processing/OCR bits. I think that is one of the main selling points of FlexiCapture. It abstracts the OCR bits so that the system designer can think about the problem itself, design a spec and think about the logic the logic. Do you need deskewing of documents? Click a button and you get deskewing. Which brings me to the second point. The products sold by ABBYY/Nuance are meant to be used by integrators (no programming needed other than the occasional VB.net script), not image processing specialists/developers. In my (biased) opinion, it makes more sense for some businesses to go the custom route instead of investing in FlexiCapture. There is also FlexiCapture Engine that is meant for developers. This has the same problems as the other offerings by ABBYY (I don't know about Nuance but I suspect it's the same): - expensive - vendor lock-in - ridiculous extra costs for things like "cloud/VM license", exporting to PDF, etc. - limits on how many pages you can process per year or in total (complex licensing schemes) - ABBYY really wants to sell you their own cluster/cloud management services which is all proprietary - limited flexibility in implementing distributed services, costs that add up fast, you have to be trained in their own stack Can you provide an example where you think that a custom solution would not work? I'm curious.