4 ms·
ABBYY has dominated the field for many years (decades really) and still outperforms every solution out there. OmniPage by Nuance is probably the second best. P
by ocrcustomserver 9y ago
ABBYY has dominated the field for many years (decades really) and still outperforms every solution out there. OmniPage by Nuance is probably the second best.
Preprocessing the images (OCR pipeline) is very important for OCR. For generic scanned PDF documents Finereader does a pretty good job.
There is a lot of stuff going on in a OCR engine. Layout analysis, dewarping, binarization, deskewing, despeckling (and others) and then there's the OCR itself.
With Tesseract you have to do a lot of things yourself, you have to provide it with a clean image. The commercial packages do that for you automatically. ABBYY and other solutions also use NLP to augment/check the OCR results from a semantic analysis perspective.
Also, there is no "one size fits all" OCR. It is highly specific to the nature of the application. Consider the following use cases:
- scanned PDF document
- scanned document with a non-standard font (e.g. Fraktur script in a historic book)
- photo of scanned document acquired with a mobile phone's camera
- passport OCR (MRZ)
- credit card OCR
- text appearing in natural image (e.g. store sign)
These are all "OCR projects" but they require very different approaches. You cannot just throw any input image at an OCR engine and expect it to work. It often requires a mix of computer vision/image processing, machine learning and OCR engine.
There is a growing number of papers using deep learning that get submitted to ICDAR (the premier OCR conference) and the other OCR conferences.
One of the problems is the lack of a universal dataset/competition like ImageNet.
The SmartDoc competition (documents captured from smartphones) was cancelled this year due to an insufficient number of participants.
If anyone is doing work with OCR + deep learning, I'd love to discuss!