5 ms·
Ask HN: What is the best open source OCR software supporting multiple languages?
- mynewtb 10y agoTesseract
- pgodzin 10y agoI've used Tessarect with Tess4J Java wrappers, which has been pretty good.
- postila 10y agoNot that good for Russian. For English, it's not the best as well – too many mistakes for some fonts.
- pgodzin 10y agoSorry, thought you meant multiple programming language support. Yes, definitely ran into some font issues and noise turning '1' into 'L' or 'T', etc. As people have been saying though, it may not be great out of the box for you but you can train it on the font you want.
- deleted 10y ago[deleted]
- meeper16 10y agoI think this is a darn good question.
- deedubaya 10y agoSorry to hijack, but what about the best OCR service? I'd much rather farm the OCR work out to another service than trying to do it myself.
- danso 10y agoI bought ABBYY FineReader for Mac for abut $99. I find it to be pretty amazing. My new scanner also came with it and I generally expect it to do a reasonable text translation of whatever I throw it, whether it be newsprint articles or crumpled receipts. If you need to do OCR that also preserves table structure -- which is what I bought FineReader for in the first place, I don't think there's any open source alternative, and FineReader does a very capable job. Here's an example of FineReader in action: OCRing the docs released by the FBI on Clinton's email system. I've also included the pdftotext output showing how FineReader's text conversion also attempts to preserve the physical layout of the text characters: https://github.com/dannguyen/clinton-hillary-email-fbi-investigation-docs https://github.com/dannguyen/clinton-hillary-email-fbi-inves...
- walterbell 10y agoThey also have a Linux CLI for servers, it seems to be licensed by annual page volume, https://www.abbyy.com/ocr-sdk-linux https://www.abbyy.com/ocr-sdk-linux & http://www.ocr4linux.com/en:pricing http://www.ocr4linux.com/en:pricing
- postila 10y agothank you, I'll check it out. Know this company very well, we originated from the same university, but I've completely forgotten to check their new products now. Thanks.
- msandford 10y agoDo you want to OCR several human languages, or do you want bindings/libraries in several programming languages? The question as written is a little ambiguous.
- postila 10y agoI need to process millions of images and extract texts from them as better as possible. Primary language is Russian, but some texts are English. Also, interested in other languages (Spanish, German, etc) for future needs. What I've tried so far (including Tesseract) is either bad for Russian texts or cannot work with mixed texts (e.g.Russian with some English words). Or both. Programming languages/platform don't matter, but smth Linux-compatible is better of course.
- contingencies 10y agoTesseract[0] is a system that is broken in to different parts, at least one does layout analysis and another does the actual OCR. Output is a different layer again. I believe it is an open source adaptation of what Google used for its books project. The interface was less than polished a few years ago, to the point where getting it running at all was rather difficult. However, for multilingual work (including Chinese) it is probably ideal.[1] Note that if you are scanning books there are now some interesting open hardware systems appearing online that turn pages and take photos with cameras, so you can scan books - without cutting them up - to a high resolution. [0] https://github.com/tesseract-ocr/tesseract https://github.com/tesseract-ocr/tesseract [1] https://github.com/tesseract-ocr/langdata https://github.com/tesseract-ocr/langdata
- deleted 10y ago[deleted]
- kidsil 10y agoTesseract can give you nightmares, but unfortunately it's the only solid OCR library out there.
- postila 10y agoWhat kind of nightmares?
- titanix2 10y agoMy master thesis[1] was about OCRing a multilingual dictionary (Vietnamese-French using Nôm characters) and I got the following results, using built-in language model: - OCR for French is very good - OCR for Vietnamese (latin characters) was dispointing - OCR for Chinese indeed outout Chinese characters but not the good ones - OCR for Chinese didn't handle letters in Chinese script - no support for Vietnamese Nôm (not suprising) The good news is: - you can train new language model - the software is built on a (clean?) API you can use yourself - there is some usable documentaition to get started & also for training models - headers files are well commented So this enable one to build a solution on Tesseract, but this is by no way an out-of-the-box solution especially if you have very specific needs. All in one it is the best free (libre, 0$) OCR engine out there. [1] http://crim.fr/sites/default/files/memoire_m2_Lecailliez_final_v1.00.pdf http://crim.fr/sites/default/files/memoire_m2_Lecailliez_fin... [French]
- dogma1138 10y agoTesseract can (and has to be) trained, so it can effectively support anything. OCR isn't limited to language usually unless you are doing some really high end stuff when it does linguistic prediction but you only need that if you are working with really poor (image) quality sources. But overall OCR is "language" agnostic, it is however usually not type set agnostic so what you would want to do is train it for whatever fonts are common for a particular language. This gets slightly tricky if you have to do handwritten transcription or very stylized fonts but in those cases the "language" again is not an issue because your OCR program doesn't understand language to begin with.
- postila 10y agoThe best tool would be something that I can iteratively improve using some ML methods, that I would run on Linux and integrate into my programs. And open source, of course. I know, I want too much :)
- avmich 10y agoNo, not too much, at least I'd agree with you. A lot of information about current Tesseract is there - https://github.com/tesseract-ocr/docs/tree/master/das_tutorial2016 https://github.com/tesseract-ocr/docs/tree/master/das_tutori... . Tesseract is trainable, even though the bulk of capabilities came from algorithms designed well before deep learning became popular. One of slides mentions that it's puzzling that Tesseract is "winning" over modern ML attempts to solve OCR. However, latest developments - adding LSTM networks to Tesseract - are reported to be promising. Wonder when they'll become available on the Github...
- postila 10y agoThis looks interesting, thank you
- kondro 10y agoIs there anything great (even if potentially pricey) for ICR (individual handwritten characters, usually separated by boxes) or handwriting? Preferably as a service.
- frik 10y agoBeside Tesseract which was a state-of-the-art OCR software by HP in the early nineties and recovered by Google a few years ago and is open source. There is Cuneiform, a former main competitor to ABBYY Finereader. CuneiForm got open sourced a view years ago, though in a sad state (project files where in VS C++ 6 ('98), comments in Russian), but a community fixed that and ported it to Linux. It's also probably the best one for Russian language. It also has an UI and some advanced features that only ABBYY amd Cuneiform have, but non of the other competitors (certainly no other open spurce OCR package). https://en.wikipedia.org/wiki/CuneiForm_(software) https://en.wikipedia.org/wiki/CuneiForm_(software)
- acd 10y agoCaffee Deep learning possibly outperforms Tesseract. https://christopher5106.github.io/computer/vision/2015/09/14/comparing-tesseract-and-deep-learning-for-ocr-optical-character-recognition.html https://christopher5106.github.io/computer/vision/2015/09/14...