4 ms·
Nice idea, would it be possible to do some OCR at the end?
by Synergyse 13y ago
Nice idea, would it be possible to do some OCR at the end?
- contingencies 13y agoTake a look at Tesseract, which last time I looked was the best codebase to use in this area. It's part of Google's open source multilingual OCR suite, which is in two parts (layout analysis and actual OCR code being segregated): https://code.google.com/p/tesseract-ocr/ https://code.google.com/p/tesseract-ocr/
- arjie 13y agoIs there a guide to using tesseract? When I last used it, I had trouble accurately recovering text from an image I just created with Times New Roman. There are probably some settings I'd need to change to get it to work properly.
- contingencies 13y agoNot sure. I looked at it years ago, during its early open source stages, and similarly concluded that it was nontrivial to get going. Should be easier now. IIRC in those days it had a very high volume mailing list for user support.
- thejosh 13y agoAt the end the script needs to upload the image to mechanical turk for analysis.
- lelandbatey 13y agoI did experiment with using this to bring out the text in photos of books. Here's an example: Input image - http://i.imgur.com/6o5FwxG.jpg http://i.imgur.com/6o5FwxG.jpg Output image - http://i.imgur.com/7OIOxfO.png http://i.imgur.com/7OIOxfO.png I tried to use vanilla Tesseract on it, but I had no luck getting anything usable out of it.
- jbondeson 13y agoYou have to compensate for the 3D deformation present in the captured image. There has been some impressive work in this area recently. Some commercial OCR engines such as Nuance Capture SDK have built in functionality for this.