4 ms·
Around couple of years ago I am working on a home project and utilised Tesseract and Laptonica for OCR. Storage and search HDFS, HBase and SolrCloud on extracte
by mkjmkumar 7y ago
Around couple of years ago I am working on a home project and utilised Tesseract and Laptonica for OCR. Storage and search HDFS, HBase and SolrCloud on extracted text. You can find the details here on my website. I was very impressed with conversion of hand written pdf docs with 90% readable accuracy. I have named it as Content Data Store(CDS) http://ammozon.co.in/headtohead/?p=153 http://ammozon.co.in/headtohead/?p=153 . Source code is open and you may find steps on installation and how to run here.
http://ammozon.co.in/headtohead/?p=129 http://ammozon.co.in/headtohead/?p=129
http://ammozon.co.in/headtohead/?p=126 http://ammozon.co.in/headtohead/?p=126
A short demo
http://ammozon.co.in/gif/ocr.gif http://ammozon.co.in/gif/ocr.gif
I didnot get time to enhance it further but planning to containerize the whole application. See if you find it useful in its current form.
- hylian 7y agoI had a similar problem and ended up using AWS' Textract tool to return the text as well as bounding box data for each letter, then overlayed that on a UI with an SVG of the original page, allowing the user to highlight handwritten and typed text. I plan to open source it so if anyone's interested let me know. Not a fan of the potential vendor lock in though, so it's only really suitable for those in an already AWS environment not worried about them harvesting your data.
- rwojo 7y agoVery interested to see this as I was about to work on the same thing!