3 ms·
Care to share resources/lessons learned for training tesseract with custom data? I'm using it for a side project and would love to hear about your insights.
by dclusin 6y ago
Care to share resources/lessons learned for training tesseract with custom data? I'm using it for a side project and would love to hear about your insights.
- sireat 6y agoI followed the resources here: https://github.com/tesseract-ocr/tessdoc/blob/master/TrainingTesseract-4.00.md https://github.com/tesseract-ocr/tessdoc/blob/master/Trainin... Also this: https://github.com/UB-Mannheim/tesseract/wiki https://github.com/UB-Mannheim/tesseract/wiki The original data was here: https://github.com/tesseract-ocr/langdata_lstm https://github.com/tesseract-ocr/langdata_lstm I did use another data source from Manheim but can't locate it right now. Using vanilla Ubuntu 18.04 I looked at the example training files and made a small script to convert my own labeled data to fit the format that tesseract requires. I did do a bit of pre-processing adjusting contrast. All the data munging was done on Python (Pillow for image processing, Flask for collecting data into a simple SQLite DB before converting back to format that Tesseract requires). Python was not necessary just something that felt most comfortable to me. I am sure someone could do it using bash scripts or node.js or anything else. EDIT: To make life easier for my curators I did run Tesseract first to generate prelabeled data for my training set. It was about 90% accurate to start with. So the process was: Tesseract OCR on some documents to be trained -> hand curation (2 months)-> train (took about 12 hours) -> 99% (on completely separate test set)