3 ms·
We do position based text extraction. We add however an 'unpaper' function which tries to correct misalignments and increases the quality of the scan.
by chezmo 10y ago
We do position based text extraction. We add however an 'unpaper' function which tries to correct misalignments and increases the quality of the scan.
- ComodoHacker 10y agoWhat OCR library do you use? What languages it supports?
- chezmo 10y agoFor scanned images we use https://github.com/tesseract-ocr/tesseract https://github.com/tesseract-ocr/tesseract. For text based PDFs we pull the text directly from the file and all languages are supported.