3 ms·
For scanned images we use https://github.com/tesseract-ocr/tesseract https://github.com/tesseract-ocr/tesseract. For text based PDFs we pull the text directly f
by chezmo 10y ago
For scanned images we use https://github.com/tesseract-ocr/tesseract https://github.com/tesseract-ocr/tesseract. For text based PDFs we pull the text directly from the file and all languages are supported.