2 ms·
> so resorted to using GhostScript to extract the pages to pngs which I then put through tesseract I understand you use this to extract text from non OCR-ed PD
by undebuggable 6y ago
> so resorted to using GhostScript to extract the pages to pngs which I then put through tesseract
I understand you use this to extract text from non OCR-ed PDFs, especially consisting of low quality scans or photos (e.g. low resolution, JPEG artifacts).
Ocassionally passing higher resolution to ImageMagick when converting a page to TIFF helped, but this sounds like a reasonable fallback as well.