3 ms·
To extract text from photos and non OCR-ed PDFs Tesseract[1] with language specific model[2] never fails me. I use my shell utility[3] to automate the workflow
by undebuggable 6y ago
To extract text from photos and non OCR-ed PDFs Tesseract[1] with language specific model[2] never fails me.
I use my shell utility[3] to automate the workflow with ImageMagick and Tesseract, with intermediate step using monochrome TIFFs. Extracting each page into separate text file allows to ag/grep a phrase and then find it easily back in the original PDF.
Having greppable libraries of books on various domains and not having to crawl through the web search each time is very useful and time-saving.
[1] https://tesseract-ocr.github.io/ https://tesseract-ocr.github.io/
[2] https://github.com/tesseract-ocr/tessdata https://github.com/tesseract-ocr/tessdata
[3] https://github.com/undebuggable/pdf2txt https://github.com/undebuggable/pdf2txt
- jdc 6y agoIf you're interested in grepping PDFs (among other formats) another option is ripgrep-all. https://github.com/phiresky/ripgrep-all https://github.com/phiresky/ripgrep-all
- for_your_info 6y agoYour bash script is totally broken. It doesn't properly parse command line arguments. Ignores the page range set on -t -f
- undebuggable 6y agoThis should work now, thanks.
- Jaruzel 6y agoI struggled to get tesseract to OCR my image based PDFs directly, so resorted to using GhostScript to extract the pages to pngs which I then put through tesseract. As an added bonus though, I gained the ability to have a thumbnail png for the search front end.
- bufferoverflow 6y agoI tried using Tesseract to OCR just some numbers on an almost plain background, and it failed around 2-3% of the time. Which made the whole thing useless, because I needed 100% correctness.
- undebuggable 6y ago> so resorted to using GhostScript to extract the pages to pngs which I then put through tesseract I understand you use this to extract text from non OCR-ed PDFs, especially consisting of low quality scans or photos (e.g. low resolution, JPEG artifacts). Ocassionally passing higher resolution to ImageMagick when converting a page to TIFF helped, but this sounds like a reasonable fallback as well.