3 ms·
Scanned PDFs only work well if they already have an OCR layer. There's some optional integration of rga with tesseract, but it's pretty slow and less good than
by phiresky 6y ago
Scanned PDFs only work well if they already have an OCR layer. There's some optional integration of rga with tesseract, but it's pretty slow and less good than external OCR tools.
ripgrep-all can do the same regexes as rg on any filetypes it supports. So you can could do something like --multiline and foo(\w+[\s\n]+){,20}bar
It won't work exactly like this, but something similar should do it:
--multiline enables multiline matching
* foo searches for foo
* \w+ searches for at least one word character
* [\W]+ searches for at least one space/nonword character like sentence marks
* {,20} searches for at most 20 iterations of the word-space combination
bar searches for bar