3 ms·
Fortunately, the PDFs already contain the plain text as metadata. I believe they are what are known as a searchable image PDFs. The code posted here isn't doin
by halter73 11y ago
Fortunately, the PDFs already contain the plain text as metadata. I believe they are what are known as a searchable image PDFs.
The code posted here isn't doing any OCR, but whatever generated the PDFs (Acrobat?) might have.
- brc 11y agoHow did they get the plain text as metadata? Was the scanning equipment doing OCR and setting that?
- danso 11y agoI believe the U.S. agencies use ABBYY FineReader, which does a pretty good job with OCR and text resolution. The U.S. Senate used it when releasing the CIA torture docs awhile ago: http://www.nytimes.com/interactive/2014/12/09/world/cia-torture-report-document.html?_r=0 http://www.nytimes.com/interactive/2014/12/09/world/cia-tort...