3 ms·
Handling PDFs and OCR is always troublesome. But in this case, if I understand you correctly, the original source documents were PDF? That they were distribut
by brc 11y ago
Handling PDFs and OCR is always troublesome.
But in this case, if I understand you correctly, the original source documents were PDF? That they were distributed as PDF for all the reasons you state?
That doesn't apply in this case - the original documents were not PDF, and had no reason to be PDF except to make it more difficult.
- halter73 11y agoFortunately, the PDFs already contain the plain text as metadata. I believe they are what are known as a searchable image PDFs. The code posted here isn't doing any OCR, but whatever generated the PDFs (Acrobat?) might have.
- brc 11y agoHow did they get the plain text as metadata? Was the scanning equipment doing OCR and setting that?
- danso 11y agoI believe the U.S. agencies use ABBYY FineReader, which does a pretty good job with OCR and text resolution. The U.S. Senate used it when releasing the CIA torture docs awhile ago: http://www.nytimes.com/interactive/2014/12/09/world/cia-torture-report-document.html?_r=0 http://www.nytimes.com/interactive/2014/12/09/world/cia-tort...
- spikels 11y agoAccording to news reports Hillary handed over the emails to the Dept. of State as hard copies. These were then processed by them into PDFs including doing OCR. http://money.cnn.com/2015/03/11/technology/security/hillary-email-paper/index.html http://money.cnn.com/2015/03/11/technology/security/hillary-... Edit: Hilarious part was where State says this is not a problem because "it would have to print Clinton's emails in the normal review process."
- A_COMPUTER 11y agoShe also handed over Outlook database files on a thumb drive to her lawyer, so it's hard to argue that this whole process was somehow necessary.