3 ms·
I can relate to the document scanning issue. In my case, I digitize all documents I get on paper. For this, I've written this wrapper around scanimage and ocrmy
by dbrgn 5y ago
I can relate to the document scanning issue. In my case, I digitize all documents I get on paper. For this, I've written this wrapper around scanimage and ocrmypdf: https://github.com/dbrgn/pydigitize https://github.com/dbrgn/pydigitize
It does the following steps:
1. Scan a document with any scanner that supports SANE (ADF supported), 2. straightening and cleaning of scanned documents, 3. run OCR on PDF so that it becomes searchable, 4. generate PDF/A file for archival, 5. add keywords to the PDF file
I've probably saved many hours with this script, even when taking into account the time it took me to write it.
Tiny projects don't always save time, but they sure are gratifying when they work as intended!
- chrisweekly 5y agoCool project! Related tangent: I haven't looked into OCR, but for a simple "iPhone camera to 'scanned' PDF", the Dropbox iOS app has a surprisingly good implementation.
- dbrgn 5y agoI prefer not to upload all my potentially private documents into a cloud service :) OCRmyPDF (https://ocrmypdf.readthedocs.io/ https://ocrmypdf.readthedocs.io/) actually does a pretty good job! It also handles deskewing and all other stuff that's necessary for good OCR results.
- chrisweekly 5y agoWell yeah, that's legit. My point was that for those of us who already choose to rely on Dropbox, it's a great little feature to be aware of.
- jwong_ 5y agoI use the "Files" application from Apple, and it works pretty well too. I was surprised, and have largely made it my ingress for receipts/paperwork.
- mindslight 5y agoI wrote a keypress-driven graphical utility that's basically a wrapper around scanadf, that allows me to call scanadf repeatedly, preview, delete pages, etc. For instance if a page misfeeds, I pull the remaining stack out, delete the bad scan, and restart from where it went wrong. When finished, the graphical window disappears, and it writes all the current pages out as an archive - the masters are checksummed, compressed with FLIF, converted to some low quality JPGs, and the whole thing is stuck in a ZIP archive with extension .cbz (viewable with evince). The eventual goal is to transcode all these masters into nicer OCRed PDFs, but I've been making do with the low quality JPGs just fine. Your script seems like a great starting point to actually get this done!
- zem 5y agoher script reminded me that i never did get my scansnap working properly under linux :( will have to give it another go sometime.
- BeetleB 5y agoI'll try your version out. I went through this pain when I bought a document scanner some years ago - there was no good solution on Linux that would let me from the command line scan, clean, OCR, and output a PDF. I found lots of scripts like yours, but none that was complete. I finally took an existing Perl script and hacked it to my needs. Features one should have: 1. Output to PDF. 2. Option to OCR (optional) 3. Clean up (e.g. skip blank pages, straighten pages, etc) 4. Allow one to specify quality/dpi 5. Select grayscale vs color 6. Duplex vs single page 7. Dynamically recognize the size of the page. The last one is the one I'm missing - if I scan something long, it trims it to fit a Letter size page.
- cranium 5y agoFunny, I also wrote some code to solve one of my printer problems. It doesn't do double-sided scanning so the script waits for two recently PDFs in a folder (while trying to be smart about not picking incompatible PDFs) and merge them in the right way: https://github.com/RomainGehrig/PDFCollate https://github.com/RomainGehrig/PDFCollate I could have worked the same amount of hours for a client and use the money to buy a duplex scanner, but where's the fun in that ?