21 ms·
Compressing and enhancing hand-written notes (2016)
- BlackLotus89 9y agoLooks interesting. Normally when I'm "cleaning" up scans I use unpaper, but although there is some overlap in functionality it doesn't do the same. Anyway very nice writeup and I will add it to my arsenal and give it a closer look later. Could be useful for my document archive+ocr solution. Edit: too bad seems like it didn't see any activity in the last year
- eadmund 9y ago> Edit: too bad seems like it didn't see any activity in the last year That's not necessarily bad: sometimes a piece of software can be done, or nearly so.
- BlackLotus89 9y agoYup but a project like this would have an empty issue tracker this one not so much ;) (which doesn't mean that it's a bad project or that I wont use it. it means that I will probably start to work on it)
- foleac 9y agoThere actually was some activity in three different branches in january: https://github.com/kskyten/noteshrink/network https://github.com/kskyten/noteshrink/network
- BlackLotus89 9y agoYeah on a branch of a fork.
- krick 9y agoCould you expand on your archive+ocr? I long wanted to start doing something like this, but never got to. I guess reading others' experience can be useful.
- mkjmkumar 9y agoI have made some progress on this as my home project using same compression and scan. I call it DFA - digital file analytics where data/images/scanned documents are sent remotely using Kafka to Hadoop and then run OCR to extract text and compression. If the document is more then 10MB go to HBase otherwise HDFS. Near real-time streaming using Spark and Flink is done too. Visualization using Banana dashboard is not so cool as it shows word counts, storage location, images and tags. Analytics on top of extracted data using ML would like to do next. More you can find at https://medium.com/@mukeshkumar_46704/digital-files-ingestion-platform-dfip-a-real-time-data-ingestion-platform-163f25f87dd https://medium.com/@mukeshkumar_46704/digital-files-ingestio...
- pjc50 9y agoNice to see a bit of k-means clustering. I was worried that this might attempt to be "smart" by converting to symbols, replicating the "Xerox changes numbers in copied documents" bug, but it's pure pixel image processing. Very clean results. In some ways it's a smarter version of the "posterize" feature.
- deleted 9y ago[deleted]
- wmu 9y agoHere is the link to the blogpost that describes the problem: http://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_are_switching_written_numbers_when_scanning http://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_...?
- anc84 9y agoI wonder if it might improve by using a better color space than HSV. Maybe CieLAB?
- jerf 9y agoIs there much room for improvement? Looks pretty good to me. It seems to me that the inaccuracies/inefficiencies/errors/whatever you like in using RGB are basically truncated out of existence by the very, very harsh binning that is occurring. I wouldn't expect any visible differences to emerge from any alternate color space.
- trurl42 9y agoThis reminds me of my time in university, when I saved all my lecture notes as DjVu [1] files. It's a great file format for space-efficient archiving of scans like that, with a bit of scripted preprocessing. [1]: https://en.wikipedia.org/wiki/DjVu https://en.wikipedia.org/wiki/DjVu
- dunham 9y agoI like the idea, but DjVu seems to be very proprietary / single vendor and not in widespread use. This has made me reluctant to use it for archival purposes (vs say PDF, which has its own issues, but feels slightly more future proof to me). I think PDF can cover pretty much the same ground with JBig and Jpeg2k. (And I believe archive.org is doing that.) But I don't know of any open source code to do the segmentation / encoding. (You have to split the bitmap from the background for jbig / jpeg encoding.)
- pwg 9y agoThere is an open-source DjVu library as well: http://djvu.sourceforge.net/ http://djvu.sourceforge.net/ Whether that makes the format "widespread" enough for your use case is of course your decision to make.
- mcguire 9y agoThe major source for DjVu files I have run across is the Internet Archive's book scanning. (Weirdly, I can't find any examples.) They're usually smaller than PDFs.
- jonathanyc 9y agoThe Image Capture app on macOS is surprisingly good at this, and one of the things I really miss on Linux and Windows, so this is neat to have. It’s also interesting for me to think about how this is a generalization of converting a scan to black and white for clarity :)
- dontyouremember 9y agoCan be incrementally improved by using a more human-focused color model than HSV, like CIECAM02 or CIELAB.
- davidzweig 9y agoAlso, see ScanTailor by Joseph Artsimovich. Excellent tool.
- Softcadbury 9y agoI wonder if your technique could remove some lines for the paper we use in France [1]. I never really understood why they were so many lines... [1]: https://images-na.ssl-images-amazon.com/images/I/815WQQdAHBL._SL1500_.jpg https://images-na.ssl-images-amazon.com/images/I/815WQQdAHBL...
- John_KZ 9y agoIs this really standard writing paper? I assume it would be useful for calligraphy or learning how to write (as you can use the subdivision to draw letters to the correct height) but I find it weird for it to be standard issue paper.
- sp332 9y agoI hadn't seen it before, but it seems to be pretty common for French-ruled paper. https://getfrenchbox.com/what-is-french-ruled-paper-seyes-school/ https://getfrenchbox.com/what-is-french-ruled-paper-seyes-sc...
- pfalke 9y agoIt is. It's called "French-ruled paper" and is the school standard in France. [0] [0] https://getfrenchbox.com/what-is-french-ruled-paper-seyes-school/ https://getfrenchbox.com/what-is-french-ruled-paper-seyes-sc...
- bb88 9y agoMy handwriting in school would have been better if I could have used this.
- l9k 9y agoIt helps children learning how to write. Lowercase letters start from the thicker line to the first thinner line above. Uppercase letters and taller lowercase letters like "t" or "d" go to the second line. And the tails of letters like "g" or "y" go the first line below.
- neelkadia 9y agoThis is GEM!
- alister 9y agoI have an observation about scanning documents that results in good quality and smaller files, but I can't satisfactorily explain why it works. Consider these two cases: (1) Scan document at very high resolution as a JPG and then use a third-party program (like Photoshop or whatever) to re-encode the JPG at your preferred low resolution. (2) Scan document at your preferred low resolution as a JPG straight away. Don't re-encode afterward. Intuition says that the results of #1 vs #2 should be identical, or that #1 should be worse because you're doing two passes on source material. But I always get better results with case #1 (i.e., high-res scan and re-encoding afterward) regardless of the type or model of scanner, or whether the scanner does the JPG encoding on-board the device itself or through a Windows/Linux/Mac driver bundled with the scanner. My theory is that scanner manufacturers are deliberately choosing the JPG encoding profile that gets them the fastest result. They want to brag about pages per minute which is an easily measured metric. Quality of JPG encoding and file size take effort to compare, but everyone understands pages per minute. If anyone has contrary experience I'd like to hear it. I've been seeing this for years with different document scanners and flatbed scanners -- regardless of how I tweak the scanner's settings, I can always get good quality in a small file by re-encoding afterward.
- dtech 9y ago> My theory is that scanner manufacturers are deliberately choosing the JPG encoding profile that gets them the fastest result. This is more-or-less correct. The chips in the printers have a lot less power than your CPU, and the algorithms are a lot worse than those in Photoshop.
- derefr 9y agoI'm surprised there aren't any high-quality JPEG encoder ASICs. Would scanners benefit from just using some of the plentiful, cheap, excellent-quality H.264 encoder ASICs—and treating the output as HEIF?
- deleted 9y ago[deleted]
- keenerd 9y agoA simpler way of achieving the same thing is to duplicate the layer, blur the top layer heavily, and then set it to "divide".
- donquichotte 9y agoReally? Do you care to explain? What is the dividend and what is the divisor? Why can dividing a image by its low pass filtered version (or vice versa) be used to "clean up" the image, i.e. subtract the background, find main colors and cluster similar colors with k-means? What if the divisor has pixels near zero?
- deleted 9y ago[deleted]
- keenerd 9y agoAreas of low contrast become whiter and areas of high contrast become more saturated. It is also more robust than k-means. The author's algo will only work on scanned images. Photographed pages from a book will often have a slight shadow on half the page from the curvature. Blur-divide will clean this up. K-means will think you've used a lot of gray and not figure out that there are multiple background colors.
- 333c 9y agoI can confirm that the author's approach doesn't work well for photographed pages. I took a photograph[0] of a page of notes, and due to the shadow, the results[1] were very unsatisfactory. [0]: https://i.imgur.com/CLZHshT.jpg https://i.imgur.com/CLZHshT.jpg [1]: https://i.imgur.com/rrwca0m.jpg https://i.imgur.com/rrwca0m.jpg
- tkp 9y agointeresting trick, thanks for sharing ! [edit] Quick test here : https://imgur.com/a/6xOz1 https://imgur.com/a/6xOz1
- 9y ago
- imrankhan2601 9y agokeep the good work up
- reaperducer 9y agoGreat job with that. I've only just started taking notes by hand once again, after being keyboard-only for many years. In your scenario, since you have assigned "scribes" taking the notes, you might be able to streamline the process with a "smart pen." There are several on the market. The one I got as a hand-me-down from a family member lets you write dozens of pages of notes, then Bluetooth them to a smartphone app that exports to PDF, Box, Google Drive, etc... Or it can actually copy the notes to the app in real time. Combined with a projector, this might be useful for the other students during class. It's supposed to be able to OCR the notes, too, but I haven't bothered to figure out how. But there's a cool little envelope icon in the corner of each notebook page that if you put a checkmark on, it will automatically e-mail the page to a pre-designated address. Again, there are several models on the market. Mine retails for about $100. Notebooks come in about 15 different sizes and cost about the same as a regular quality notebook. Just some thoughts.
- inetknght 9y agoI have found that my Galaxy Note 2014 is pretty much hands-down the best note taking tablet in my opinion. It's better than the crap that Microsoft and Apple are trying to hawk off. It doesn't have as many fancy apps but for _strictly_ note taking, sharing notes via email, and book reading, it's pretty awesome. I just wish its price would come down. It's still full price from four years ago :| and even getting more expensive because it's so old
- rripken 9y agoYou are referring to the 10.1 tablet? I know its not the same but I had a Note4 and was pleased with its note taking ability. I imagine that with the extra screen real-estate the tablet was even better. It looks like they are $453 on Newegg! I agree that does seem crazy for an old device. If you can put up with a used device there are two on swappa for around $200 https://swappa.com/buy/samsung-galaxy-note-101-2014-wifi https://swappa.com/buy/samsung-galaxy-note-101-2014-wifi If it truly is the hands-down best note taking tablet you might as well buy them both and standardize to the platform.
- kazinator 9y agoI can get seemingly comparable results with a couple of simple operations in Gimp. Here is a casual job on the first image: https://i.imgur.com/Sy2rvsU.png https://i.imgur.com/Sy2rvsU.png The steps: 1. Duplicate the layer. 2. Gaussian-blur the top layer with big radius, 30+. 3. Put the top layer in "Divide" mode. Now the image is level. 4. Merge the layers together into one. 5. Use Color->Curves to clean away the writing bleeding through from the opposite side of the paper. 6. To approximate the blurred look of Matt Zucker's result, apply Gaussian blur with r=0.8. Notes: The unblurred image before step 6 is here: https://i.imgur.com/RbWSUnD.png https://i.imgur.com/RbWSUnD.png Here is approximately the curve used in step 5: https://i.imgur.com/lvfqCNK.png https://i.imgur.com/lvfqCNK.png I suspect Matt worked at a higher resolution; i.e. the posted images are not the original resolution scans or snapshots.
- fouc 9y agoYou can get the higher resolution one here: https://github.com/mzucker/noteshrink/blob/master/examples/notesA1.jpg https://github.com/mzucker/noteshrink/blob/master/examples/n... BTW I'm curious how you'd fare on the graph one. I didn't like his results for it. https://github.com/mzucker/noteshrink/blob/master/examples/graph-paper-ink-only.jpg https://github.com/mzucker/noteshrink/blob/master/examples/g...
- deleted 9y ago[deleted]
- kazinator 9y agohttps://imgur.com/a/dsLhk https://imgur.com/a/dsLhk Note how the grid is completely gone, the Sharpie strokes are fuller and the ghosting around the red ink is gone. (The word "Red" seems to have been written faintly, like with a non-working ball point pen, and then written over properly.) The thing is, I took a completely different approach here. I won't give a complete step-by-step recipe, but the gist of it is this: 1. Create a copy layer of the image. 2. Optionally level the intensity with the divide trick; I didn't bother. 3. Convert this copy to grayscale. 4. Threshold it to black and white, such that the grid is eliminated, but the writing remains solid. 5. Blur the writing (radius 3-4). 6. Threshold again. 7. Now you have a black and white version that is a bit thicker than the original. TURN THIS INTO A LAYER MASK. An inverted one which passes through the writing, and renders everything else transparent. 8. Apply this mask to the original image. This requires transfering a layer mask between layers. 9. Slide a white background under the masked layer. Now you have the lettering clean on white. 10. Play with simple Color->Brightness-Contrast. I ended up with something like brightness -66, contrast +88. In the final step, because of the layer mask that is in effect, these controls affect only the writing: the white coming from the unaffected layer below stays white no matter what you do with the contrast and brightness controls. Why the different approach: I first tried the original approach and the result was good. But I thought you wouldn't like it either. It was similar to Matt's. I did a better job of eliminating the grid, but the writing was less vivid. (Likewise I also preserved the yellow tint of the paper.) I wanted the grid completely gone, with vivid writing. Playing with the intensity transfer curves was not quite doing it; there was poor separation between vanshing the grid while preserving the ink. This attempt can be seen here: https://imgur.com/a/ldrBN https://imgur.com/a/ldrBN The green writing is particularly unsatisfactory.
- dracodoc 9y agoI used to use a free software "ComicEnhancerPro" (The author is Chinese, there is English version but may not easy to find reliable download site) specially designed to enhance scanned comics. You can remove the background very effectively by dragging a curve with preview. You almost always need to preview and adjust some parameters, unless you have a template for similar cases.
- nayuki 9y agoOn the top image, I see that the back side of the page has clearly leaked through. In my experiences with scanning paper, I found a trick that essentially eliminates any visible backside content: Using a flatbed scanner, I would scan with the lid open, and the room darkened. The worst thing to do is to scan with the lid closed, with a lid that has a white background. This would increase the reflection from the backside of the page.
- MagerValp 9y agoYou achieve the same effect with a black paper on top of the document you’re scanning, or in between the pages if it’s a book. As a bonus you can leave the light on :)
- goerz 9y agoIn terms of compression for scanned notes, I haven't found anything that comes close to what even an older version of Adobe Acrobat yields, due to the use of the JBIG2 codec. Has anybody found any way to compress PDF files with JBIG2 on Linux/Mac? It's pretty much the only reason I have to find a Windows machine with Acrobat installed a couple of times a year, to postprocess a batch of scanned PDFs.
- ramses0 9y ago`German and Swiss regulators have subsequently (in 2015) disallowed the JBIG2 encoding in archival documents.[19]`
- goerz 9y agoYeah, but I'm not archiving for the German or Swiss government. For scanned handwritten notes, JBIG2 still beats out anything else by an order of magnitude at least
- haikuginger 9y agoI wonder if it'd be possible to do automatic detection and removal of notebook lines via an FFT (frequency domain) transform.
- freecodyx 9y agointeresting, i was working on something similar, (get color palette from an image).
- herpaderp_33 9y agoIt would be interesting to see this paired with potrace.
- Myrmornis 9y agoThis is awesome, and a depressingly large factor better than any blog post I’ll ever write. I totally identify with the need for this. I also want to archive images of notes and whiteboards, and they must be kept small as so far my life fits in google drive and github. Currently I use Evernote to do this. I don’t use any other functionality in Evernote but the “take photo” action does processing and size reduction very like the blog post.
- andimai 9y agoHow does this method compare to adaptive thresholding or Otsu's binarization method? https://docs.opencv.org/3.4.0/d7/d4d/tutorial_py_thresholding.html https://docs.opencv.org/3.4.0/d7/d4d/tutorial_py_thresholdin...
- amai 9y agoA lot of apps out there which can do this for you: http://uk.pcmag.com/cloud-services/86200/guide/the-best-mobile-scanning-apps-of-2018 http://uk.pcmag.com/cloud-services/86200/guide/the-best-mobi...
- krsree 9y agoThe link says that generated pdf is a container for the png or jpg image. Is it possible to get a true pdf from the scan? Specifically so that i can search inside the pdf.
- eltoozero 9y agoFor anyone having issues getting this to work on macOS with homebrew dependencies, I was able to get it to work after finally getting an old version of numpy installed using the following command. sudo pip install --upgrade --ignore-installed --install-option '--install-data=/usr/local' numpy==1.9.0 If you don't use the numpy==1.9.0 you'll get the 1.14.2 version which is also broke. The rest of the options allow pip to soft-override the macOS built-in numpy 1.8.0 which is immutable in the /System/ directory. Anyway, after I did all that I was able to start playing with the app, I had previously been using a kludge workflow to get a nice output in black and white by using the imagemagick convert -shave option to remove the scanned edges of images, then doing a -depth 1 to force the depth down (which only works well on really clean scans), then I can -trim to clear the framing white pixels and re-center using the -gravity center -extent 5100x6600 to frame the contents centered inside a 600dpi image. Rough but it works, I was hassling with trying to isolate "spot colors" for another thing, but this might actually do the trick!!!