3 ms·
Image Scaling using Deep Convolutional Neural Networks
- msoad 11y agoThis is very cool! Next in DNN adventures should be a network trained with lots of videos for animating still pictures!
- hlfw0rd 11y agoYou might like this paper. Multi-view Face Detection Using Deep Convolutional Neural Networks : http://arxiv.org/abs/1502.02766 http://arxiv.org/abs/1502.02766
- cafebeen 11y agoInteresting stuff! It would certainly benefit from a comparison to other super-resolution techniques, e.g. Glasner et al. "Super-resolution from a single image" Freeman et al. "Example-based super-resolution"
- deepnet 11y agoJohn Resig used a Convnet to upscale japanese prints, waifu2x http://ejohn.org/blog/using-waifu2x-to-upscale-japanese-prints/ http://ejohn.org/blog/using-waifu2x-to-upscale-japanese-prin...
- darkmighty 11y agoNeat, specializing for Japanese prints must have improved the outcome.
- tn13 11y agoWhat would have made this article interesting was the examples of different images being scaled.
- buymorechuck 11y agoThere are several examples in the article that show original, bicubic and DeCNN upscaling side by side for comparison.
- rasz_pl 11y agothey arent exactly super convincing, I wouldnt be surprised if something like super2xsai + GS4xHqFilter beat those examples :/
- larsiusprime 11y agowaifu2x is based on the same basic principle as this, and it beats the pants off of those algorithms, for the kind of images its best at (waifu2x was designed for, and therefore trained on, anime/manga images)
- buymorechuck 11y agoFWIW, waifu2x was inspired by this paper from researchers at Chinese University of Hong Kong and Microsoft Research Asia http://arxiv.org/abs/1501.00092v3 http://arxiv.org/abs/1501.00092v3 waifu2x is a great demonstration of this approach applied to a specific domain. By coincidence, Flipboard's DNN approach was developed around the same time as the MSRA research in summer 2014. I'm excited to see future research in applying deep learning to generative tasks. Some of the CNN music composition work is quite impressive.
- JulianMorrison 11y agoWhat's intriguing about this is that the output isn't really real. The best place to see this is the bark patterns on the trees in the last 3-way comparison. The output is convincing and yet not quite right. The neural net didn't know, so it guessed plausibly. Keep scaling and I bet you'd see Google inceptionism style dream details slipping in.
- fredophile 11y agoI'll admit that I skimmed the article but I have the feeling this CNN didn't learn what they intended it to learn. Looking at the examples shown they started with a full resolution image and applied some downsampling algorithm to get the lower resolution to apply their algorithm to. Their algorithm has learned to undo the downsampling that they applied. This doesn't mean it will perform well on images that haven't been downsampled or images that have been downsampled in a different way.
- mining____ 11y agoThis is possibly true, but it's pretty much the only practical way to do this. If we look at most of the literature around upscaling, this method is used pretty frequently. For a more comprehensive look at using CNNs for image upscaling, see e.g. http://research.microsoft.com/en-us/um/people/kahe/publications/eccv14srcnn.pdf http://research.microsoft.com/en-us/um/people/kahe/publicati...
- buymorechuck 11y agoThere is a more recent version of this paper published here: http://arxiv.org/abs/1501.00092v3 http://arxiv.org/abs/1501.00092v3 At Flipboard, we did not have time to do a full comparison of related upscaling research, but we were happy with the low amount of error our CNN achieved.
- oh_sigh 11y agoIs it possible to differentiate a downsampled image from an image captured natively at a given resolution? My gut tells me no, but this certainly isn't my field.
- jfoutz 11y agoNot my field either. If a picture is just a big 2d array, then i don't think so, you're deleting bits, and information is lost. If the picture is a collection of summed sin waves, maybe. If the big picture is just sampling more frequently, then maybe it's cheating by looking at the encoding. the smaller resolution will have sampling problems, it'll lose higher frequency data, because it's not sampled enough. I dunno. I can see the op's point. Maybe there are artifacts introduced by scaling down. Still, regardless of the mechanism, information theory tells us there's no lossless compression. Information is lost, and the NN needs to make something up to fill in the blanks. Looks better than bicubic to me!