12 ms·
Colorizing black and white photos with deep learning
- iluvmylife 9y agoThe averaging problem in colorization is interesting. If it learns that an apple can be red, green and even yellow - how does it know how to color it? A HN user in an earlier thread suggested to use a fake/real colorization classifiers as a loss function. [1] But I still feel that it would not solve the averaging problem. It would hop between different colors and probably converge to brown. I haven’t come across a plausible solution so far. [1] https://news.ycombinator.com/item?id=10864801 https://news.ycombinator.com/item?id=10864801
- emilwallner 9y agoIt could try to classify the apple tree or the context, but it would require a lot of training data. If it's out of context, it should select a color based on probability. But it's hard to solve this with just input and output data. The simple solution is to use noncontradictory training data, i.e. only having green apples. I have an urge to teach it simple logic. Instead of making it brown, it selects the color with the highest probability from a range of colors. However, I haven't come across a deep learning implementation like this to mimic.
- zardo 9y ago>But I still feel that it would not solve the averaging problem. It would hop between different colors and probably converge to brown. At least to the extent that GANs work, it works. They will alternate between the observed colours based on the noise vector. They do not simply converge to averages, because the discriminator easily recognizes brown apples as fakes.
- taneq 9y ago> If it learns that an apple can be red, green and even yellow - how does it know how to color it? I dunno, does it look more like a red apple, a green apple or a yellow apple?
- fizx 9y agoGoogle "learning xor".
- matt4077 9y agoYou could adjust the error function. The common Root Mean Square error pushes predictions to the average. If you use absolute errors, or even a logistic function instead, you'll encourage the model to commit to a decision on a multimodal distribution. Alternatively, use a discrete colour space and consider colours as categorical data not implying any ordinal scale.
- zan2434 9y agoAlthough it is quite egregious here - this is not a problem inherent to colorization but rather to generative models in general. Using something akin to a variational autoencoder would solve this problem, because it learns a distributional approximation rather than a single point estimate of the color, and then the random noise vector input allows one to sample from this output distribution. Similarly, Mixture Density Networks allow you to model a distribution and then sample from it.
- terrabytes 9y agoThis is refreshing. I’ve been learning machine learning through Kaggle. Recently and I’m a bit tired with the “tuning hyperparameter” culture. It rewards people that have the pockets to spend on computing power and the time to try every parameter. I’m starting to find problems that don’t have a simple accuracy metric more interesting. It forces me to understand the problem and think in new ways, instead of going down a checklist of optimizations.
- wrangler99 9y agoHa, it reminds me of what Andrej Karpathy said "Kaggle competitions need some kind of complexity/compute penalty. I imagine I must be at least the millionth person who has said this." It would be interesting to collaborate/compete on more creative tasks and have different metrics for success. [1] https://twitter.com/karpathy/status/913619934575390720 https://twitter.com/karpathy/status/913619934575390720
- ReDeiPirati 9y agoSo true. Another reason to put constraints in Kaggle competition is due to production environment. How many winner models have been used in production? I suspect this number is near zero. High accuracy with a delayed time makes a ML/DL artefact not usable in production, because from users point of view speed is much more valuable than the difference between 97% and 98% in accuracy.
- emilwallner 9y agoI'm also starting to follow people and communities that work with deep learning in new ways. Here are some of my favorites: [1] http://colah.github.io/ http://colah.github.io/ [2] https://iamtrask.github.io/ https://iamtrask.github.io/ [3] https://distill.pub https://distill.pub [4] https://experiments.withgoogle.com/ai https://experiments.withgoogle.com/ai
- perturbation 9y agoYou can be a little less brute force if you use something like hyperopt (http://hyperopt.github.io/hyperopt/ http://hyperopt.github.io/hyperopt/) or hyperband (https://github.com/zygmuntz/hyperband https://github.com/zygmuntz/hyperband) for tuning hyperparameters (Bayesian and multi-armed bandit optimization, respectively). If you're more comfortable with R, then caret supports some of these types of techniques as well, and mlr has a model-based optimization (https://github.com/mlr-org/mlrMBO https://github.com/mlr-org/mlrMBO) package as well. These types of techniques should let you explore the hyperparameter space much more quickly (and cheaply!), but I agree - having money to burn on EC2 (or access to powerful GPUs) will still be a major factor in tuning models.
- flsantos 9y agoInteresting! If nowadays, pictures are colorized by hand in photoshop, it wouldn't be practical to colorize a full black and white movie. I guess this deep learning approach would solve this problem and colorize old black and white classic movies.
- zimpenfish 9y agoI think colorising b&w movies must already be fairly practical given the size of this list: https://en.wikipedia.org/wiki/List_of_black-and-white_films_that_have_been_colorized https://en.wikipedia.org/wiki/List_of_black-and-white_films_...
- saip 9y agoAgreed. I imagine this has applications in compression as well. You could stream a movie (or a football game) in black and white and enable each device to color it on the spot. A similar technique could also be done for HD/3D/VR.
- mholmes680 9y agothat is an amazing idea.
- zardo 9y agoYes, you provide a handful of full data keyframes and reconstruct the details of the stream from the middle out.
- kiliankoe 9y agoThat middle out compression has some fantastic Weissman scores I believe.
- roywiggins 9y agoColoring football uniforms might be nearly impossible though...
- tzahola 9y ago
- ReDeiPirati 9y agoIt could be really interesting if it could return different coloured versions and provide a way to explore this different style.
- houqp 9y agoWhat's even cooler is adding support for human annotations so users can selectively give colorization hints for different parts of the image to customize the output.
- saip 9y agoA GAN style approach to learning and generating variants could be interesting as well. It could generate a couple of hundred plausible versions. Then you have another network that is trained to differentiate between fake and real colored photos which picks the best version.
- whatrocks 9y agoI'd like to train this on color comic strips and then run something traditionally black and white like xkcd through it. Seems like it could make the colorization part of hand drawn animation much easier.
- web007 9y agoYou'll probably want something closer to a GAN like pix2pix - https://phillipi.github.io/pix2pix/ https://phillipi.github.io/pix2pix/ An example implementation would look something like edges2cats https://affinelayer.com/pixsrv/ https://affinelayer.com/pixsrv/
- whatrocks 9y agoedges2cats is already way too fun. Thank you for the link!
- doppenhe 9y agocolorization applied to video: http://demos.algorithmia.com/video-toolbox/ http://demos.algorithmia.com/video-toolbox/
- bootcat 9y agoThank you, really good read and great product idea !!
- vwcx 9y agoAs a professional photo editor and historian, colorized photos really agitate me. I'm all for the creation of new ways to get people to engage with historical primary documentation, but the nuance that these colorizations are interpretations gets lost immediately. Do an image search for "D-Day in color" and try to tell me which results are original color negatives and which are colorizations made by teenagers. I'm also a little confused as to why colorizations always aim to restore color to the equivalent of a faded color negative, with muted tonality and grain. Human logic is funny.
- nerdponx 9y agoI'm also a little confused as to why colorizations always aim to restore color to the equivalent of a faded color negative, with muted tonality and grain. Human logic is funny. I always assumed that if you tried to use "full color" it would look weird, since the photos themselves usually are quite faded and grainy.
- jacobolus 9y agoMore than that, it would look weird because restoring the color to a photo in a way that looks plausibly photorealistic is really hard. If you make something more stylized, the viewer doesn’t have as much reference to compare and realize that there’s something wrong. To the grandparent: people have been trying to colorize black and white photos since the 1840s; complaining isn’t going to stop them now, https://en.wikipedia.org/wiki/Hand-colouring_of_photographs https://en.wikipedia.org/wiki/Hand-colouring_of_photographs
- vwcx 9y agoTo the grandparent: people have been trying to colorize black and white photos since the 1840s True. But there was also tons of skepticism in the medium throughout the mid- to late 1800s. Oliver Wendell Holmes' writing on veracity of photography is a neat reminder that it took the public decades to come to terms that the photographic process was a somewhat-veritable facsimile of "real" life.
- 9y ago
- quotemstr 9y agoThe suggestion of using a classification network as a loss function is brilliant! I love how we can, in general, elevate the sophistication of ML models by having different models interact and train each other.
- strin 9y agoLive demo! http://beta.moxel.ai/models/strin/colorization/latest http://beta.moxel.ai/models/strin/colorization/latest