6 ms·
trained on celebA, so no, but you could for sure train this on a more varied dataset
by erwannmillon 3y ago
trained on celebA, so no, but you could for sure train this on a more varied dataset
- Eisenstein 3y agoWould it be as simple as feeding it a bunch of decolorized images along with the originals?
- atorodius 3y agoyes, so infinite training data. but the challenge will be scaling to large resolutions and getting global consistency
- drapado 3y agoI guess you can always use a two-stage process. First colorize, then upscale
- atorodius 3y agoyeah, you can use SOTA super res, but that tends to be generative too (even diffusion based on its own, or more commonly based on GANs). it can be a challenge to synthesize the right high res details. but that’s basically the stable diffusion paper (diffusion in latent space plus GAN superres)
- erwannmillon 3y agoYeah, if you have a high res image, you can get color info at super low-res and then regenerate the colors at high res with another model. (though this isn't an efficient approach at all) https://github.com/TencentARC/T2I-Adapter https://github.com/TencentARC/T2I-Adapter i've also seen a controlnet do this.
- jrockway 3y agoIs that challenging? Humans have awful color resolution perception, so even if you have a huge black-and-white image, people would think it looks right with even with very low-resolution color information. Or, if the AI hallucinates a lot of high frequency color noise, it wouldn't be noticable. Wikipedia has a great example image here: https://en.wikipedia.org/wiki/Chroma_subsampling https://en.wikipedia.org/wiki/Chroma_subsampling. Most people would say all of them looked fine at 1:1 resolution.
- atorodius 3y agoI meant more from a comoute standpoint, the models are expensive to run full res
- jrockway 3y agoI see what you mean. I think that you can happily scale the B&W image down, run the model, and then scale the chroma information back up. Something I was thinking about after writing the comment is that the model is probably trained on chroma-subsampled images. Digital cameras do it with the bayer filter, and video cameras add 4:2:0 subsampling or similar subsampling as they compress the image. So the AI is probably biased towards "look like this photo was taken with a digital camera" versus "actually reconstruct the colors of the image". What effect this actually has, I don't know!
- atorodius 3y agogood point, I hadn’t realized that you only need to predict chroma! That actully greatly simplifies things re. chroma subsampling in training data: this is actually a big problem and a good generative model will absolutely learn to predict chroma subsampled values (or JPEG artifacts even!). you can get around it by applying random downscaling with antialiasing during training.
- erwannmillon 3y agobasically the training works as follows: Take a color image in RGB. Convert it to LAB. This is an alternative color space where the first channel is a greyscale image, and two channels that represent the color information. In a traditional pixel-space (non latent) diffusion model, you noise all the RGB channels and train a Unet to predict the noise at a given timestep. When colorizing an image, the Unet always "knows" the black and white image (i.e the L channel). This implementation only adds noise to the color channels, while keeping the L channel constant. So to train the model, you need a dataset of colored images. They would be converted to LAB, and the color channels would be noised. You can't train on decolorized images, because the neural network needs to learn how to predict color with a black and white image as context. Without color info, the model can't learn.
- coldtea 3y ago>You can't train on decolorized images, because the neural network needs to learn how to predict color with a black and white image as context. Without color info, the model can't learn. I think the parent means with delocorized images used to test the success and guide the training (since they can be readily compared with the colored image they resulted from which would be the perfect result). Not to use decolorized images alone to train for coloring (which doesn't even make sense).
- atorodius 3y agoYou can take arbitrary images and convert them to grayscale for training, and do conditional diffusion
- bemusedthrow75 3y agoBut convert them to grayscale how? Black and white film doesn't have one single colour sensitivity. Play around with something like DxO FilmPack sometime (it has excellent measurement-based representations of black and white film stocks). It's a much more complex problem than it might seem on the surface.
- 3y ago