3 ms·
RGB<->YCbCr is a linear operation, so a deep enough network shouldn't have trouble learning to do the transform if it's useful. Gamma correction might be import
by johntb86 4y ago
RGB<->YCbCr is a linear operation, so a deep enough network shouldn't have trouble learning to do the transform if it's useful. Gamma correction might be important, though.
- dTal 4y agoSure, but surely you'd rather not? I imagine you want the network encoding the details of the specific image, not wasting weights to learn about easy wins that apply to every image. You could burn a whole layer converting to YCbCr, and the net in the article only had 3.
- toxik 4y agoYou don't know if that is actually the case, the whole appeal, originally, is that you don't need this kind of handcrafting of features. Now, clearly, people preprocess their data all the time so that point is kind of moot. However, if YCbCr was meaningfully better for the network, I think it would be pretty well-known as a preprocessing step. Also, the actual sensor data is RGB.
- limbicsystem 4y agoBut the human visual system is insensitive to high resolution yb and red/green information so you can essentially down sample those layers to almost nothing before you even start. That's a fundamental trick which is also used in another way by jpeg but presumably not by this RGB algorithm.
- bob1029 4y agoIf you extend the JPEG analogy all the way and consider the 4:2:0 subsampling mode, then you could reduce your input parameter counts by exactly 50% by consuming a pre-converted image.