3 ms·
Interessting. Why is not the "raw" sensor format, for example Bayer BG8, used more in deeplearning? Would it not contain a little bit more correct information?
by dosshell 8y ago
Interessting. Why is not the "raw" sensor format, for example Bayer BG8, used more in deeplearning? Would it not contain a little bit more correct information? and if used, we could skip the Bayer-> RGB conversion completly.
- ovi256 8y agoBecause most models are trained from files, which store in RGB. Using raw sensor format would make sense for inference, but if your model is created for RGB, why bother.
- h4b4n3r0 8y agoPretty much. And also, models almost never operate on full-size image, and if you e.g. pixel bin a bayer image you will get a de-mosaiced image, so that's not very interesting. What is interesting is, it'd be pretty cool to train on linear HDR images for instance (i.e. direct 14-stop sensor output without contrast curves and color correction), or on visible spectrum + IR + UV, but I'm not aware of any large public datasets like this. The reason why this is interesting is because a lot of vision models suffer from the lack of robustness. Their performance on the train/validation datasets might be OK, but connect them to a real live video feed and they turn to shit. To some extent this can be addressed by data augmentation during training, but you have to understand that this is an imperfect process: you're operating on the data that's already pre-cooked, so to speak, for _human_ consumption, with color correction, gamma curves, etc already applied, and there's not a whole lot you can do to make things more robust to different lighting conditions, noise in the feed when camera bumps the gain in its ADCs, color drift due to e.g. sodium lamps, over/under exposure, etc. Unsurprisingly, this is especially noticeable in efficient, quantized models, because they don't really have much leeway built in to work around the deficiencies in the input data.
- ovi256 8y agoRegarding multi-spectral images, there was a recent Kaggle competition classifying ship-or-iceberg from radar multiple band data. So yeah, CNNs work well for that too. https://www.kaggle.com/c/statoil-iceberg-classifier-challenge https://www.kaggle.com/c/statoil-iceberg-classifier-challeng...
- lovelearning 8y agoPublic raw image datasets are rare, probably because most cameras, except DSLRs, don't support storing raw images. Also, the nature of the inference may not require high fidelity raw images since they're going to be down-sampled anyway. Unless research suggests that raw images improve inference metrics, nobody's going to invest effort into building a raw image dataset. One notable recent exception to this is the "Learning to see in the dark" paper on low light image enhancement[1] - they found that training on raw images gives much better enhancement than JPEG. The data availability problem may start getting solved in the near future because mobiles (atleast Android) support storing and processing raw images, and on-device learning and inference capabilities are getting better all the time on mobiles. [1]: https://www.youtube.com/watch?v=qWKUFK7MWvg&feature=youtu.be https://www.youtube.com/watch?v=qWKUFK7MWvg&feature=youtu.be
- dosshell 8y agoInteresting link! Yhea, I guess JPEG destroys a lot information low light images. > probably because most cameras, except DSLRs, don't support storing raw images This may be true for consumer cameras but not for machine vision cameras. Most color machine vision cameras do support the Bayer format. Basler, FLIR (Point Grey), IDS and so on all support the Bayer format. If you roll your own camera you definitely have to deal with the Bayer format from the sensor you are using. And if you do not want any compression on your data-set images the Bayer format also takes less space than the RGB format. I'm not arguing it is a good idea, I'm just interested if it works :)
- lovelearning 8y agoIt's useful information about industrial machine vision cams, since I don't have any experience with them. You are right about custom cameras - in fact, I have some camera modules for experimenting with surveillance, night vision and machine learning, and some of their datasheets mention registers that, when set, output Bayer format images. Machine learning in surveillance is still a rather under-serviced market, but it's potentially a rich source of raw image datasets. Raw format training and inference definitely work. The obstacle is only data availability.
- 8y ago