3 ms·
You can decompose an entire scene into a collection of 3D GANs, each of which is fit by GAN Inversion to the object instance which best explains what is seen in
by dougabug 4y ago
You can decompose an entire scene into a collection of 3D GANs, each of which is fit by GAN Inversion to the object instance which best explains what is seen in the images.
Also, we absolutely can tell the difference between trash cans with and without lids, as well as brick walls vs painted walls. This falls under many areas of ML-based Computer Vision, such as metric, contrastive, and zero shot machine learning, also open world object detection / recognition. These frame level semantic discrepancies can be detected by discriminative DNNs, which we can use to provide a temporal consistency loss to guide the generator to produce temporally consistent videos.