4 ms·
I think that these images are less convincing than "thispersondoesnotexist". The network did well for the extent a tree is similar to a person: long palm leaves
by srg0 5y ago
I think that these images are less convincing than "thispersondoesnotexist". The network did well for the extent a tree is similar to a person: long palm leaves look are like green hair, and look passable.
The network did bad where it had to take into account global illumination, or distance to the object and its level of detail: 1) tree trunks are mostly flat and do not depend on the possible sun location; 2) shadows don't match the trees and the rocks; 3) nearby rocks often lack texture and are not properly illuminated (they look flat).
It's not surprising given that the method is fundamentally local (CNN-based). I think better results will be achieved when a generator creates a scene graph rather than a raster. Are there works which actually create a 3D scene (or use it as an internal representation)?
- andrewmcwatters 5y agoGiven that these scenes are produced from existing rasters, I don’t know how you would do that without either data classification or time of day/GI inference. You’d have to do spacial inference to produce the 3D scenes which is more complicated, and then at that point if you’re using the generated raster output as textures, you’d have to normalize them to their albedos and then do global illumination from there.
- srg0 5y agoI do not follow recent CV research anymore, but there were various "3D from a single image" techniques even 10-15 years ago. I suppose some progress has been made. A quick search turns up this paper https://openaccess.thecvf.com/content/CVPR2021/papers/Zhang_Holistic_3D_Scene_Understanding_From_a_Single_Image_With_Implicit_CVPR_2021_paper.pdf https://openaccess.thecvf.com/content/CVPR2021/papers/Zhang_... So it's possible to infer not only depth map, but also an accurate scene graph. Estimating illumination and texture synthesis have also been done in the past. Maybe it is not necessary to go all the way into proper 3D and ray tracing, but it can be possible to use 3D representation as one of the intermediate layers. My point is these generators should embed some kind of domain model to look more realistic.