4 ms·
One thing that jumps out to me is that the white fur on the animals has a strong green tint due to the reflected light from the green surfaces. I wonder if the
by jetrink 3y ago
One thing that jumps out to me is that the white fur on the animals has a strong green tint due to the reflected light from the green surfaces. I wonder if the model learned this effect from behind the scenes photos of green screen film sets.
- diggan 3y agoIt's just diffuse irradiance, visible in most real (and CGI) pictures although not as obvious as that example. Seems like a typical demo scene for a 3D renderer, so I bet that's why it's so prominent.
- zero_iq 3y agoThe models do a pretty good job at rendering plausible global illumination, radiosity, reflections, caustics, etc. in a whole bunch of scenarios. It's not necessarily physically accurate (usually not in fact), but usually good enough to trick the human brain unless you start paying very close attention to details, angles, etc. This fascinated me when SD was first released, so I tested a whole bunch of scenarios. While it's quite easy to find situations that don't provide accurate results and produce all manner of glitches (some of which you can use to detect some SD-produced images), the results are nearly always convincing at a quick glance.
- astrange 3y agoOne thing they don't so far do is have consistent perspective and vanishing points. https://arxiv.org/abs/2311.17138 https://arxiv.org/abs/2311.17138
- orbital-decay 3y agoAs well as light and shadows, yes. It can be fixed explicitly during training like the paper you linked suggests by offering a classifier, but it will probably also keep getting better in new models on its own, just as a result of better training sets, lower compression ratios, and better understanding of the real world by models.
- awongh 3y agoI think you have to conceptualize how diffusion models work, which is that once the green triangle has been put into the image in the early steps, the later generations will be influenced by the presence of it, and fill in fine details like reflection as it goes along. The reason it knows this is that this is how any light in a real photograph works, not just CGI. Or if your prompt was “A green triangle looking at itself in the mirror” then early generation steps would have two green triangle like shapes. It doesn’t need to know about the concept of light reflection. It does know about composition of an image based on the word mirror though.
- mlsu 3y agoIt does make sense though. Accurate global illumination is very strongly represented in nearly all training data (except illustrations) so it makes sense that the model learned an approximation of it.
- samstave 3y agoWow - is it doing pre-render-ray-tracing?
- samstave 3y agoEDIT: Wrong window folks.... What if you can | a scene to a model and just have it calc all the ray-paths and then | any color/image... if you pre-calc various ray angles, you can then just map your POV and allow for the volume as it pertains to your POV be mapped with whatever overlay you want. Here is the crazy cyberpunk part: IT (whatever 'IT' is) keeps a lidar of everything EVERYONE senses in that space and can overlap/time/sequence anything about each experience and layer (baromoter/news/blah tied to that temporal marker) Micro resolution of advanced lidar is used in signature creation to ensure/verify/detect fake places vs IRL. Secret nodes are used to anti-lidar the sensors... so a place can be hidden from drones attempting to map it. These anonolies are detectable thou, and GIS experts with terra forming skills are the new secOPs. Fn dorks. -- so, you already have an asset, lets say its a CUBOID room - with walls and such of wood texture_05.png
- smoldesu 3y agoI think you've read too far into this. Ray tracing is not a useful real-world primitive for extracting information from most scenes. Sure, "everything is shiny", but most surfaces are diffuse and don't contain useful visual information besides the object they illuminate. Many supposedly "pure" reflections like mirrors and glass are actually subtle caustics that introduce too much nuance to account for. Also, "pipe" isn't considered harmful terminology (yet) just FYI. I was confused seeing the "|" mononym in it's place.
- samstave 3y agoThanks for that - I like | . I was being lazy.... But I realize you are correctin the mirroring - I immediately thought it was ray tracing the green hue from the reflection onto a surface that could see it... Inference is far more efficient - however - it would be really interesting to know HOW an AI 'thinks' about such reflections? Whats the current status of AIs documenting themselves?