7 ms·
“Photo of a red sphere on top of a blue cube. Behind them is a green triangle, on the right is a dog, on the left is a cat” https://pbs.twimg.com/media/GG8mm5v
by subzel0 3y ago
“Photo of a red sphere on top of a blue cube. Behind them is a green triangle, on the right is a dog, on the left is a cat”
https://pbs.twimg.com/media/GG8mm5va4AA_5PJ?format=jpg&name=large https://pbs.twimg.com/media/GG8mm5va4AA_5PJ?format=jpg&name=...
- Workaccount2 3y agoNot bad, I'm curious of the output if you ask for a mirrored sphere instead.
- svenmakes 3y agoThis is actually the approach of one paper to estimate lighting conditions. Their strategy is to paint a mirrored sphere onto an existing image: https://diffusionlight.github.io/ https://diffusionlight.github.io/
- jetrink 3y agoOne thing that jumps out to me is that the white fur on the animals has a strong green tint due to the reflected light from the green surfaces. I wonder if the model learned this effect from behind the scenes photos of green screen film sets.
- diggan 3y agoIt's just diffuse irradiance, visible in most real (and CGI) pictures although not as obvious as that example. Seems like a typical demo scene for a 3D renderer, so I bet that's why it's so prominent.
- zero_iq 3y agoThe models do a pretty good job at rendering plausible global illumination, radiosity, reflections, caustics, etc. in a whole bunch of scenarios. It's not necessarily physically accurate (usually not in fact), but usually good enough to trick the human brain unless you start paying very close attention to details, angles, etc. This fascinated me when SD was first released, so I tested a whole bunch of scenarios. While it's quite easy to find situations that don't provide accurate results and produce all manner of glitches (some of which you can use to detect some SD-produced images), the results are nearly always convincing at a quick glance.
- astrange 3y agoOne thing they don't so far do is have consistent perspective and vanishing points. https://arxiv.org/abs/2311.17138 https://arxiv.org/abs/2311.17138
- orbital-decay 3y agoAs well as light and shadows, yes. It can be fixed explicitly during training like the paper you linked suggests by offering a classifier, but it will probably also keep getting better in new models on its own, just as a result of better training sets, lower compression ratios, and better understanding of the real world by models.
- awongh 3y agoI think you have to conceptualize how diffusion models work, which is that once the green triangle has been put into the image in the early steps, the later generations will be influenced by the presence of it, and fill in fine details like reflection as it goes along. The reason it knows this is that this is how any light in a real photograph works, not just CGI. Or if your prompt was “A green triangle looking at itself in the mirror” then early generation steps would have two green triangle like shapes. It doesn’t need to know about the concept of light reflection. It does know about composition of an image based on the word mirror though.
- mlsu 3y agoIt does make sense though. Accurate global illumination is very strongly represented in nearly all training data (except illustrations) so it makes sense that the model learned an approximation of it.
- samstave 3y agoWow - is it doing pre-render-ray-tracing?
- samstave 3y agoEDIT: Wrong window folks.... What if you can | a scene to a model and just have it calc all the ray-paths and then | any color/image... if you pre-calc various ray angles, you can then just map your POV and allow for the volume as it pertains to your POV be mapped with whatever overlay you want. Here is the crazy cyberpunk part: IT (whatever 'IT' is) keeps a lidar of everything EVERYONE senses in that space and can overlap/time/sequence anything about each experience and layer (baromoter/news/blah tied to that temporal marker) Micro resolution of advanced lidar is used in signature creation to ensure/verify/detect fake places vs IRL. Secret nodes are used to anti-lidar the sensors... so a place can be hidden from drones attempting to map it. These anonolies are detectable thou, and GIS experts with terra forming skills are the new secOPs. Fn dorks. -- so, you already have an asset, lets say its a CUBOID room - with walls and such of wood texture_05.png
- smoldesu 3y agoI think you've read too far into this. Ray tracing is not a useful real-world primitive for extracting information from most scenes. Sure, "everything is shiny", but most surfaces are diffuse and don't contain useful visual information besides the object they illuminate. Many supposedly "pure" reflections like mirrors and glass are actually subtle caustics that introduce too much nuance to account for. Also, "pipe" isn't considered harmful terminology (yet) just FYI. I was confused seeing the "|" mononym in it's place.
- samstave 3y agoThanks for that - I like | . I was being lazy.... But I realize you are correctin the mirroring - I immediately thought it was ray tracing the green hue from the reflection onto a surface that could see it... Inference is far more efficient - however - it would be really interesting to know HOW an AI 'thinks' about such reflections? Whats the current status of AIs documenting themselves?
- Hugsun 3y agoThat's very impressive!
- yreg 3y agoIt is! This isn't something orevious models could do.
- iamgopal 3y agoInteresting is that Left and right taken from viewer’s perspective instead of red sphere’s perspective
- ebertucc 3y agoHow do you know which way the red sphere is facing? A fun experiment would be to write two prompts for "a person in the middle, a dog to their left, and a cat to their right", and have the person either facing towards or away from the viewer.
- leumon 3y ago"When in doubt, scale it up." - openai.com/careers
- Filligree 3y agoThat's _amazing_. I imagine this doesn't look impressive to anyone unfamiliar with the scene, but this was absolutely impossible with any of the older models. Though, I still want to know if it reliabily does this--so many other things are left to chance, if I need to also hit a one-in-ten chance of the composition being right, it still might not be very useful.
- Feuilles_Mortes 3y agoWhat was difficult about it?
- lucidrains 3y agoprevious systems could not compose objects within the scene correctly, not to this degree. what changed to allow for this? could this be a heavily cherrypicked example? guess we will have to wait for the paper and model to find out
- bbor 3y agoFrom the original paper with this technique: We introduce Diffusion Transformers (DiTs), a simple transformer-based backbone for diffusion models that outperforms prior U-Net models and inherits the excellent scaling properties of the transformer model class. Given the promising scaling results in this paper, future work should continue to scale DiTs to larger models and token counts. DiT could also be explored as a drop-in backbone for text-to-image models like DALL E 2 and Stable Diffusion. Afaict the answer is that combining transformers with diffusers in this way means that the models can (feasibly) operate in a much larger, more linguistically-complex space. So it’s better at spatial relationships simply because it has more computational “time” or “energy” or “attention” to focus on them. Any actual experts want to tell me if I’m close?
- deleted 3y ago[deleted]
- lucidrains 3y ago
- npunt 3y agoWe're getting to strong holodeck vibes here
- heyoni 3y agoNow try “a highway being held up by an airplane” Tried all morning and ChatGPT could not do it.
- 8n4vidtmkvmk 3y agoThat's hard for me to parse as a human. Do you mean the plane is on the highway and causing a traffic jam? Or is the highway literally being held by a humanoid plane?
- bamboozled 3y agoWould it be a highway resting on top of a plane, like the plane is a pillar ?