10 ms·
They love saying things like "generative AI doesn't know physics". But the constraint that both eyes should have consistent reflection patterns is just another
by jmmcd 2y ago
They love saying things like "generative AI doesn't know physics". But the constraint that both eyes should have consistent reflection patterns is just another statistical regularity that appears in real photographs. Better training, larger models, and larger datasets, will lead to models that capture this statistical regularity. So this "one weird trick" will disappear without any special measures.
- actionfromafar 2y agoWouldn't also the adverserial model training have to the take "physics correctness" into account? As long as the image detects as "<insert celebrity> in blue dress", why would it care about correct details in eyes if nothing in the "checker" cares about that?
- Filligree 2y agoCurrent image generators don’t use an adversarial model. Though the ones that do would have eventually encoded that as well; the details to look for aren’t hard-coded.
- actionfromafar 2y agoInteresting. Apparently, I have much to learn.
- themoonisachees 2y agoGP told you how they don't work, but not how they do: Current image generators work by training models to remove artificial noise added to the training set. Take an image, add some amount of noise, and feed it with it's description as inputs to your model. The closest the output is to the original image, the highest the reward function. Using some tricks (a big one is training simultaneously on large and small amounts of noise), you ultimately get a model that can remove 99% noise based only on the description you feed it, and that means you can just swap out the description for what you want the model to generate and feed it pure noise, and it'll do a good job.
- 101008 2y agoI read this description of the algorithm a few times and I find it fascinating because it's so simple to follow. I have a lot of questions, though, like "why does it work?", "why nobody thought of this before", and "where is the extra magical step that moves this from 'silly idea' to 'wonder work'"?
- ikari_pl 2y agoanswer to the 2nd and 3rd question is mostly "vastly more computing power available", especially the kind that CUDA introduced a few years back
- aitchnyu 2y agoDid anybody prompt a GenAI to get this output?
- spywaregorilla 2y agoIt wouldn't work. The models could put stuff in the eyes but it wouldn't be able to do so realistically, consistently or even a fraction of the time. The text describing the images does not typically annotate tiny details like correct reflections in the eyes so prompting for it is useless.
- krembo 2y agoJust like the 20 fingers disappeared
- sholladay 2y agoAgreed, but the tricks are still useful. When there are no more tricks remaining, I think we must be pretty close to AGI.
- layer8 2y agoThat still won’t make them understand physics. This all reminds me of “fixing” mis-architected software by adding extra conditional code for every special case that is discovered to work incorrectly, instead of fixing the architecture (because no one understands it).
- tossandthrow 2y agoThis is more a comment to the word "understand" than "physics". Yes, the models output will converge to being congruent with laws of physics by virtue of deriving that as a latent variable.
- deleted 2y ago[deleted]
- falcor84 2y ago>That still won’t make them understand physics I would assume that larger models working with additional training data will eventually allow them to understand physics to the same extent as humans inspecting the world - i.e. to capture what we call Naive Physics [0]. But the limit isn't there; the next generation of GenAI could model the whole scene and then render it with ray tracing (no special casing needed). [0] https://en.wikipedia.org/wiki/Na%C3%AFve_physics https://en.wikipedia.org/wiki/Na%C3%AFve_physics
- layer8 2y agoThere seems to be little basis for this assumption, as current models don’t exhibit understanding. Understanding would allow to apply it to situations that don’t match existing patterns in the training data.
- aj7 2y agoThat’s not large models “understanding physics.” Better, giving output “statistically consistent” with real physical measurements. And no one, to my knowledge, has yet succeeded in a general AI app that reverts to a deterministic calculation in response to a prompt.
- Someone 2y agoBut we don’t know how much larger the models will have to be, how large the data sets or how much trianing is needed, do we? They could have to be inconceivably large. If you want to correct for this particular problem you might be better off training a face detector, an eye detector and a model that takes two eyes as input and corrects for this problem. Process then would be: - generate image - detect faces - detect eyes in each face - correct reflections in eyes That is convoluted, though, and would get very convoluted when you want to correct for multiple such issues. It also might be problematic in handling faces with glass eyes, but you could try to ‘detect’ those with a model that is trained on the prompt.
- rocqua 2y agoI feel like a GAN method might work better, building a detector, and training the model to defeat the detector.
- bastawhiz 2y ago> They could have to be inconceivably large. The opposite might also be true. Just having better, well curated data goes a long way. LAION worked for a long time because it's huge, but what if all the garbage images were filtered out and the annotations were better? The early generations of image and video models used middling data because it was the only data. Since then, literally everyone with data has been working their butts off to get it cleaned up to make the next generation better. Better data, more intricate models, and improvements to the underlying infrastructure could mean these sorts of "improvements" come mostly "for free".
- wruza 2y agoADetailer does exactly that. Feels like this large thread above is non-practicing for the most part. There’s no eyes module in it by default, but it’s trivial-ish to add, and a hires eyes dataset isn’t hard to collect either. Just found eyes model on https://civitai.com/models/150925/eyes-detection-adetailer https://civitai.com/models/150925/eyes-detection-adetailer (seems anime only)
- amelius 2y agoShouldn't a GAN be able to use this fact immediately in its adversarial network?
- godelski 2y agoUnfortunately no. The GAN always need to be in balance and contention with the generator. You can swap out the discriminator later, but you also got to make sure your discriminator is able to identify these errors. And ML models aren't the best at noticing small details. And since they too don't understand physics, there is no reason to believe that they will encode such information, despite every image in real life requiring consistency. Also remember that there is a learning trajectory, and most certainly these small details are not learned early on in networks. The problem is that this information is post hoc trivial to identify errors, but it isn't a priori. It is also easy for you because you know physics innately and can formulate causal explanations.
- johnsutor 2y agoI know there are murmurs that synthetic data (i.e. using rendering software with 3D models) was used to train some generative models, including OpenAI Sora; seems like it's the only plausible way right now to get the insane amounts of data needed to capture such statistical regularities.
- sangnoir 2y ago> Better training, larger models, and larger datasets, will lead to models that Hypothetically, with enough information, one could predict the future (barring truly random events like radioactive decay). Generative AI is also constrained by economic forces - how much are GenAI companies willing to invest to get eyeball reflections right? Would they earn adequate revenue to cover the increase in costs to justify that feature? There are plenty of things that humanity can technically achieve, that don't get done because the incentives are not aligned- for instance, there is enough food grown to feed every human on earth and the technology to transport it, and yet we have hunger, malnutrition and famines.
- stevenwalton 2y ago> how much are GenAI companies willing to invest to get eyeball reflections right? Willing to? Probably not much. Should? A WHOLE LOT. It is the whole enchilada. While this might not seem like a big issue and truthfully most people don't notice, getting this right (consistently) requires getting a lot more right. It doesn't require the model knowing physics (because every training sample face will have realistic lighting). But what underlines this issue is the model understanding subtleties. No model to date accomplishes this. From image generators to language generators (LLMs). There is a pareto efficiency issue here too. Remember that it is magnitudes easier to get a model to be "80% correct" than to be "90% correct". But recall that the devil is in the details. We live in a complex world, and what that means is that the subtleties matter. The world is (mathematically) chaotic, so small things have big effects. You should start solving problems not worrying about these, but eventually you need to move into tackling these problems. If you don't, you'll just generate enshitification. In fact, I'd argue that the difference between an amateur and an expert is knowledge of subtleties and nuance. This is both why amateurs can trick themselves into thinking they're more expert than they are and why experts can recognize when talking to other experts (I remember a thread a while ago where many people were shocked about how most industries don't give tests or whiteboard problems when interviewing candidates and how hiring managers can identify good hires from bad ones).
- dwaltrip 2y agoGetting the eyeballs correct will correlate with other very useful improvements. They won’t train a better model just for that reason. It will just happen along the way as they seek to broadly improve performance and usefulness.
- stevenwalton 2y ago> But the constraint that both eyes should have consistent reflection patterns is just another statistical regularity that appears in real photographs Hi, author here of a model that does really good on this[0]. My model is SOTA and has undergone a third party user study that shows it generates convincing images of faces[1]. AND my undergrad is in physics. I'm not saying this to brag, I'm giving my credentials. That I have deep knowledge in both generating realistic human faces and in physics. I've seen hundreds of thousands of generated faces from many different models and architectures. I can assure you, these models don't know physics. What you're seeing is the result of attention. Go ahead and skip the front matter in my paper and go look at the appendix where I show attention maps and go through artifacts. Yes, the work is GANs, but the same principles apply to diffusion models. Just diffusion models are typically MUCH bigger and have way more training data (sure, I had access to an A100 node at the time, but even one node makes you GPU poor these days. So best to explore on GANs ): I'll point out flaws in images in my paper, but remember that these fool people and you're now primed to see errors, and if you continue reading you'll be even further informed. In Figures 8-10 you can see the "stars" that the article talks about. You'll see mine does a lot better. But the artifact exists in all images. You can also see these errors in all of the images in the header, but they are much harder to see. But I did embed the images as large as I could into the paper, so you can zoom in quite a bit. Now there are ways to detect deep fakes pretty readily, but it does take an expert eye. These aren't the days of StyleGAN-2 where monsters are common (well... at least on GANs and diffusion is getting there). Each model and architecture has a different unique signature but there are key things that you can look for if you want to get better at this. Here's things that I look for, and I've used these to identify real world fake profiles and you will see them across Twitter and elsewhere: - Eyes: Eyes are complex in humans with lots of texture. Look for "stars" (inconsistent lighting), pupil dilation, pupil shape, heterochromia (can be subtle see Figure 2, last row, column 2 for example), and the texture of the iris. And also make sure to look at the edge of eyes (Figs 8-10) and - Glasses: look for aberrations, inconsistent lighting/reflections, and pay very close attention to the edges where new textures can be created - Necks: These are just never right. The skin wrinkles, shape, angles, etc - Ears: These always lose detail (as seen in TFA and my paper), lose symmetry in shape, are often not lit correctly, if there are earrings then watch for the same things too (see TFA). - Hair: Dear fucking god, it is always the hair. But I think most people might not notice this at first. If you're having trouble, start by looking at the strands. Start with Figure 8. Patches are weird, color changes, texture, direction, and more. Then try Fig 9 and TFA. - Backgrounds: I make a joke that the best indicator to determine if you have a good quality image is how much it looks like a LinkedIn headshot. I have yet to see a generated photo that has things happening in the background that do not have errors. Both long-range and local. Look at my header image with care and look at the bottom image in row 2 (which is pretty good but has errors), row 2 column 4, and even row 1 in column 4's shadow doesn't make sense. - Phase Artifacts: This one is discussed back in StyleGAN2 paper (Fig 6). These are still common today. - Skin texture: Without fail, unrealistic textures are created on faces. These are hard to use in the wild though because you're typically seeing a compressed image and that creates artifacts too and you frequently need to zoom to see. They can be more apparent with post processing though. There's more, but all of these are a result of models not knowing physics. If you are just scrolling through Twitter you won't notice many of these issues. But if you slow down and study an image, they become apparent. If you practice looking, you'll quickly learn to find the errors with little effort. I can be more specific about model differences but this comment is already too long. I can also go into detail about how we can't determine these errors from our metrics, but that's a whole other lengthy comment. [0] https://arxiv.org/abs/2211.05770 https://arxiv.org/abs/2211.05770 [1] https://arxiv.org/abs/2306.04675 https://arxiv.org/abs/2306.04675
- ken47 2y ago> So this "one weird trick" will disappear without any special measures. > Better training, larger models, and larger datasets But "better training" here is a special measure. It would take a lot of training effort to defeat this check. For example, you'd need a program or group of people who would be able to label training data as realistic/not based on the laws of physics as reflected in subjects' eyeballs.
- deleted 2y ago[deleted]
- archgoon 2y agoOne does not get Newton by adding more epicycles.
- dheera 2y agoExactly. Notably, in my experiments, diffusion models based on U-Nets (e.g. SD1.4, SD2) are worse at capturing "correlations at a distance" like this in comparison to newer, DiT-based methods (e.g. SD3, PixArt).
- thesz 2y agoHere's the link about neural scaling law: https://en.wikipedia.org/wiki/Neural_scaling_law https://en.wikipedia.org/wiki/Neural_scaling_law Can you make a napkin calculation of how much better training should be, how larger models and datasets should be to overcome difference in widely separated but related pixels of the image? I did that and results are not promising: 200 billions parameter model will not perform much better than 100 billions' one. The loss is already too small. Also, the phenomena exemplified in the article is a problem of relation between distant entities in generated media. The same can be seen in LLM where they have inconsistencies at the beginning and end of sentences. If generated eyes start to lie, there will be different separated but related objects to rely on.
- bjackman 2y agoThis was my first thought too... You can already see in the examples in the article that the model understands that photos of people tend to have reflections of lights in their eyes. It understands that both eyes tend to reflect the same number of lights. It's already modelling that there's a similarity relationship between these areas of the image (nobody has dichromia in these pictures). I can remember when it was hard for image generators to model the 3d shape of a table, now they can very easily display 4 very convincing legs. I don't have technical expertise here but it just seems like a natural starting point to assume that this reflection thing is a transient shortcomimg.
- KingOfCoders 2y agoOr the times when the model would create six fingers. No longer.
- ec109685 2y agoSometimes this is addressed not by fixing one model, but instead running post processing models that are specialized to fix particular know defects like oddities with fingers.
- KingOfCoders 2y agoAh!
- croes 2y ago>will lead to models that capture this statistical regularity. That's not guaranteed, AI does find statistical regularities we miss but also miss some we find.
- MaxBarraclough 2y agoTreating knowing (or understanding) as binary is a common failing in discussions about AI.
- ec109685 2y agoA simpler process is to automatically post process the image to “fix” the eyes. Similar techniques are used to address deformities with hands and other localized issues.