4 ms·
In the Teacup example I should've been more clear. I didn't mean general "real world" teacup recognition. I meant as a pedagogic example test case where your go
by quantadev 2y ago
In the Teacup example I should've been more clear. I didn't mean general "real world" teacup recognition. I meant as a pedagogic example test case where your goal was only to recognize that EXACT object, BUT from any view ANGLE. That's a far simpler test case than real-world, and can even be done on trivially small parameter-count MLPs too. Yes for recognizing real world objects you need large numbers of examples of real-world images sure. I was merely getting at the fact that synthetic data can be perfectly valid training data, in research scenarios where we're just doing these kinds of experiments, to probe MLP learning capabilities.
What I'm trying to get at is if you have two sets of training data, that are equal in every way, except that one is synthetic and the other is real-world and training always fails on the synthetic data one, to me that's as "impossible" (i.e. astounding) as the Slit-Experiment that proves wave/particle duality, but it seems that this is indeed the case.
- kerkeslager 2y agoI don't think you're understanding what it means for training a model to "fail". Sure, you can train a model to recognize a CGI teacup, but nobody cares. That's like testing if your scissors can cut air or if your car can move at 0mph. The goal of training on synthetic data is to be able to have the trained model operate on real world data, and the test is whether it can operate on real-world data. And it's unsurprising when a model trained on synthetic data fails to operate on real-world data. Yes, it would be surprising if you trained an AI model on a CGI model and it failed to operate on the same CGI model. But that's not what's being tested, because that's trivial. That's not "probing MLP learning capabilities"--we know that works, and can even tune parameters to control exactly how well it works. We know exactly how complex the CGI model is, so we know exactly how much complexity we need to capture and how much complexity is lost at each step of the training process so we can calculate exactly how well the AI model will operate on that. You don't even need AI for that. What we don't know is how complex the real world is. This presents a bunch of unknowns: 1. Is our training dataset large enough to capture most of the complexity of the real world? 2. Are our success metrics measuring the complexity of the real world? 3. Which parts of our training dataset are observed complexity (signal) and which parts are merely random (noise)? > What I'm trying to get at is if you have two sets of training data, that are equal in every way, except that one is synthetic and the other is real-world and training always fails on the synthetic data one, to me that's as "impossible" (i.e. astounding) as the Slit-Experiment that proves wave/particle duality, but it seems that this is indeed the case. No, that is not the case. We DON'T have two sets of training data that are equal in every way except that one is synthetic and the other is real world. That doesn't exist, and will never exist, because it cannot exist. This idea needs to be deleted from your thinking because it is objectively, mathematically, immutably, physically, literally, specifically, absolutely, inherently impossible. It is unsurprising that training on synthetic data fails. Again: "fails" in this case, means that the model trained on synthetic data fails to operate on real world data--nobody cares if your model operates on the exact data it was trained on. The reason it is unsurprising that training on synthetic data fails to operate on real-world data is that synthetic data is inherently a loss of information from the understanding of real-world data that was used to generate it. No matter how many CGI models of teacups you generate, your CGI models of teacups will never capture all the complexity of real-world teacups. So training an AI model on CGI models of teacups will always fail the only test that matters: operating on real world data containing teacups.
- quantadev 2y ago> synthetic data is inherently a loss of information That statement is exactly what I disagree with. Here's why: Thought experiment: Imagine a human infant who had only ever seen a pure white teacup, but never any other color, and only on the blue background of it's crib sheets. They can learn to understand "Teacup as a Shape" completely independent of any texture, lighting, background, etc. MLPs also can train like this, because vision AIs are generating "understandings" of shapes. If you filtered all training data (from a real-world dataset, for example) to contain ONLY white 3D rendered teacups the AI would still be able to learn Teacup shape (just like the infant), and it would recognize all teacups of all colors, even if the training data only contained synthetic-generated white ones. The following can be (and is) true at the same time: To get best results on real world objects, the best thing you can train on is real-world imagery, because a diverse set of images helps the learning. But no single "synthetic" image is "bad" (or even less useful) just because it's synthetic and not photographic.
- kerkeslager 2y ago"Thought experiment" is just a rebranding of "some shit I made up". Calling it an "experiment" belies the fact that an experiment involves collecting observations, and the only thing you're observing here is the speculation of your own brain. No part of your thought "experiment" is evidence for your opinion. > They can learn to understand "Teacup as a Shape" completely independent of any texture, lighting, background, etc. Or, maybe they can't. I don't know, and neither do you, because neither of us has performed this experiment (not thought experiment--actual experiment). Until someone does, this is just nonsense you made up. What we do know is that human infants aren't blank slates: they've got millions of years of evolutionary "training data" encoded in their DNA, so even if what you say happens to be true (through no knowledge of your own, because as I said, you don't know that), that doesn't prove that an AI can learn in the same way. This is analogous to what we do with AIs when we encode, for example, token processing, in the code of the AI rather than trying to have the AI bootstrap itself up from raw training on raw bytestreams with no understanding. You could certainly encode more data about teacups this way to close some of the gap between the synthetic and real-world data (i.e. tell it to ignore color data in favor of shape data in the code), but, that's adding implicit data to the dataset: you're adding implicit data which says that shape is more important than color when identifying teacups. And that data will be useful for the same program run against real-world data: the same code trained against a real world teacup dataset will still outperform the same code trained against a synthetic dataset when operating on real-world data. This isn't a thought experiment: it's basic information theory. A lossy function which samples its input is at most only as accurate as the accuracy of its input. But no image AIs I know of work this way because it would be a very limiting approach. The dream of AI isn't recognizing teacups, it is (in part) recognizing all sorts of objects in visual data, and color is important in recognizing some object categories. Frankly, it's clear you lack the prerequisite background in information theory to have an opinion on this topic, so I would encourage you to admit you don't know rather than spread misinformation and embarrass yourself. If you want to know more, I'd look into Kolmogorov complexity and compression and how they relate to AI. I won't be responding further because it's not worth my time to educate people who are confident that their random speculations are facts.