4 ms·
But I think we can write computer programs to generate infinite amounts of "correct" data. For example, imagine you want to train AI to recognize Tea Cups. We c
by quantadev 2y ago
But I think we can write computer programs to generate infinite amounts of "correct" data. For example, imagine you want to train AI to recognize Tea Cups. We can generate using computer graphics (not even AI) infinite numbers of "correct" example images to train on, simply by rotating a CGI model thru every possible viewpoint on a 3D sphere. If that doesn't work, it means there's got to be some deep physics about why. If training works on "natural" correct data but not "synthetic" correct data, that's telling us something DEEP about physics.
The same is true with LLM (language data) factual statements. We can generate infinite numbers of true factual statements to train on. If the AI refuses to learn from synthetic data despite that training data being correct, that's bolstering my view that perhaps there's more deep physics going on related to causality chains and even crazy concepts like multiverses.
- kerkeslager 2y ago> But I think we can write computer programs to generate infinite amounts of "correct" data. For example, imagine you want to train AI to recognize Tea Cups. We can generate using computer graphics (not even AI) infinite numbers of "correct" example images to train on, simply by rotating a CGI model thru every possible viewpoint on a 3D sphere. That's not training AI to recognize tea cups, that's training AI to recognize your model of a teacup. Of course this will fail, because your generated data doesn't contain everything that a picture of a teacup contains. Look at real teacup pictures [1]. None of the teacups I have at home have flowers on them, so I wouldn't have thought to put flowers on my model, but most of those pictures have flowers on them. But even if you thought to put flowers on your teacup, are you now generating a bunch of varieties of flower patterns? And that's not even starting in on images like a teacup with tea in it [2], a teacup with a dog in it [3], or a teacup with a teabag[4]. In short, your 3d model ISN'T CORRECT in any useful sense of the word. [1] https://duckduckgo.com/?t=h_&q=teacup&iax=images&ia=images https://duckduckgo.com/?t=h_&q=teacup&iax=images&ia=images [2] https://thumbs.dreamstime.com/b/tea-cup-tea-upside-wood-table-white-porcelain-67100891.jpg https://thumbs.dreamstime.com/b/tea-cup-tea-upside-wood-tabl... [3] https://img-s-msn-com.akamaized.net/tenant/amp/entityid/BB1kIbdB.img?w=768&h=511&m=6 https://img-s-msn-com.akamaized.net/tenant/amp/entityid/BB1k... [4] https://upload.wikimedia.org/wikipedia/commons/thumb/a/ad/Tea_bag.JPG/1200px-Tea_bag.JPG https://upload.wikimedia.org/wikipedia/commons/thumb/a/ad/Te...
- quantadev 2y agoIn the Teacup example I should've been more clear. I didn't mean general "real world" teacup recognition. I meant as a pedagogic example test case where your goal was only to recognize that EXACT object, BUT from any view ANGLE. That's a far simpler test case than real-world, and can even be done on trivially small parameter-count MLPs too. Yes for recognizing real world objects you need large numbers of examples of real-world images sure. I was merely getting at the fact that synthetic data can be perfectly valid training data, in research scenarios where we're just doing these kinds of experiments, to probe MLP learning capabilities. What I'm trying to get at is if you have two sets of training data, that are equal in every way, except that one is synthetic and the other is real-world and training always fails on the synthetic data one, to me that's as "impossible" (i.e. astounding) as the Slit-Experiment that proves wave/particle duality, but it seems that this is indeed the case.
- kerkeslager 2y agoI don't think you're understanding what it means for training a model to "fail". Sure, you can train a model to recognize a CGI teacup, but nobody cares. That's like testing if your scissors can cut air or if your car can move at 0mph. The goal of training on synthetic data is to be able to have the trained model operate on real world data, and the test is whether it can operate on real-world data. And it's unsurprising when a model trained on synthetic data fails to operate on real-world data. Yes, it would be surprising if you trained an AI model on a CGI model and it failed to operate on the same CGI model. But that's not what's being tested, because that's trivial. That's not "probing MLP learning capabilities"--we know that works, and can even tune parameters to control exactly how well it works. We know exactly how complex the CGI model is, so we know exactly how much complexity we need to capture and how much complexity is lost at each step of the training process so we can calculate exactly how well the AI model will operate on that. You don't even need AI for that. What we don't know is how complex the real world is. This presents a bunch of unknowns: 1. Is our training dataset large enough to capture most of the complexity of the real world? 2. Are our success metrics measuring the complexity of the real world? 3. Which parts of our training dataset are observed complexity (signal) and which parts are merely random (noise)? > What I'm trying to get at is if you have two sets of training data, that are equal in every way, except that one is synthetic and the other is real-world and training always fails on the synthetic data one, to me that's as "impossible" (i.e. astounding) as the Slit-Experiment that proves wave/particle duality, but it seems that this is indeed the case. No, that is not the case. We DON'T have two sets of training data that are equal in every way except that one is synthetic and the other is real world. That doesn't exist, and will never exist, because it cannot exist. This idea needs to be deleted from your thinking because it is objectively, mathematically, immutably, physically, literally, specifically, absolutely, inherently impossible. It is unsurprising that training on synthetic data fails. Again: "fails" in this case, means that the model trained on synthetic data fails to operate on real world data--nobody cares if your model operates on the exact data it was trained on. The reason it is unsurprising that training on synthetic data fails to operate on real-world data is that synthetic data is inherently a loss of information from the understanding of real-world data that was used to generate it. No matter how many CGI models of teacups you generate, your CGI models of teacups will never capture all the complexity of real-world teacups. So training an AI model on CGI models of teacups will always fail the only test that matters: operating on real world data containing teacups.