5 ms·
I was just saying it's counter-intuitive that the "source" of any training data would ever matter as much as the "correctness" of the data; but you're right, th
by quantadev 2y ago
I was just saying it's counter-intuitive that the "source" of any training data would ever matter as much as the "correctness" of the data; but you're right, that was very sloppy wording on my part, sorry.
Here's a longer, related post, from me (albeit also confusing, haha):
https://news.ycombinator.com/item?id=42352759 https://news.ycombinator.com/item?id=42352759
- kerkeslager 2y agoI think you're trying to separate two inseparable concepts. The ONLY means we have of verifying the correctness of data is by comparing it with observation, i.e. data from the real world. Real world data is correct data, and synthetic data is inherently only as correct as its correlation with real world data. There is of course some real world data that is more correct than other real world data, based on collection methods, sample sizes, etc., but again the only way we know that is, again, real world data.
- quantadev 2y agoBut I think we can write computer programs to generate infinite amounts of "correct" data. For example, imagine you want to train AI to recognize Tea Cups. We can generate using computer graphics (not even AI) infinite numbers of "correct" example images to train on, simply by rotating a CGI model thru every possible viewpoint on a 3D sphere. If that doesn't work, it means there's got to be some deep physics about why. If training works on "natural" correct data but not "synthetic" correct data, that's telling us something DEEP about physics. The same is true with LLM (language data) factual statements. We can generate infinite numbers of true factual statements to train on. If the AI refuses to learn from synthetic data despite that training data being correct, that's bolstering my view that perhaps there's more deep physics going on related to causality chains and even crazy concepts like multiverses.
- kerkeslager 2y ago> But I think we can write computer programs to generate infinite amounts of "correct" data. For example, imagine you want to train AI to recognize Tea Cups. We can generate using computer graphics (not even AI) infinite numbers of "correct" example images to train on, simply by rotating a CGI model thru every possible viewpoint on a 3D sphere. That's not training AI to recognize tea cups, that's training AI to recognize your model of a teacup. Of course this will fail, because your generated data doesn't contain everything that a picture of a teacup contains. Look at real teacup pictures [1]. None of the teacups I have at home have flowers on them, so I wouldn't have thought to put flowers on my model, but most of those pictures have flowers on them. But even if you thought to put flowers on your teacup, are you now generating a bunch of varieties of flower patterns? And that's not even starting in on images like a teacup with tea in it [2], a teacup with a dog in it [3], or a teacup with a teabag[4]. In short, your 3d model ISN'T CORRECT in any useful sense of the word. [1] https://duckduckgo.com/?t=h_&q=teacup&iax=images&ia=images https://duckduckgo.com/?t=h_&q=teacup&iax=images&ia=images [2] https://thumbs.dreamstime.com/b/tea-cup-tea-upside-wood-table-white-porcelain-67100891.jpg https://thumbs.dreamstime.com/b/tea-cup-tea-upside-wood-tabl... [3] https://img-s-msn-com.akamaized.net/tenant/amp/entityid/BB1kIbdB.img?w=768&h=511&m=6 https://img-s-msn-com.akamaized.net/tenant/amp/entityid/BB1k... [4] https://upload.wikimedia.org/wikipedia/commons/thumb/a/ad/Tea_bag.JPG/1200px-Tea_bag.JPG https://upload.wikimedia.org/wikipedia/commons/thumb/a/ad/Te...
- quantadev 2y agoIn the Teacup example I should've been more clear. I didn't mean general "real world" teacup recognition. I meant as a pedagogic example test case where your goal was only to recognize that EXACT object, BUT from any view ANGLE. That's a far simpler test case than real-world, and can even be done on trivially small parameter-count MLPs too. Yes for recognizing real world objects you need large numbers of examples of real-world images sure. I was merely getting at the fact that synthetic data can be perfectly valid training data, in research scenarios where we're just doing these kinds of experiments, to probe MLP learning capabilities. What I'm trying to get at is if you have two sets of training data, that are equal in every way, except that one is synthetic and the other is real-world and training always fails on the synthetic data one, to me that's as "impossible" (i.e. astounding) as the Slit-Experiment that proves wave/particle duality, but it seems that this is indeed the case.