11 ms·
Some argue that synthetic data can make AI systems better
- WalterBright 5y agoIf a computer program is generating the training data, aren't you just training the AI to do the same thing as the already existing computer program does?
- throwawaynay 5y agoGenerating a realistic city and a self driving AI are two wildly different tasks
- SomewhatLikely 5y agoEven a simpler task like image classification such as: does the picture contain a lion. Imagine you have 3d model of a lion. You can render it from lots of different angles, lighting conditions, backgrounds, stretched out, curled up, etc. You know the ground truth classification on all renderings is that the picture contains a lion, but being able to generate images of lions is a very different task from recognizing lions in images.
- hahajk 5y agoThe potential issue with using synthetic data to simulate the problem (like image classification) is that recognizing a lion in synthetic imagery and recognizing a lion in real imagery may also be very different tasks to a computer.
- throwawaynay 5y agonot that much actually because the image are handled at a lower resolution and with a smaller set of colors, to make the models train faster(a lot faster) so they look pretty much the same to a computer vision algorithm
- ausbah 5y agoat least in the case of reinforcement learning, no. just because you can simulate the problem doesn't mean you know how to optimally solve it - ex driving a car
- btdmaster 5y agoNot necessarily: https://en.wikipedia.org/wiki/Data_augmentation https://en.wikipedia.org/wiki/Data_augmentation
- ynfnehf 5y agoYou can train an AI to do the inverse of the existing program (as is the case for the self-driving described in the article.) Take some input, generate output using the existing program, and then train the AI with the input/output reversed.
- tomp 5y agoNo. We had realistic 3D graphics 20 years ago. I'm not aware of any true 3D computer vision system that could reliably play those games (from just vision).
- csee 5y agoNot at all. The existence of the Gran Turismo game (a simulator) is not the same thing as an AI that can play Gran Turismo.
- stingraycharles 5y agoBut a better analogy would be if an AI generated a computer game that another computer can learn to play. In the end, it’s more like an anonymization layer than anything. If a computer is trained to generate input data for other computers to train with, there’s not a lot special going on.
- manmal 5y agoGenerating an environment and acting within it are wildly different things. Eg Tesla generates virtual camera footage of traffic situations they want their vehicles to handle correctly. The footage generator is basically a scripted video game director, while the trained AI is one of the most complex software projects ever.
- csee 5y ago"But a better analogy would be if an AI generated a computer game that another computer can learn to play." Doesn't the latter AI have a policy which contains novel information that does not exist in the former AI? Even if what you say is true in some abstract information theory sense (and I would question that), there is a world of practical difference in the usefulness of a trained self-driving AI and the game engine within which that AI functions.
- taeric 5y agoBut.... if you were to train a model on the simulator of the game, you would expect it to pick up on the rules programmed into the simulator. This is really no different than it picking on the rules embedded in the data gathering of real world data. Any implicit and hidden decisions in that space would be expected to find their ways into the ML.
- redytedy 5y ago
- 8note 5y agoas a contrast to the other comments, you will still get biases from where the simulation differs from reality. eg. if your simulated traffic lights dont blink at 60hz, the model trained on it wont know to handle it
- wolverine876 5y agoI thought the same thing ... Simulations are daily parts of life, in research and development from real-world simulations in research (e.g., climate) and industry (often in spreadsheets); to the theory of gravity (gravity simulated with mathematics) - and every other theory of science, social science, and humanities; to develping your iPhone app on your laptop or just reading the train schedule or using a mapping program. So what is the difference? Those simulations were built from reality. The Theory of Gravity was built from and confirmed with empirical observations, not from someone else's simulation of gravity! That essential foundation of science, reality (it's not science otherwise), is what is missing. Also, we already have the problem of our biases and preconceived notions infecting training data, and AI becoming a simulation of that rather than reality. By then training on 'simulated data' (yikes!), we seem to create more of a loop.
- rapjr9 5y agoIf machine learning (ML) is trained on human behavior it can never be better (for some measure of better) than people are. So if racism is widespread, it will influence decisions trained into a ML algorithm. That raises the question, can we train an ML algorithm to make the decisions we _want_ rather than base its decision making on what people currently do. Training on generated data might be one way to do that, to build in implicit biases because we want more fairness in decision making than people currently exhibit. But then who gets to choose what those "goal" biases are? Someone could add a bias that improves life for people of male gender and worsens it for everyone else for example. This seems like a really important problem that has no clear answer. It's very much related to electing politicians where the choice of politician is also a choice of future goals. We don't know of any objective way to always decide what goals are best for everyone (voting certainly is not objective since many voters are not aware of all possible policy/goal implications and can be tricked into voting against their own interests). Yet it seems certain that researchers are working to introduce selective biases into algorithms, if just to adjust for known biases that are problematic. As an opaque input into an opaque algorithm, biases become invisible in deployments and could become very difficult to reverse or fix later. Even with continuous learning, if people's behavior is the input, the result will never be better than people are currently, which can create problems for the future. Some kind of intentional bias aimed towards goals seems necessary, yet also seems very dangerous since it can introduce biases decided by a very small set of people.
- RicoElectrico 5y agoThe fact that you can compute y = f(x) doesn't imply you already know x = f⁻¹(y).
- WalterBright 5y agoIf you want an AI that can recognize cats, and train it with computer generated pictures of cats, you may wind up with an AI that only recognizes cat pictures from that particular generator.
- dmabuf 5y agoThis tradeoff is known as synthetic domain shift and is still an active area of research. https://paperswithcode.com/task/synthetic-to-real-translation https://paperswithcode.com/task/synthetic-to-real-translatio...
- renewiltord 5y agoAn obvious counterexample. I was young in 2005 when Tesseract was open sourced. I wanted to use it to do something. But I decided to trial it out first by writing something in Notepad and then screenshotting it and trying. Synthetic data! But no, I didn’t know how to make the existing computer program turn image into text.
- SomewhatLikely 5y agoPretty clickbaity. Lots of "some argue", "some say", "has estimated", and "striving to", but not much substance about actual successes. I believe both Tesla and Cruise are working in this direction but there are serious issues to be worked out. I also vaguely remember some work on pose estimation being helped by generating renderings. Going over real successes would make for a more convincing article.
- sockpuppet69 5y ago
- chestervonwinch 5y agohttps://en.wikipedia.org/wiki/Weasel_word https://en.wikipedia.org/wiki/Weasel_word
- MaxBarraclough 5y agoMy thought exactly. Also, HackerNews submissions are generally meant to use the title from the article.
- fxtentacle 5y agoOdd, my impression is that everything moves in the opposite direction. Good synthetic data is expensive. Recording the real world is free. Just a variational autoencoder with a discrete latent space is good enough to learn a usable phoneme recognition and pronunciation model from raw WAV files with unsupervised learning. And clip shows is just for far you can get without supervision. So what's the point in paying for artificial data if you can solve the problem without it?
- zhobbs 5y agoRecording real world data can be cheap for some use cases, but often labeling it is very expensive.
- krapht 5y agoWhat. That really depends on your task. In my field, real data is extremely expensive and if we could generate believable synthetic data, we'd save a ton of money.
- hervature 5y ago> if we could generate believable synthetic data, we'd save a ton of money. It sounds like you are saying good synthetic data is more expensive?
- nomel 5y ago> Recording the real world is free. Waiting for edge cases to occur, in the real world, is definitely not free.
- vhold 5y agoWith synthetic data you can also generate the thing you are trying to solve with the AI in the first place, like generating a depth map, and an object classification map to go with a simulated image. With real world data a human will have to label it.
- fxtentacle 5y ago
- daenz 5y agoSome years ago I worked at a startup that was doing OCR on paper receipts. As part of my application to the company, I wrote a synthetic training data generator[0] to generate a range of CG receipts, along with pixel-perfect accuracy of labeled XY bounding boxes for each letter. Generating synthetic training data allows for a high degree of flexibility to the shape of your data. It allows you to focus on strengthening edge cases where you just don't have enough real world data. 0. https://www.arwmoffat.com/work/synthetic-training-data https://www.arwmoffat.com/work/synthetic-training-data
- nomel 5y agoWhen developing anything in the real world, you don't wait for edge cases to happen naturally. You force them to happen, and look at the response. This is the rigor of engineering. Has that been lost?
- snek_case 5y agoThere's this dogmatic idea in the machine learning field that only real world data is valuable.
- cinntaile 5y agoData augmentation is a thing within machine learning (deep learning) so that's a big generalization.
- snek_case 5y agoData augmentation is a big thing in applied machine learning. Deep learning practitioners use data augmentation because it works incredibly well. However, deep learning researchers tend to view data augmentation as some kind of dirty trick that wouldn't be necessary if you just had a bigger dataset.
- a_bonobo 5y agoI guess the field is constantly making a decision between 'do we want to outperform humans, but lose interpretability' and 'we should always be interpretable' Human-generated artificial data will always contain the human's assumptions. But this data might not contain the 'superhuman' element that leads to the ML-system outperforming humans, because we don't know what that is (yet). Receipts are a good example for artificial data, because they're human-made. But a lot of what we do with ML systems is using data that doesn't come from humans, images of wildlife, satellite images, biological data, etc
- bpodgursky 5y agoThis is relevant to healthcare startups where it is extremely difficult to get your hands on enough real Protected Health Information to do any interesting ML work (unless you are already part of an enormous company with PHI... and even then it is harder than you'd think).
- rosenjcb 5y agoI worked at a health insurance company and had access to a lot of data (for ML research). It was impossible to lend that data to contractors for nearly any reason and this frustrated us.
- sbrother 5y ago"Some argue"? We've been doing this for the last ten years and that's just as far as my career goes back.
- rosenjcb 5y agoYou're right. This isn't new or controversial. I don't know why HN is so weak on data science. Maybe software devs are weak on math in general.
- fhood 5y agoWhat a strange article. I think that a company like Nvidia could probably provide a great deal of additional value to the world of scientific computer modeling. And I think that simulations can work really well to assist in training models, I don't really understand why that would be up for debate, they already do. What I don't understand is talking about "simulating the entire world down to atomic interactions" or teleporting to Mars via data collected from..."sensors". Little sections of this, particularly wrt pushing towards more standardized systems for building computer models make sense and seem like a worthwhile goal, but most of this reads like nonsense to me.
- freddealmeida 5y agoOne of my companies has a patent in this space (Neuri) (rather defunct now). It worked exceptionally well for time series data. Used similar work at Ascent.ai for robotics data (visual mostly) and that worked very well. In fact I don't think we ever really saw an issue with this approach.
- NalNezumi 5y agoAscent as in the startup in Tokyo? Didn't the self driving car approach fail so miserably they had to 180 degree pivot to P&P robotics applications? Interesting, but Sim2real have had some lab success for the past few years. Making it work in real application seems to be way more trickier, especially profitably
- YeGoblynQueenne 5y ago>> The solution is to just have more data and better data. Nah, sorry, that is just trying to put out the fire by throwing fuel at it. The big, big weakness of neural networks right now is their reliance on gigantic datasets that require gigantic computational resources. Neural nets need those because they can't generalise to unseen data. So people try to get them to see as much data as possible during training. They still overfit, but if they can overfit to a diverse enough dataset, then they can be useful in practice, even if that's only to solve narrow, specific instances of a problem (like in the domino recognition system in the article). To address this weakness what is needed is to find ways to make neural nets less reliant on data, not to find ways to make more data. Make neural nets capable of generalising robustly to unseen data, from few training instances. Then you don't need to train in a simulation. Of course that would require a radical rethink of how deep neural nets are trained (even perhaps whether they remain "deep", or whether they are trained using gradient descent, the sources of their data-hunger). Trying to make more data by simulation is only kicking the can down the road and the only effect it can possibly have is to push the time at which the real limitation must be really addressed even further down the line so that an other generation of researchers has to deal with it while the current generation can keep getting their papers published and their grants granted. See Vladimir Vapnik's challenge to the machine vision community: https://youtu.be/bQa7hpUpMzM?t=492 https://youtu.be/bQa7hpUpMzM?t=492 To summarise: learn to identify MNIST digits from 60 examples of each class, rather than 6000, while retaining current accuracy. (My words now:) Improve sample efficiency to improve neural nets. Neural nets have shown a remarkable ability to work well when large amount of resources are available. Now, do like everyone else does in computer science and try to make them (sample) efficient.
- teruakohatu 5y ago> learn to identify MNIST digits from 60 examples of each class, rather than 6000 More like 60,000 per class once augmentation has generated a bunch of new samples from each orginal!
- treesprite82 5y ago> To summarise: learn to identify MNIST digits from 60 examples of each class, rather than 6000, SOTA accuracy on a similar but more challenging problem (5-shot 20-way rather than 60-shot 10-way) appears to be around 99.6%: https://paperswithcode.com/sota/few-shot-image-classification-on-omniglot-5-1 https://paperswithcode.com/sota/few-shot-image-classificatio... > while retaining current accuracy. Depends how strict you're being with this. There's room for the gap to shrink, but I think on average classifiers (whether organic or machine) with a large number of examples to go off of will always perform at least marginally better than classifiers with fewer examples.
- naveen99 5y agoDreams are basically synthetic data that train our brains while we sleep.
- hoppla 5y agoHow cool would it not be if the computers would not only use synthetic data for training, but also simulating the outcome of their own actions before taking it. But not only it’s own actions - like a game of chess it would also include all the possible immediate actions of other actors
- ouid 5y agowell that just sounds like programming with extra steps.
- ddingus 5y agoDoes this not depend on the source data? Some datasets are easy to render, generate, whatever. In that scenario, sure! Seems like a solid case can be made, particularly where an analytic approach can speak to the data elements needed. Other data needs to be sourced from the world. That's harder, and it's extremely likely artificial data is either too expensive to create at the fidelity, or lack of, to make economic sense, or the cases are too numerous for an analytic approach to be inclusive.