4 ms·
TL;DR I think its complementary. Synthetic data is particularly valuable when even the unlabelled data is expensive to obtain. For example if you want to train
by razcle 6y ago
TL;DR I think its complementary.
Synthetic data is particularly valuable when even the unlabelled data is expensive to obtain. For example if you want to train a driverless car, you may never see an ambulance driving at night in the rain even if you drive for thousands of miles. In that case, being able to synthesise data makes a lot of sense and we have lots of tools for computer graphics that make this easy.
Synthetic data can also be useful to share data when there are privacy concerns but my own feeling here is that there are better approaches to privacy preservation, like federated learning and learning via randomised response (https://arxiv.org/pdf/2001.04942.pdf https://arxiv.org/pdf/2001.04942.pdf).
In general though, outside of some vision applications, I'm pretty sceptical of synthetic data. For synthetic data to work well, you need a really good class conditional generator. E.g "generate a tweet that is a negative sentiment review" but if you have a sufficiently good model to do this, then you can probably use that model to solve your classification task anyway.
For most settings, I think synthetic data will work for data augmentation as a regulariser but will not be a substitute for all labelled data.
For the labelled data, active learning should still help.
- pgao 6y agoFormer self-driving engineer here. I'm also pretty skeptical about synthetic data. For the scenario you described, it turns out that if you drive enough, you'll eventually see some examples of ambulances at night in the rain. If it's really that rare, it's often easier to rent your own ambulance, drive around and do some staged data collection, and annotate the results than it is to set up a synthetic data pipeline. At the end of the day, even in vision applications, real data is always better than synthetic data if you can get it. Things like sensor noise or interference are hard to replicate in synthetic data. Most teams turn to synthetic data for simulation purposes or as a last resort.