4 ms·
The answer is no, since when you generate data, you either * Have a set of data as a "basis" 1. This diminishes the "equalization" factor since you need
by kory 7y ago
The answer is no, since when you generate data, you either
* Have a set of data as a "basis"
1. This diminishes the "equalization" factor since you
need a lot of data to get a good approximation of the
distribution anyways.
2. You need to create a model based off of the set,
which mathematically should be close to the same problem
as just building the target model.
* No or small training set to use
1. You need to create a model that probably uses some
stat distribution to generate. Your target model will just
learn that distribution.
2. Your initial assumptions create a distribution, and
that is not going to be the same distro of real-world
data. Maybe painfully off-base. I've worked on this
problem for months and it's fairly difficult to get
perfect in an easy (modeled by simple stat distributions)
scenario.
There are problems where generating data can work, but they're specific problems or can only be used for rare edge-cases that don't show up enough in a dataset. For the most difficult problems it is probably just as difficult to generate "correct" data as it is to generate a model without real-world data.
- gilbaz 7y agoYeah but you're assuming that you know how their creating this data. Just throwing an option out there - what if they created a latent space of 3D people and iteratively expanded it with GANs and 2D real image datasets. That would generalize. Just a thought, not sure what's really going on there, I just know that they probably have something interesting they're cooking up! This is a really crazy vision
- kory 7y agoTraining a GAN to generate people without a significantly large dataset (if that's even possible) is probably just as difficult of a problem as just building the model you want in the end without sufficient data. Assuming those image sets are small they will create a model with a large bias. If you're talking about fine-tuning an existing model with small datasets, this is done already and works fairly well if not overused. It all comes down to: to create the first "data-generating" model you need a lot of data and compute. Expanding it is a different story, but that isn't where the problem lies. We come full-circle back to the same problem as what we started with: big players can afford to build these models and small players can't.