6 ms·
Information theoretically speaking, how do you generate a "synthetic" dataset (as the article calls it) with the same fidelity as an original dataset without ha
by throwawaymath 7y ago
Information theoretically speaking, how do you generate a "synthetic" dataset (as the article calls it) with the same fidelity as an original dataset without having access to a critical basis set of the original? What would you do to obtain that fidelity? Extrapolate from sufficiently many independent conclusions drawn from the original?
And as a followup, if you can generate a synthetic dataset by extrapolating from sufficiently many independent conclusions drawn from the original (as opposed to having access to the original itself), would you still need to use such a dataset for training?
Things like Monte Carlo simulation can be used to approximate real world conditions, but they can't typically capture the full information density of organic data. For example, generating a ton of artificial web traffic for fraud analysis or incident response only captures a few dimensions of what real world user traffic captures.
The author talks about simulating data to focus on edge cases or avoid statistical bias, but I don't see how simulated data actually achieves that.
- acollins1331 7y agoInterpolation between known sets? You could have an envelope of real world conditions and create thousands of samples that are variations in between for purposes of training.
- ska 7y ago> Information theoretically speaking, how do you generate a "synthetic" dataset (as the article calls it) with the same fidelity as an original dataset without having access to a critical basis set of the original? You don't. It's not useless as a technique, but it is limited. IMO more limited than the article's author presents, but they do touch on some of the useful bits.
- gilbaz 7y agoCool points - "...original dataset without having access to a critical basis set of the original?" I think that they're not trying to copy existing datasets but are trying to generate new datasets that solve various computer vision use-cases. Looks lke they're using 3D photorealistic models and environments to then generate 2D data. It is a cool idea, if they had the ability to synthesize a large amount of 3D people and objects and insert them into 3D environment in ways that made sense and then run motion simulation, they could hypothetically create an incredible amount of high-quality data. Sounds pretty hard to do honestly... I think Monte Carlo is used for something very different than computer vision / machine learning. Monte Carlo is usually used to estimate an average result given many dependent variables and a simplified model of the problem. So if I want to estimate how far my paper airplane will fly and I have a simulator, I would vary the paper thickness, folds and wind. Each time I would run the simulator, get a result and then I can estimate the average distance the paper airplane would go! (actually sounds like a fun project lol). Anyway this is just different. Simulation is good for edge cases because you can simulate them disproportionally to their prevalence in the real world. So let's say that we're in a smart store and we want to recognize when an elderly person falls on the floor to send human help to the correct location. This happens maybe one in 5 year in a given store. If we were to gather data we may get 10 examples. If they can simulate this, they could simulate 100k elderly people falling and then train models to recognize it! Kind of crazy really.
- deehouie 7y agoBut then how do you simulate, or imagine all the possible ways of falling and all the possible places this could happen? You have one sample, that's all. Ultimately, you have to use domain knowledge, but domain knowledge comes from observed data. High fidelity comes from having a lot of data. This takes you back to ground one.
- TrackerFF 7y agoHasn't that been a thing with at least car/vehicle detection etc. for a while now? Generating tons of data from simply using decent 3D renderings, made with game engines etc.
- cbrun 7y agoAll good points. In this case, the original dataset is created from real world body scans. You collect enough scans in this "base collection of scans" to have a "real" distribution of the world. You can then span a latent space on top of this initial distribution and use GANs to further scale it. This isn't as good as real yet, but it generates results that are better than limited quantities of real data alone. Agree with your point around the Monte Carlo simulation. Synthetic data is not the be-all end-all to train neural networks.
- yshcht 7y agoI think the work on domain randomization for visual data is something that can be worth exploring. https://lilianweng.github.io/lil-log/2019/05/05/domain-randomization.html https://lilianweng.github.io/lil-log/2019/05/05/domain-rando...
- paggle 7y agoIt's not about creating information in the information-theoretic sense. It's that nobody knows how neural networks work. So even if, say, all of the knowledge to recognize that this is a screwdriver is present in the 3D model, it's not like we know how to train a neural network to detect those features from a 2D camera image. So we just generate a gazillion 2D images in various rotations and lighting conditions and let the neural network use its black magic. With no "backchannel" into the "brains" of the neural network, we can't tell it to recognize black people's emotions the same way it recognizes white people's emotions. So if we don't have enough black people in our training set, we build a way to simulate an image of a black person and hammer it into the neural network's brainstem.
- jeromebaek 7y ago> Information theoretically speaking, how do you generate a "synthetic" dataset (as the article calls it) with the same fidelity as an original dataset without having access to a critical basis set of the original? Information theoretically speaking, this is impossible. The synthetic dataset will always have exploitable mathematical properties in a way non-synthetic dataset will not. It will open up the trained model to easy adversarial attacks.
- proverbialbunny 7y agoHow you do it is similar to this talk: https://youtu.be/MiiWzJE0fEA https://youtu.be/MiiWzJE0fEA It's not perfect. You do need real world data. Synthetic data is great for creating edge cases. Say you want your model to classify more general use cases, but your real world data is limited. You can generate data that fits edge cases with some variance and the model will accept more variety, reducing over fitting.