4 ms·
Cool points - "...original dataset without having access to a critical basis set of the original?" I think that they're not trying to copy existing datasets b
by gilbaz 7y ago
Cool points -
"...original dataset without having access to a critical basis set of the original?"
I think that they're not trying to copy existing datasets but are trying to generate new datasets that solve various computer vision use-cases. Looks lke they're using 3D photorealistic models and environments to then generate 2D data. It is a cool idea, if they had the ability to synthesize a large amount of 3D people and objects and insert them into 3D environment in ways that made sense and then run motion simulation, they could hypothetically create an incredible amount of high-quality data. Sounds pretty hard to do honestly...
I think Monte Carlo is used for something very different than computer vision / machine learning. Monte Carlo is usually used to estimate an average result given many dependent variables and a simplified model of the problem. So if I want to estimate how far my paper airplane will fly and I have a simulator, I would vary the paper thickness, folds and wind. Each time I would run the simulator, get a result and then I can estimate the average distance the paper airplane would go! (actually sounds like a fun project lol). Anyway this is just different.
Simulation is good for edge cases because you can simulate them disproportionally to their prevalence in the real world. So let's say that we're in a smart store and we want to recognize when an elderly person falls on the floor to send human help to the correct location. This happens maybe one in 5 year in a given store. If we were to gather data we may get 10 examples. If they can simulate this, they could simulate 100k elderly people falling and then train models to recognize it! Kind of crazy really.
- deehouie 7y agoBut then how do you simulate, or imagine all the possible ways of falling and all the possible places this could happen? You have one sample, that's all. Ultimately, you have to use domain knowledge, but domain knowledge comes from observed data. High fidelity comes from having a lot of data. This takes you back to ground one.
- TrackerFF 7y agoHasn't that been a thing with at least car/vehicle detection etc. for a while now? Generating tons of data from simply using decent 3D renderings, made with game engines etc.