3 ms·
The idea that the synthetic generator model must have solved the actual modelling problem is a attractive idea that doesn’t correspond to what people want data
by thruflo 7y ago
The idea that the synthetic generator model must have solved the actual modelling problem is a attractive idea that doesn’t correspond to what people want data for: they want to eyeball it, see what’s in it, test some algorithms, figure out how they might approach the problem. You can do that very well on realistic synthetic data (much better than any other privacy tech) even if the synthetic data has lost some utility through its statistical approximations.
The idea that training on synthetic data is a “charade” misunderstands the usefulness of having realistic, “drop in compatible” data that works with your existing code or models.
The ideas of training models as a service and also of working directly with synthetic data generators to “extract“ from them are great but incompatible with (a) the complexity of real world DS workflows in regulated industries and (b) data scientists current code / workflows / techniques.
If I’m a bank, I’m not going to give you my fraud rules and you’re not going to solve my problems with xgboost. Access to models is just as locked down as access to data.
This is why it’s useful having an intermediary. Like ... a generative model that you can train where the data is and then copy over to where the modelling is happening.
- ska 7y agoI think there are a number of different use cases that people want data for, so there aren't any blanket answers to "is this a good approach". As an intermediary like you describe is far different than "I don't have enough real data for what I want to do", for example.