2 ms·
While generally I agree with your conclusion (synthetic data doesn’t have better privacy guarantees, probably will hurt training if you use it naively), I would
by vladf 7y ago
While generally I agree with your conclusion (synthetic data doesn’t have better privacy guarantees, probably will hurt training if you use it naively), I wouldn’t be so pessimistic.
At the risk of digressing from TFA, Candes’ knockoffs, for instance, are an example of a (theoretically) successful use of synthetic data for model robustness. Still need original data, of course.
Basically, the broader point is that you don’t need to solve the full problem of joint likelihood estimation to use generative models effectively, e.g., GANs are another example.