6 ms·
Hey, I wrote this! Happy to answer questions.
by tasdfqwer0897 7y ago
Hey, I wrote this! Happy to answer questions.
- clickok 7y agoYou mention possible score functions in Problem 5, so I have a query about that. This has been bugging me for a bit-- is there a characterization of applicable loss functions for GANs? I'm curious about their effects on the results, but also if there's room for different losses. Or do we capture most desirable behavior using the listed score functions?
- tasdfqwer0897 7y agoHmm, I'm not sure what you mean by applicable loss functions? I'll answer what I think you're asking and you can tell me if I got it wrong: There's been a lot of effort spent on coming up with different loss functions for GANs, but https://arxiv.org/abs/1711.10337 https://arxiv.org/abs/1711.10337 shows that, according to the metrics in problem 5, they don't really improve results compared to the original GAN loss function. There's something called a Wasserstein GAN https://arxiv.org/abs/1701.07875 https://arxiv.org/abs/1701.07875 that you may have heard about, but IMO the useful thing to take away from that paper is their 'gradient penalty' technique and not the new loss function.
- woliveirajr 7y agoI'm interested in using GAN as a method to make authorship attribution more robust. Using complex NN models (let's say, LSTM) doesn't give much better results than simple neurons with deep layers. Seems that the model extraction (how to represent some author style) is the only point that matters, and it then easily forged (i.e., not hard to reproduce the desired style on purpose). I'm trying to use GAN as a mean of desconstruction, of removal of the ease patterns to see if some NN can use subtle characteristics to identify the author.
- tasdfqwer0897 7y agoSo you are worried that your existing attribution method is too focused on 'obvious' attributes and you want to see if you can make it focus on less obvious things? IIUC, that's something that's been looked at in the ML Fairness literature. See this paper for example: http://www.aies-conference.com/wp-content/papers/main/AIES_2018_paper_162.pdf http://www.aies-conference.com/wp-content/papers/main/AIES_2...
- formalsystem 7y agoSo let's say I'm in a regime where I'd like to train something like a language model RNN on a private text dataset. If I train the RNN directly on the data then I'm essentially leaking private data to the output. Do you know of any approaches to generate fake data using GANs, train a language model on that fake data and still get a good classifier that does well on test data while still quantifying how much private data is being leaked? Related: do GANs work well as a data augmentation technique and can their exact contribution to how much a model could potentially be improved be quantified. EDIT: Added some clarifications
- p1esk 7y agoYou can train rnn on encrypted data
- formalsystem 7y agoSure but the output would still leak data. E.g: You give the RNN the phrase "Chase Bank" and it outputs "Sell the stock". Encrypting the data secures the data loading part but it's not like your model didn't learn anything.
- p1esk 7y agoThe output would be encrypted of course. You’d decrypt it on your end. Whoever hosts the model can’t know what it learned, without breaking your encryption.
- yorwba 7y agoUnless you use fully homomorphic encryption (way too expensive for machine learning), the model can't learn anything without breaking your encryption. So you fixed the leak only at the cost of making the model completely useless.
- p1esk 7y agoSource? Unless the encryption actually destroys information I don’t see why it would necessarily make the job harder for the rnn.
- RSchaeffer 7y agoWhy is a VAE not a generative model?
- govg 7y agoThey are generative models.