3 ms·
Their extraction: (1) assumes the attacker knows the caption for some training images, and (2) primarily works on images duplicated 100x-3000x in the training d
by mxwsn 4y ago
Their extraction:
(1) assumes the attacker knows the caption for some training images, and
(2) primarily works on images duplicated 100x-3000x in the training dataset.
Their attack does not succeed for any singleton images. Deduplicating can be challenging on internet-scale datasets, but their work as presented does not appear to be a major concern for releasing diffusion models trained on other smaller datasets.
On memorization - I suspect this is a great thing for downstream performance, and a positive indicator that diffusion models are actually better generative models than prior methods (VAEs, GANs, etc). This mirrors the finding that feedforward neural networks can memorize randomly labeled data very well. Intuitively it feels like memorization is a quantifiable behavior that is a foundational activity in information processing - it is one type of optimal usage of observed data - that superpowers downstream performance.
- PartiallyTyped 4y ago> are actually better generative models than prior methods (VAEs, GANs, etc). Diffusion models are VAEs and follow the same variational framework. You could imagine that VAEs are diffusion models with a single step in the forward and backward processes ;). They actually optimize the same VLB objective, but with diffusion models the objective is a trajectory instead of a single step, however, when training we are optimizing single step transitions. This is possible because the objective ends up being a sum of logarithms, thus there is no dependence between terms. In practice we solve a simplified objective which looks a lot like as we do with standard AutoEncoders ;) The key component that differentiates the two is in what we expect of the underlying neural network. It is far easier to parameterize small changes than large ones, with VAEs you ask the decoder to produce a large change in the latent variable, whereas with diffusion, we generally split them into 4000 smaller changes assuming you are using the DDPM approach and not the DDIM one. Because we are improving with very small steps, we avoid the blurriness of VAEs, and we don't go out of distribution when sampling random noise. VAEs are often difficult to synthesize because even with KLD in the objective, the encoder produces a low variance distribution, and so when we sample noise from a high(er) variance gaussian, we are out of distribution rather quickly.
- mxwsn 4y agoI agree it's illuminating to understand diffusion models in relation to VAEs, but I personally consider them different models, but the line in the sand is definitely subjective. I think this because (reasons I'm sure you're familiar with) - Diffusion model is closest to a hierarchical VAE, but hierarchical VAEs were significantly less popular than regular VAEs - The variational objective in diffusion models in practice is weighted - Diffusion models require unchanging latent dimension while VAEs aren't restricted to this - Historically, diffusion models grew out of score-based approaches, not from VAEs
- PartiallyTyped 4y agoYou raise good points, if anything, it'd have probably been more accurate of me to express that DDPMs and probabilistic variants fit within the same Bayesian framework as VAEs but with the posterior and likelihood functions simply being markov chains instead. This allows us to separate non probabilistic diffusion models e.g. cold diffusion. But then again, what's the difference between a deterministic model and sampling from a delta function? ;)
- Zacharias030 4y agoSuper interesting post. Tyvm! Which three sources would you recommend for someone fluent in ML to read up on to arrive at your conclusions presented here (or their own)?
- PartiallyTyped 4y agoThe Variational Auto Encoder paper [1], and the DDPM paper[2] are pretty much all you need for this, [6[ is also good but covered by [2]. Going through the derivations helped solidify things for me. I haven't read [9] but looks very promising, authors include Jonathan Ho, and D. Kingma who authored [2] and [1] respectively. From there [3,4] show improvements to DDPMs, [5] shows that diffusion models can be very general. [7,8] show diffusion models from the view of score matching. [1] AutoEncoding Variational Bayes [2] Denoising Diffusion Probabilistic Models [3] Denoising Implicit Models [4] Improved Denoising Diffusion Probabilistic Models [5] Cold Diffusion [6] Deep Unsupervised Learning using Nonequilibrium Thermodynamics [7] Generative Modeling by Estimating Gradients of the Data Distribution [8] Score-Based Generative Modeling through Stochastic Differential Equations [9] Variational Diffusion Models