3 ms·
The paper itself actually goes to show that this is insufficient (but should still be done, they went from regenerating 1 in 819 to 1 in 1063 images in a small
by babel_ 4y ago
The paper itself actually goes to show that this is insufficient (but should still be done, they went from regenerating 1 in 819 to 1 in 1063 images in a small corpus litmus test), and that diffusion is weaker to image extraction attacks than GANs (and seems to be weaker the better it gets).
As to why? Lack of care and vigilance, most likely. If they're willing to spend millions of GPU hours training some of these big models, then this wouldn't be a huge cost, and has near linear complexity over the dataset that parallelises trivially. Sadly, vigilance is a finite resource, and that kind of data cleaning is often dismissed in favour of doubling down on the training and assumming it'll be big enough to handle it (however, as has been shown, that's not the case).
(Edit: thanks FeepingCreature. The deduplication test used for the numbers above was a separate test in the paper for just this, not indicative of their extraction attack in general, with a small corpus and compared a diffusion model trained on both. So I liken it to a litmus test for the efficacy of deduplication.)
- FeepingCreature 4y agoNote: they went from regenerating 1 in 819 to 1 in 1063 on a relatively small corpus.
- f38zf5vdt 4y agoSD2's dataset was deduplicated prior to training. That's why the paper is about SD1, which was a prototype model completed before Stability even had VC raises.