3 ms·
Until somebody comes out with a better way to trace back output images to training data, this is the best "data" we have so far. Before this, there were anecdo
by semicolon_storm 4y ago
Until somebody comes out with a better way to trace back output images to training data, this is the best "data" we have so far.
Before this, there were anecdotal examples of SD outputting an image with some Shutterstock watermark or very similar to some artist's work, but the prompts also seemed highly specific or were asking for something in that artist's style.
This tool at least lets us start to trace back the average image, and so far it does seem SD is adding something novel to its outputs.
- gwern 4y ago> Until somebody comes out with a better way to trace back output images to training data, this is the best "data" we have so far. There are many better ways, in the sense that they actually do something like estimate the causal effect of a specific training datapoint, like leave-one-out cross-validation training or surrogates to Shapley value, or using nonparametric models to trace backwards. This is a whole subfield of ML research.* (The primary summary is: "it's hard and the easy approaches don't work." Which why he's not doing any of those but an easy incorrect thing.) * I'm not entirely sure why anyone cared so much... Research topics can be kinda arbitrary. But in this case, I think there was something of a fad around 2017 that there were going to be 'data marketplaces' where you would be trying to estimate the value of each datapoint to price it. This turned out to not exist as a business model: you either used big public data for free for generic model capabilities, or you had small proprietary data you'd die rather than sell to a competitor.