5 ms·
> Clearly seeing at least one photo of me (and AFAIK it was only trained on thousands of copies of that single photo) was absolutely crucial to the construction
by throwaway1851 4y ago
> Clearly seeing at least one photo of me (and AFAIK it was only trained on thousands of copies of that single photo) was absolutely crucial to the construction of this image, and yet this website isn't finding any
I think there are two separate ideas that are being conflated here. The first idea is that there is a mapping between text input and a joint text/image embedding space. For that mapping, yes, your profile photo (paired with some caption text that says saurik) would have to be the most important training input because, as you say, how else does SD know what “saurik” looks like?
But the second idea is the mapping from the shared latent space to an image output. It is NOT necessarily true that your profile photo is the most significant training example for the generated image. That’s because the text “saurik” can map to a neighborhood in the latent space that has all the photos this tool retrieved — photos which are indeed very similar to the generated image.
As an example, let’s say that I resemble Ben Affleck. And you type in “throwaway1851” and you get back an image that does look like me, but also looks an awful lot like Ben Affleck. In this case I really wouldn’t be surprised if my one profile photo didn’t contribute much to the generated image as compared to a dozen different photos of Ben Affleck. Perhaps the original mapping just pointed to Affleck anyway, because I wasn’t significant enough to end up taking space in the model’s parameters.
- saurik 4y agoI think a key question in attribution is whether the model would have been able to generate the same result without access to the input, and then how much it would have lost having been restricted from that input. If you remove from the mechanism all of the copies of my profile picture (and there are a lot of them...), I guess I am willing to believe that it might still have enough text descriptions of saurik to come up with sort of what I look like, but I doubt it? On the other side, if you remove any random handful of these profile pictures of fat hairy nerds (aka, people who look a bit like me) I doubt it really needed all of them to figure out what it needed to know. To take your example: let's say its knowledge that "throwaway1851" looks like Ben Affleck comes from your one profile picture--the only time the multi-modal embedding model was ever able to associate that word and Ben Affleck's photo into a similar location in the vector space--then if it wasn't allowed to see that photo during its training there is no way that would have happened. It simply doesn't matter if it seems to mostly rely on its knowledge of Ben Affleck to conjure up a photo of you: it has so many photos of Ben Affleck that none of them really matter anymore, but that single photo of you that made it even realize you looked like Ben Affleck in the first place is absolutely critical and probably deserves most of the attribution.
- throwaway1851 4y agoThat makes sense in a “but for” type of causality. However, I don’t think that’s what this website is aiming for (however flawed or misleading it might be so far as its methodology is concerned). I think the idea of attribution here is more of a visual concept: ie, “which images contributed visual features in the output image?” For that, if you wrote down the latent embedding for “saurik”, then retrained CLIP and Stable Diffusion from scratch without any of saurik’s profile pics in the training data, it is quite possible that you could generate an image from the embedding you wrote down and it would look the same. Pure speculation on my part - perhaps your profile pic is the real source material and this website is junk. I just think it’s an interesting and worthy question the site is trying to answer, even if it’s not possible to answer it with much certainty.