4 ms·
We are working on a compensation model for Stock contributors for the training content, and will have more details by the time we release. The training is base
by mesh 4y ago
We are working on a compensation model for Stock contributors for the training content, and will have more details by the time we release.
The training is based on the licensing agreement for Adobe contributors for Adobe Stock.
(I work for Adobe)
- egypturnash 4y agoThanks! I am delighted to know that Adobe's got plans on that front.
- roughly 4y agoI would be very, very interested to see a compensation system that took into account the outputs of the trained model - as in, weights derived from your work are attributable to X% of the output of this system, and therefore you are due Y% of the revenue generated by it. It sounds like Adobe is taking seriously the question of artist compensation, and I'd love to see someone tackle the "Hard Problem" of actual attribution in these types of systems.
- brookst 4y agoI've looked a few times, but have not seen any research on assigning provenance to the weights used in a particular inference run. It's a super interesting space for a bunch of reasons. But the naive approach of having a table of how much each individual training item influenced every weight in the model seems impossibly big. For DALL-E 2's 6.5B parameters and 650m training items, that's 4.2 quadrillion associations. And then you have to figure out which weights contributed the most to an output. I would love to see any research or even just thinking that anyone's done on this topic. It seems like it will be important in the future, but it also seems like a crazy difficult scale problem as models get bigger.
- bash-j 4y agoCould you not use tags used to label the image? If your image contains more niche tags that match the user input, your revenue share will be higher. Depending how much extra people earn for certain tags, it might incentivise people to upload more images of what is missing from the training data.
- brookst 4y agoThat's interesting, but I'm not sure it works. I think that works out to "for any given prompt, distribute credit to every source image that has a keyword that appears in the prompt, proportional to how many other source images had that same keyword". If I include the tag "floor", do I get some (tiny) percentage of every image that uses "floor" in the prompt, even if the bits from my image did not end up affecting model weights much at all in training? Worse, for tags like "dramatic lighting", it's likely that the important source images will depend on the other words in the prompt; "sunset, dramatic lighting" will probably not use the rely on the same weights or source images as "theater interior, dramatic lighting". And then you get the perverse incentives to tag every image with every possible tag :) I'd love to be convinced otherwise, but I'm not seeing prompt-to-tag association working.
- bash-j 4y agoThe tags could be added by a model rather than the user submitting the image. Maybe do both and verify the tags with a model? Users could get a rating based on how reliably they tag their pictures and are trusted to add more niche tags at higher ratings. You could even help tag other pictures to improve your rating.
- gradys 4y agohttps://arxiv.org/abs/1703.04730 https://arxiv.org/abs/1703.04730 > How can we explain the predictions of a black-box model? In this paper, we use influence functions -- a classic technique from robust statistics -- to trace a model's prediction through the learning algorithm and back to its training data, thereby identifying training points most responsible for a given prediction.
- natch 4y ago> And then you have to figure out which weights contributed the most to an output. Why do you have to, though? What do you hope to trace back to exactly?
- brookst 4y agoThe business problem is “which items from the training corpus contributes the most to this inference”
- astrange 4y agoThat is impossible. You might be able to do it if you invented a completely different method of image generation, but the amount of original images present in a diffusion model is 0% with reasonable training precautions, and attributing its weights to particularly any of its input is nearly arbitrary. (Also, it's entirely possible that eg a model could generate images resembling your work without "seeing" any of your work and only reading a museum website describing some of it. Resemblance is in the eye of the beholder.)
- jamilton 4y agoI think the naive approach of just dividing revenue equally across all contributors could be acceptable, and would have lower overhead costs.
- pabs3 4y agoWhat about the openly licensed content you mentioned above?
- easyThrowaway 4y agoHope I can get better rates from this compared than those offered by Spotify to my musician friends. How much are we talking about? 0.00001¢ per licensed generated image?