3 ms·
If I take a copy of your art to hang on my wall, I've violated your copyright. But if I "copy" the experiential knowledge of your art into my brain by viewing
by alexwebb2 3y ago
If I take a copy of your art to hang on my wall, I've violated your copyright.
But if I "copy" the experiential knowledge of your art into my brain by viewing it, I'm not violating your copyright. My brain doesn't contain a copy of the art, it's just been influenced by viewing it, and I might be more capable of producing art that mimics your style.
What these models are doing feels, to me, vastly more like the second case.
- ToucanLoucan 3y agoBut you are not capable of reproducing from your mind something that is in that artists' style, after that viewing. If you show someone a painting and ask them to recreate it, even setting aside the skill gap, you do not get the same painting back. You get a different painting, with some overlap, with the "focus" of it being what the person paid the most attention to during their viewing. In this way, ML training is just not the same as a person being inspired by or even being asked to recreate another person's creative work. Over the course of making something, an artists' "voice" would be best characterized I feel as the tiny choices they make along the way that all point to and reinforce a larger point or purpose to the piece. This "voice" shows up in all creative output, not just spoken word. That is what I feel people are feeling is lacking in generated art: because a machine-learning model does not have a voice, it does not have intent, it has a mandate from a third party from which it tries to draw from, and instead of making numerous, tiny but contributory choices, it instead decides on a weighted average of all the choices made in the art that the model was trained upon, which is simply not the same thing. It makes all generated art have this very sterile, soulless feeling to it because these tiny choices that would otherwise be made by a person trying to illicit an effect are instead just the machine sort of shrugging and being like "well in most things I've seen where a woman is sitting this way, her hand is tilted this way" but it doesn't know why the hand is tilted or what that means for the subject, which means the hand-tilt might be applied to subjects for whom it makes absolutely no sense at all to tilt the hand.
- spywaregorilla 3y ago> But you are not capable of reproducing from your mind something that is in that artists' style, after that viewing. I am capable of doing this. > If you show someone a painting and ask them to recreate it, even setting aside the skill gap, you do not get the same painting back. You get a different painting, with some overlap, with the "focus" of it being what the person paid the most attention to during their viewing. same for ai models. in fact, more true for ai models. You'll have a much harder time recreating an image from popular models than from human memory. > It makes all generated art have this very sterile, soulless feeling to it because these tiny choices that would otherwise be made by a person trying to illicit an effect are instead just the machine sort of shrugging and being like "well in most things I've seen where a woman is sitting this way, her hand is tilted this way" but it doesn't know why the hand is tilted or what that means for the subject, which means the hand-tilt might be applied to subjects for whom it makes absolutely no sense at all to tilt the hand. this will surely age well
- Kim_Bruning 3y ago> this will surely age well People are already experimenting and/or deploying systems that run an LLM before doing text2image [1]. This helps a lot with 'understanding'. The next low-hanging fruit would then be to go back and improve the labeling used for training the image-generation models in the first place (using new multi-modal LLMs). [1] eg. the current GPT+ standard model will do Prompt -> GPT-4 -> DALL-E