6 ms·
Systems like Copilot and Dall-E and so on turn their training data into anonymous common property. Your work becomes my work. This may appeal to naive people (
by bugfix-66 4y ago
Systems like Copilot and Dall-E and so on turn their training data into anonymous common property. Your work becomes my work.
This may appeal to naive people (students, hippies, etc.), for whom socialist/communist ideas are attractive, but it's poison in the real world because it eliminates the reward system that motivates most creative work. People work hard for credit or respect, if they're not working for money.
Ask yourself, why does the MIT License (https://opensource.org/licenses/MIT https://opensource.org/licenses/MIT) contain the following text?
Copyright <YEAR> <COPYRIGHT HOLDER>
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
These systems are a mechanism that can regurgitate (digest, remix, emit) without attribution all of the world's open code and all of the world's art.
With these systems, you're giving everyone the ability to plagiarize everything, effortlessly and unknowingly. No skill, no effort, no time required. No awareness of the sources of the derivative work.
My work is now your work. Everyone and his 10-year old brother can "write" my code (and derivatives), without ever knowing I wrote it, without ever knowing I existed. Everyone can use my hard work, regurgitated anonymously, stripped of all credit, stripped of all attribution, stripped of all identity and ancestry and citation.
It's a new kind of use not known (or imagined?) when the copyright laws were written.
Training must be opt in, not opt out.
Every artist, every creative individual, must EXPLICITLY OPT IN to having their hard work regurgitated anonymously by Copilot or Dall-E or whatever.
If you want to donate your code or your painting or your music so it can easily be "written" or "painted", in whole or in part, by everyone else, without attribution, then go ahead and opt in. Most people aren't so totally selfless.
But if an author or artist does not EXPLICITLY OPT IN, you can't use their creative work to train these systems.
All these code/art washing systems, that absorb and mix and regurgitate the hard work of creative people must be strictly opt in.
I say this as a person who writes deep-learning parallel linear algebra kernels professionally.
We've crossed a line here.
- deleted 4y ago[deleted]
- gpderetta 4y agoExactly, any coder and artist should learn from scratch without absolutely any exposure to existing code, work of art or even idea. Anything else is outright STEALING! Excuse me when I make an apple pie from scratch.
- qull 4y agoThats a reducto ad absrudium at best. While you have a point, schools and even museums are generally compensated for providing these training models to the public, to look at it in a ml way.
- gpderetta 4y agoSure, I also had to pay for the books I studied from, but Dr. Tanenbaum is yet to knock at my door to assert copyright on all the code I have written.
- CharlesW 4y agoThis is a perfect example because, depending on the apples you're using, growing them may have required a license and adherence to licensing requirements. https://mnhardy.umn.edu/apples/licensing https://mnhardy.umn.edu/apples/licensing https://provarmanagement.com/cosmic-crisp/ https://provarmanagement.com/cosmic-crisp/
- polotics 4y agoHave you maybe possibly previously been exposed to the concept of an argumentation straw man? Feeding actual works of art into an approximation machine, and no expecting the output of said machine to not be owned by the author of the art is making a big assumption I think. There is the word copy in "copyright" and the model did definitely got a copy of the original at source. No matter the dilution, copyright is being breached, as I understand it.
- ysavir 4y agoOut of curiosity, how would you feel if someone fed your HN comment history into a ML model, then used that to respond on every HN topic and conversation under the username "othergpderetta"?
- p0pcult 4y agoNFTs to the rescue.
- dzink 4y agoFor generative art, trademarking your name might help prevent people from using it in prompts, but for general copyright, where does the line stand between someone casually publishing every color in the rainbow, every note combination, every letter in the alphabet, and claiming anyone else is infringing on their copyright? If someone copies your thesis, abstract, poem word for word, that is clear violation of your IP, but we are all remixing words that everyone uses, colors, brush strokes, API terms, programming language keywords, and notes. Copyright law has the fair use doctrine and transformative use is explicitly allowed to allow iteration. There is some level of granularity that is essential to creativity - otherwise one entity can copyright all possible combinations and prevent any creativity from happening legally. If AI goes below that threshold, all of humanity has a chance to iterate far faster and find new spaces and fill new needs for everyone. Humans have been able to draw in the style of Picasso or Monet for centuries. A program doing it is not infringement, just much faster iteration.
- burkaman 4y ago> A program doing it is not infringement, just much faster iteration. "Much faster" is absolutely relevant, morally and legally. Visiting a website a bunch of times is not illegal, programmatically DDoSing it is. Having a private conversation with someone and writing down what they said afterwards is not illegal, but recording the conversation and perfectly reproducing it without their permission often is. Shouting at someone in public is generally ok, having a drone follow them around anytime they're in public playing a recording of whatever you shouted is probably not ok. Computers are not people. Just because it's ok for a person to do something, doesn't mean it's ok to have a computer do the same thing a billion times per second.
- dzink 4y agoDDoSing is bad not because of the speed but because it overwhelms the infrastructure a product is designed for. Doing it to your own computer by making it crunch AI models until it runs out of memory is perfectly legal and iterative. Printing pages of a book in seconds vs dedicating lives of people to hand draw each letter in the monastery is iteration. Computers are not people. Computers are iteration tools people use to free up precious lifetime they have and bring more value to the world. If you are a human who trains 10000 hrs to invest like Paul Graham, or draw like Thomas Kincade, or play the piano, or operate as a top brain surgeon, you have spent a fraction of your life to do this fast and reap the rewards. But that fraction of your life has tremendous cost on society. Many people paid with their time and money to feed you, teach you, house you, during that time and during your upbringing which allowed you to have those 10k hrs to dedicate to this task. Now all of that work can be used by you to do exponentially more with your precious life. Instead of spending days or years making a portrait, you’d spend seconds. Now you can find higher purpose and solve much bigger problems - instead of asking for 100 to hand draw a portrait for a few hundred people in your lifetime, you could create one for every teen who needs a boost in their self esteem and raise their confidence and ability to cope with challenges in their life at massive scale. More importantly, since there is a huge scarcity of people trained to fulfill each niche need that forms a bottleneck on society’s capacity to use that. Imagine if instead of airplane we counted on a few trained supermen to fly people who needed to cross places fast by hand. How many people would die before they see the world or are taken to a doctor, etc. The world can’t survive on superheroes or super trained people. The world can do more with the time and lives of people in it.
- kleer001 4y agoAt the moment there's no legal protection for style in an of its self. Additionally there maybe (and should be) if this style-capture actually displaces the artists they're apeing. But I don't see that happening. IMHO its a tempest in a teacup. Why? Because Ai generated "art" is a soupy mess and real life human artists can speak and understand colloquial language, work quickly, and develop new styles based on new direction. But then again maybe we're looking at the death of a widespread industry like when gigantic industrial looms came on the scene, but I highly doubt it. Then again, last of all, I do see a future where AIs generate full feature length photo-real movies in minutes based on prompts and cheaply.
- polotics 4y agoI disagree that creative individuals have to do anything explicit here: copyright law is pretty clear that the burden of proof of right is with the copier, not the copied. I expect most artists won't be sending invoices for licensing fees just yet, but corps surely will bleed dry anyone that produce unlicensed derivative works that generates any income.
- kerblang 4y agoAs the artist in the article points out, the artwork in the model doesn't belong to her and by current legal standards she has no authority to give permission; of course the corporate owners do have authority, and I'm not even sure you need new laws to enforce the copyright complaint. I was complaining about all of this when the derivation was based on "the internet" and everyone was being ripped off at once. All the AI-generated art out there is doing the same thing. Of course most of this is being used to create derivations of trendy pop art, so are we really losing anything? Was there ever any hope for artistic capitalism as something that communicates in meaningful ways beyond the most local of scale?
- avereveard 4y agoyou can get copilot to regurgitate copyrighted code verbatim, but I haven't seen stable diffusion recreating copyright works yet, which is quite an important difference.
- ROTMetro 4y agoWasn't it regularly spitting out whole watermarks?
- avereveard 4y agowas the image underneath watermarked, or it just reproduced the watermark style over an unrelated image?
- dzink 4y agoAI acts like an alien more than a copier in that case. If you tell stable diffusion to make you clip-art of people in a conference room, you will probably see humans that don't have noses or fingers, or have 4 arms and some unreadable text that looks like a watermark going across them. The AI parameters have been trained on millions of clip-art images and assume the image should have whatever statistically applies to most images that have the same keywords on them. You can't make it copy an image, even if you tried. You can't even get it to fix the face or hands without additional processing with differently trained models. It sees as an alien would.
- vikingerik 4y agoSerious question: What is the difference between a human intelligence looking at a work and using concepts from it in their own, compared to an artificial intelligence doing it? If the copyright violation occured by the AI's inputs looking at the work... how is that different than an image of the work landing on a human's retinas?