5 ms·
If you memorize all of harry potter word for word, or some famous solo vocal track from memory, are you committing a copyright violation? Or only if you then r
by harshreality 3y ago
If you memorize all of harry potter word for word, or some famous solo vocal track from memory, are you committing a copyright violation? Or only if you then recreate it and try to redistribute your copy?
The scenario where AI training is locked down doesn't result in 1,000,000 individuals getting paid. (What would they get paid, and by whom?) It results in Disney, Adobe, etc.—massive companies with existing licenses to use content just about however they want—training their own models and locking everyone else out of the large AI model training game, until AI gets good enough to start generating human-quality creative work on its own (the same kind of progression as alphago/lee to alphago/zero), perhaps with the addition of a small set of purely copyright-free material.
Excluding all copyrighted material would be tying an AI model's metaphorical hands behind its back, since humans, although capable of producing great works through much iterative effort in isolation, all rely on having learned from some copyrighted work. Find an author who hasn't read plenty of recent books as well as older classics, or a musician (other than classical) who hasn't listened to plenty of modern music, or a director or editor who hasn't watched tons of movies and films. Recall Newton, "[I]f I have seen further, it is by standing on the shoulders of giants." Many of those "shoulders" are copyrighted.
- kranke155 3y agoYes, and you know how humans acquire works to learn from? They pay for it. They buy the books. They buy tickets to theatre. They buy entrance to the gallery. The trick that's being done now is hey, we don't have to pay since it's not a person. (to the creator) But hey, it is just like a person when it learns! (legal system) If AI models require human training data, then they should pay for it. Easy.
- harshreality 3y agoFalse. Libraries exist. Borrowing books from neighborhood libraries or friends exists. Watching movies and TV with friends exists. Listening to music on the radio (yes, those free electromagnetic thingies) still exists. There are many, many, many free performances or accessible copies of all kinds of copyrighted content, plenty to train either a neural net or a human brain on. Books3 has separate legal concerns, but Google has a legally acquired corpus of tons of books, which they've mostly cleaned up from scans (probably far better than IA has), and have probably used to train Bard on. Their lawyers must be biting their nails waiting to see how these lawsuits turn out, though. Until AGI arrives, or some other method of training LLMs from the ground up on sparse examples by incrementally building on structural knowledge of language.... training on ridiculous amounts of copyrighted content is required. Not because anyone wants to copy those works, but because training that way fills in for a lack of real-world experience that every child gets, which includes consuming and interacting with a bunch of copyrighted content that isn't tracked because it's not practical to do so. You could train a LLM only on project gutenberg, and the LLM would churn out stilted English and the occasional iambic pentameter. That's great if you want works that seem like they were written over a century ago, but nearly useless otherwise.
- kranke155 3y agoLibraries exist? Do you think books fly onto library shelves for free? As far as I know, someone bought them. Your neighbour or friend also bought the stuff. I suspect you're not being straight here, I just have to ignore this whole line of reasoning since it seems so absurd. >Until AGI arrives, or some other method of training LLMs from the ground up on sparse examples by incrementally building on structural knowledge of language.... training on ridiculous amounts of copyrighted content is required. that's not my problem. Those AI model folk should just compensate the people they're using training data from, and they should ask for permission. >You could train a LLM only on project gutenberg, and the LLM would churn out stilted English and the occasional iambic pentameter. That's great if you want works that seem like they were written over a century ago, but nearly useless otherwise. Not my problem. Why are the problems of the wonderful AI developers suddenly human, global problems that we all have to find a way to fix? If they want access to training data - they should pay for the privielige.
- harshreality 3y agoI think humans should pay for permission to learn. Heaven forbid copyright holders don't get paid for all the material they've put out that humans are using (often stealing) to learn from in order to become useful members of society! Physical library books are governed by the doctrine of first sale. That's why google has one of the largest (maybe excluding l-bg-n and IA) corpus of books on the internet. They might have the cleanest corpus of OCR'd book content of anyone, since IA uses commercial or open source OCR and that's it, while google for a long time used recaptcha to check OCR results. For physical books, the cost per read of a library book is an order of magnitude smaller than the cost per read of privately purchased books. How can you tolerate the economic model of libraries when the net effect is a theft of maybe 80%-95% from the author and publisher? Libraries subsidize books that nobody wanted to read, but steal from authors and publishers whose books are read multiple times per physical copy. Even libraries' onerous ebook licenses are not commercial retail ebook pricing. They're just closer to retail pricing than the publishers could ever manage with physical books, because there's no pesky right of first sale which turns physical book libraries into piracy havens. I would prefer to get away from OpenAI and Facebook and all the other people using potentially tainted sources like books3. The obvious legal question for them isn't whether training was legal, but whether the acquisition of the training data was legal. That's a straightforward copyright issue, or at least as straightforward as fair use determinations can ever be. Whether we agree with copyright law as it stands, it's certain that copyright applies when books3 is transferred around the internet. How transformative it is, how much the transfer of books3 affects the market, and the other two factors, make those actions fair use, are the only questions to be considered. The training aspect is where all the difference of opinion lies: What is your position on Google using its corpus of books (legally acquired and possessed, as the content behind google books) to train a LLM? Do they need to acquire additional rights from copyright holders? Why, and under what legal theory? How would they get permission ahead of time? How would they agree to a pricing model? Would they spend tens or hundreds of millions of dollars training a model, and only then negotiate with rights holders to find out whether the license fees they want will be economically viable? We all know that most major rights holders would never grant a one-time license fee. It would be perpetual rent-seeking from AI output. I don't see how any of these LLM or image generation models would be economical if rights holders had their way. They wouldn't mind. They're notoriously slow to adopt tech, but if they did anything, they'd hire AI experts, build their own models, and license the models back to Google and Microsoft.
- beej71 3y agoI don't think this is the argument that's being made, though. They're not saying, "This is a clear cut case of piracy--pay me for that book." They're saying, "You can't consume my book in that way."
- ben_w 3y agoI had a free school education (including Shakespeare and Ethan Frome, both of which are out of copyright now though only the former when I studied it); several free libraries; and with the exception of my final year even my university tuition was free[0]; after graduation the museums I went to were also free; I watched free educational videos from Apple Developer and YouTube, and listened to free podcasts; I have learned things from reading Wikipedia; and I have done free online courses in both natural languages and programming languages. This doesn't cover everything: I did, indeed, also buy books on HTML and JS, and my first C compiler, and a licence to REALbasic[1]. But that doesn't refute the fact that I did learn a lot for free. > If AI models require human training data, then they should pay for it. Easy. You can do that if you like, but that won't stop any of the economic issues that arise. The cost of running Stable Diffusion is so low that even if you had literal slaves, and you were spending only the UN extreme poverty threshold on keeping them alive and housed, the pro-rata cost of keeping them alive for long enough to type in the prompt dominates the total cost of making images. Right now these models are still, despite their impressiveness, flawed: while an artist can use them to great effect, most of us will have our generations easily spotted by some flaw we have never trained ourselves to notice. If the models become good enough to fully replace all artists, the only way the profession called "artist" isn't going to go the same way as the profession called "computer" is if the arts are to humans as fancy tails are to peacocks: the effort being the point, extravagantly wasting effort to show you're fit enough to manage fine despite the penalty. [0] UK rules at the time, thanks to my dad's early retirement and therefore "low income" status [1] as it was so named at the time, Xojo now
- kranke155 3y agoIf we let this idea that "AI training data usage has no compensation for rights owners" to be become ensconced in the legal system, then all human endeavour will become fair game to be acquired by someone to make a Machine Intelligence out of, and remove you completely out of the profit loop of your own work. This will happen in every industry and occupation, one by one. Is this what you think is desirable? The alternative is perversely simple: PAY for the right to use training data.
- 3y ago
- deleted 3y ago[deleted]
- cmiles74 3y agoWhere is this idea that copyrighted material should be excluded from training data coming from? My understanding is that people want to be compensated when their intellectual property is used as training data for a machine. That strikes me as an entirely reasonable expectation. One person memorizing Harry Potter for their own amusement, even if they make money doing public appearances where they recite sections of the work verbatim for the amusement of the audience is not in any way similar to the process of training an LLM or of that LLM's output. The scale alone is so vastly different that it renders the comparison useless and misleading.
- creer 3y ago> people want to be compensated when their intellectual property is used as training data for a machine. That's fine that they want to. The question is whether copyright law gives them that and that's very unlikely.