3 ms·
I think humans should pay for permission to learn. Heaven forbid copyright holders don't get paid for all the material they've put out that humans are using (o
by harshreality 3y ago
I think humans should pay for permission to learn. Heaven forbid copyright holders don't get paid for all the material they've put out that humans are using (often stealing) to learn from in order to become useful members of society!
Physical library books are governed by the doctrine of first sale. That's why google has one of the largest (maybe excluding l-bg-n and IA) corpus of books on the internet. They might have the cleanest corpus of OCR'd book content of anyone, since IA uses commercial or open source OCR and that's it, while google for a long time used recaptcha to check OCR results.
For physical books, the cost per read of a library book is an order of magnitude smaller than the cost per read of privately purchased books. How can you tolerate the economic model of libraries when the net effect is a theft of maybe 80%-95% from the author and publisher? Libraries subsidize books that nobody wanted to read, but steal from authors and publishers whose books are read multiple times per physical copy.
Even libraries' onerous ebook licenses are not commercial retail ebook pricing. They're just closer to retail pricing than the publishers could ever manage with physical books, because there's no pesky right of first sale which turns physical book libraries into piracy havens.
I would prefer to get away from OpenAI and Facebook and all the other people using potentially tainted sources like books3. The obvious legal question for them isn't whether training was legal, but whether the acquisition of the training data was legal. That's a straightforward copyright issue, or at least as straightforward as fair use determinations can ever be. Whether we agree with copyright law as it stands, it's certain that copyright applies when books3 is transferred around the internet. How transformative it is, how much the transfer of books3 affects the market, and the other two factors, make those actions fair use, are the only questions to be considered.
The training aspect is where all the difference of opinion lies:
What is your position on Google using its corpus of books (legally acquired and possessed, as the content behind google books) to train a LLM? Do they need to acquire additional rights from copyright holders? Why, and under what legal theory?
How would they get permission ahead of time? How would they agree to a pricing model? Would they spend tens or hundreds of millions of dollars training a model, and only then negotiate with rights holders to find out whether the license fees they want will be economically viable? We all know that most major rights holders would never grant a one-time license fee. It would be perpetual rent-seeking from AI output. I don't see how any of these LLM or image generation models would be economical if rights holders had their way. They wouldn't mind. They're notoriously slow to adopt tech, but if they did anything, they'd hire AI experts, build their own models, and license the models back to Google and Microsoft.