4 ms·
Scene releases are in human-consumable formats. Books as lines of text are not human-consumable (by reasonable humans). There's a strong case for sharing this
by harshreality 2y ago
Scene releases are in human-consumable formats.
Books as lines of text are not human-consumable (by reasonable humans).
There's a strong case for sharing this being fair use, even if distributing hundreds of thousands of ebooks would ordinarily violate copyright.
I don't know what OpenAI's training corpus is or how they got it, but in case anyone has forgotten, Google has its own license-free collection of scanned books. It had already run that through OCR (and probably periodically re-OCRs) to generate its google books indexes, so there is no copyright argument against Google's book-trained AI models (assuming they train on books).
An argument that Books4 is piracy is an argument for an oligopoly on good AI models.
- bayindirh 2y ago> There's a strong case for sharing this being fair use, even if distributing hundreds of thousands of ebooks would ordinarily violate copyright. If you're using the resulting model trained with this for (academic) research (independently or at a university), that's true. If you're earning money with the model (cough chatGPT, Bard, Gemini, et. al cough), that's false. Edit: Original version of the above paragraph started as "If you're doing...", making the below comment true. Edit is made to clarify the point further.
- harshreality 2y agoIf you're complaining about the training, you're confusing the entity doing the sharing in this case, with the entities doing the training. ETA: If you're complaining about use of the model, this becomes a dispute over how close to an original work an output can get before it becomes a copyright violation. That seems unresolvable. Look at court cases where musicians sue over short sequences of notes. Nobody knows where the threshold should be for that, for poetry, for flash fiction, for academic papers, for novels, for paintings, or for voice samples. Courts just go with whatever they think feels right in a particular case. That is an untenable legal situation in a world with AI models like these. Also, to reiterate, Google has its own fulltext, license-free corpus of books, which it undoubtedly has used to train models. What's your view on that? Should Google be the only entity allowed to train on large books datasets because it happened to get one through a loophole of collaborating with libraries to OCR their collections?
- bayindirh 2y ago> You're confusing the entity doing the sharing... Yeah, you're right. I fixed my comment to clarify my point, thanks. > Also, to reiterate, Google has its own fulltext, license-free corpus of books, which... Loopholes doesn't invalidate/override morals. I find training of any model with a corpus without consent from their respective authors immoral, regardless of laws around it. Same for code, images, sound, text, whatever. These systems can leak their training data, or recreate them verbatim with the correct prompts. Furthermore, "claimed clean" datasets like "The Stack" are not clean by any means because the dataset is huge and tools are not good enough. Plus, even if the licenses allow this, there's still morality/consent aspect to it. BTW, Fair Use explicitly defines the use should be non-profit. If you profit from the resulting model, it's no fair use by definition. Any large artistic work which can be reproduced by these systems should be ingested with consent, period. To be clear, I use none of the AI systems available on the market today.
- harshreality 2y agoI understand your moral position. But you must realize that's not what copyright does. That's wishful thinking, the sort often manifest in modern written works on the copyright page where the publisher helpfully writes, "The moral right of the author has been asserted". Begging the question: What moral right, other than the legal copyright that was already asserted separately? Note the irony, too: It's the publisher writing about morality on behalf of the author. It's the publisher that would give consent for any AI-training use in your concept of the proper legal and moral order of things. The author probably doesn't know much about copyright law, but wants the work monetized as much as possible, because they like having a house and eating food, and sure, creating more content sometimes. To that end, they've assigned copyright to their publisher. Or independent creators would have to create a union and assign licensing power to the union leadership. But again, only licensing to big tech, because nobody else could afford it. Also see the ETA: paragraph in GP, regarding the difficulty of judging whether simple works are copyrightable, or whether complex works have actually been reproduced in a copyright-violating way. This feels like an argument over angels on the head of a pin. Look at the markets. They're all-in on AI. Big tech (and big startup) AI models will move forward, no matter what licensing agreements between publishers and AI behemoths, or copyright exceptions, end up being necessary. The only question is whether others who don't have that power or money should be able to try to train their own models (admittedly probably much lower-parameter, because they don't have datacenters full of H100s) on datasets that are at least in the same ballpark.