9 ms·
OpenAI now tries to hide that ChatGPT was trained on copyrighted books
- rvz 3y agoUnsurprising. Even worse than the Getty situation since OpenAI knew that they trained on copyrighted books for ChatGPT without the permission from the authors rather than getting a license just like they did with Shutterstock for DALLE-2. The AI grift continues. Copying the same Silicon Valley playbook with a new narrative for promoting their AI snake-oil.
- fswd 3y agoI'm not immediately recalling the Getty situation, could you perhaps enlighten me?
- ciabattabread 3y agoAI-generated images were re-creating the Getty Images watermark (a transparent grey rectangle).
- TillE 3y agoI think the conclusion is slightly wrong: they're not trying to hide the training process. They'll probably wind up vigorously defending that in court however they can, they're in big trouble if they can't train like that. They're trying to avoid reproducing copyrighted text, which is a totally separate (and arguably more clear-cut) legal question. Input vs output.
- ZeroZeroOneZero 3y ago> Input vs output. This is what I've been wondering. Does Fair Use apply here at all? Sure, the models were trained on copyrighted material. But wouldn't the generative part of the AI count as transformative?
- jameshart 3y agoI was trained on copyrighted books. I only get in trouble when I spout out paragraphs from them from memory and pass them off as my own. I know not to do that, though. Seems only fair that the same should apply to GPT.
- ZeroZeroOneZero 3y agoAgreed. Honestly, this is starting to feel like copyright holders using the Big, New, Scary AI as a strawman to attack Fair Use.
- bko 3y agoI don't get it. Isn't that what people wanted to happen? Don't produce copyrighted work in the output. Sure you can learn from it, much like a director might learn from hundreds of movies he's watched. He obviously can't copy the plot from Die Hard but he can use elements he's picked up from it to make a Christmas movie.
- SoftTalker 3y agoWhat do you mean he can't copy the plot? Any commercially successful movie is using one of about five basic plots.
- coolandsmartrr 3y agoCould you explain the five basic plots? While the number of plot structures are fairly countable, a plot can still be considered different by changing part of its contents (character, setting, etc.). The decision on whether a plot outright infringes on an existing plot varies case-by-case. An example of outright copying is "Fistful of Dollars" directed by Sergio Leone, which lifts the plot from "Yojinbo" by Akira Kurosawa.
- SoftTalker 3y ago1. David vs. Goliath 2. Romeo and Juliet 3. Robin Hood 4. Crime and Punishment 5. maybe there are only four.
- hnben 3y ago> 5. https://en.wikipedia.org/wiki/Hero%27s_journey https://en.wikipedia.org/wiki/Hero%27s_journey
- coolandsmartrr 3y agoInteresting. Is there a consensus of these story archetypes among writers?
- cypherpunks01 3y ago
- WillPostForFood 3y agoPapyrus Duplico! I read Harry Potter, am I now not allowed any magic spell names without paying a fee to JK Rowling for the training?
- FridayoLeary 3y agomy question is how spells were possible before latin was invented.
- whyenot 3y agoI was also trained on copyrighted books. What’s the problem? Isn’t the whole purpose of a book to be read?
- everly 3y agoPeople keep comparing it to a human ingesting content throughout their life and then being influenced in their own works. I'm sorry but that is not the same thing - at all. The concept of training a model with the explicit intent of selling the output of that model is inherently different. Not saying that it should be illegal. But it is clearly in violation of the spirit of existing copyright law, in my opinion. They set out with the intent to make money, using copyrighted input. Seems pretty simple. See dragonwriter's comment for the articulate version of what I'm saying
- artninja1988 3y agoIt is not illegal to make money with copyrighted content. Nor should it be
- everly 3y agoIf you don't procure the appropriate license for that content then yes, it is (exceptions for things like parody notwithstanding). For example, clearance of music samples.
- metalspot 3y agoit would seem to me that from a technical perspective the weights of an AI model trained on copyrighted material would be a reproduction of the copyrighted work. just because you combine the information from millions (or more) copyrighted works together doesn't mean that you aren't reproducing them. the process of training requires reproduction and distribution of the works internally as part of the data processing pipeline so why wouldn't you need a license for that?
- subw00f 3y agoIf I publish an article on a subject I extensively read about on books and add no new information, I’m just reproducing them. Should that be considered a violation of copyright?
- bathtub365 3y agoThe analogy would hold if a single human could read all books and then be copied an infinite number of times with little effort and respond to an infinite number of prompts simultaneously. This is fundamentally different than the example of a single person reading books and being inspired by them to produce something. The scale is important and I see this lost in a lot of discussions.
- akasakahakada 3y agoSimply no. Logic must be consistent from 1 to infinity. If someone read and speak too fast is illegal, then let's do it. Put everyone has IQ > 130 into jail. Oh Rainman cannot forget things, obviously that is illegal, shot it down.
- metalspot 3y agofalse analogy. if openai is downloading and copying and distributing copyrighted material internally and then compressing that information into a model that can reproduce it later and selling access to that model that is a very different thing.
- w10-1 3y agoThe article's focus on copyright is inflammatory, but the paper does present some big questions. Original paper: https://huggingface.co/papers/2308.05374 https://huggingface.co/papers/2308.05374 Abstract: major categories of LLM trustworthiness: reliability, safety, fairness, resistance to misuse, explainability and reasoning, adherence to social norms, and robustness. Each major category is further divided into [...] a total of 29 sub-categories. The measurement results indicate that, in general, more aligned models tend to perform better in terms of overall trustworthiness The paper shows that trustworthiness is now a design goal. It would seem good to avoid e.g., hallucinations, but is it really socially good to induce reliance? Compare AI companies targeting "trustworthiness" today with social media companies targeting "engagement" in decades past: they might increase uptake, but build significant externalities into their business model, and then lose most of the social benefits of adoption.
- ada1981 3y agoIf we don’t reign in our fetish with copyrights, China will be happy to become the world leader in AI.
- thrill 3y agoSo was I and everyone I know.
- archo 3y agohttps://archive.is/8BKRF https://archive.is/8BKRF
- eloop 3y agoIf LLMs are soon to become AGIs they will have as much right to read copyright works as anyone else. Is a post scarcity world where parasitic rent seeking would be pointless something to wish for, or do we consider it harmful?
- MagaMuffin 3y ago[dead]
- hititncommitit 3y agoI think the interesting question is- who is responsible when copyrighted material is provided as an output and is inadvertently used, openai, or the user?