6 ms·
You are probably correct, legally. But given that we're talking about Aaron Swartz, who was legally in the wrong but morally (in the classic hacker's sense) in
by CaptainFever 2y ago
You are probably correct, legally. But given that we're talking about Aaron Swartz, who was legally in the wrong but morally (in the classic hacker's sense) in the right, I meant "copyright allows you to sell artificial scarcity" in the moral sense.
I think fundamentally we have a difference in opinion on what copyright is supposed to be about. I hold the more classic hacker ideal that things should be free for re-use by default, and copyright, if there is any, should only apply to direct or almost-direct copies up to a very limited time, such as 10 years or so. (Actually, this is already a compromise, since the true ideal is for there to be no copyright laws at all.)
- angoragoats 2y agoI generally agree with you about the desired use for copyright, but what gives me pause is the scale at which training AI models uses copyrighted material. I’m not sure there’s a precedent in the past to compare it to. And the result, despite many years and billions of dollars worth of work, is something that can’t even reliably reference or summarize the material it’s trained on.
- CaptainFever 2y agoYou're right that there's not much precedent to this, which is why I do think that "training rights" is a valid proposal, as long as the proponents acknowledge that this would be an expansion of current IP laws (i.e. further away from the ideal of a limited copyright). Though... if you say "And the result, despite many years and billions of dollars worth of work, is something that can’t even reliably reference or summarize the material it’s trained on.", doesn't this imply that there's not much to worry about here? I sense that this is a negative jab, but this undermines the original argument that there is so much worth in the model that we need to create new IP laws to handle it. I mean, I'm not sure what to make of this statement in the first place. Training data should be for the model to learn language and facts, and referencing or summarizing the material directly seems to be out of scope of that. One tends to summarize the prompt, not training data.
- angoragoats 2y ago> the original argument that there is so much worth in the model that we need to create new IP laws to handle it No one argued this, to my knowledge. I think that there might be a need for new copyright laws, but the alternative in my mind is that we decide there's not a lot of worth there, meaning that we do nothing, and what OpenAI/Meta/MS/Google/Anthropic/etc are doing is simply de jure illegal. The statement I made about LLMs having major flaws is a point in support of this alternative. > Training data should be for the model to learn language and facts, and referencing or summarizing the material directly seems to be out of scope of that. I strongly disagree, as your prompt can (and for a certain type of user, often does) contain explicit or implicit references to training data. For example: * Explicit: “What is the plot of To Kill a Mockingbird by Harper Lee?” * Implicit: “How might Albert Einstein write about recent development X in physics research?”