3 ms·
You're right that there's not much precedent to this, which is why I do think that "training rights" is a valid proposal, as long as the proponents acknowledge
by CaptainFever 2y ago
You're right that there's not much precedent to this, which is why I do think that "training rights" is a valid proposal, as long as the proponents acknowledge that this would be an expansion of current IP laws (i.e. further away from the ideal of a limited copyright).
Though... if you say "And the result, despite many years and billions of dollars worth of work, is something that can’t even reliably reference or summarize the material it’s trained on.", doesn't this imply that there's not much to worry about here? I sense that this is a negative jab, but this undermines the original argument that there is so much worth in the model that we need to create new IP laws to handle it.
I mean, I'm not sure what to make of this statement in the first place. Training data should be for the model to learn language and facts, and referencing or summarizing the material directly seems to be out of scope of that. One tends to summarize the prompt, not training data.
- angoragoats 2y ago> the original argument that there is so much worth in the model that we need to create new IP laws to handle it No one argued this, to my knowledge. I think that there might be a need for new copyright laws, but the alternative in my mind is that we decide there's not a lot of worth there, meaning that we do nothing, and what OpenAI/Meta/MS/Google/Anthropic/etc are doing is simply de jure illegal. The statement I made about LLMs having major flaws is a point in support of this alternative. > Training data should be for the model to learn language and facts, and referencing or summarizing the material directly seems to be out of scope of that. I strongly disagree, as your prompt can (and for a certain type of user, often does) contain explicit or implicit references to training data. For example: * Explicit: “What is the plot of To Kill a Mockingbird by Harper Lee?” * Implicit: “How might Albert Einstein write about recent development X in physics research?”