3 ms·
This would totally derail open-source LLMs and establish a SV monopolies since they have sufficiently deep pockets to probably license an awful lot of material.
by chriskanan 3y ago
This would totally derail open-source LLMs and establish a SV monopolies since they have sufficiently deep pockets to probably license an awful lot of material.
Typically, the model is trained by "reading" each document only a single time. Should you have to license in perpetuity anything you read or see online that has a copyright because it changed your brain a little? If not, why should the model be treated any differently?
What about the Phi family of models that are not directly trained on copyrighted data but instead are trained on data produced by GPT-4, among other sources.
- chrisjj 3y ago> This would totally derail open-source LLMs Would totally derail faux open-source LLMs. Genuinely open-source LLMs woukd bw fine.
- tahnyall 3y ago> why should the model be treated any differently? Because a model isn't real, it's not alive, it's not human, it doesn't know right from wrong, it's a software machine, nothing more.
- cj 3y agoThe LLM <> Human Brain argument is so frequently used but makes no sense. You can disagree, but legally speaking it will be interesting to see the legal strategy OpenAI takes with the NYT. I'm almost certain that OpenAI won't be using the human brain argument in their defense. Legally, claiming your machine should be treated the same as a human is a losing argument.
- strogonoff 3y agoLLM can work like human brain, but it doesn’t mean it is like a human person. First, we know that personhood cannot be reduced to the brain (cf. the gut flora studies and all that). Second, more generally that claim would require solving what’s known in philosophy as the hard problem. The arguments that consciousness arises from a certain arrangement of entities we perceive in reality (or that it doesn’t exist at all) is specious: it implies that those other things exist in absolute sense, i.e. are the underlying territory, even though the only evidence for their existence is coming via consciousness. That’s a lot of entities to grant magical existence to out of nothing. Once you assume that those things are more likely to be a map, the question turns into what is the territory then, and the spotlights shift to consciousness (a.k.a. the only thing we can be absolutely sure to exist, as it is required for the “we” part to exist in the first place). And yes, of course, if OpenAI uses the “sentient LLM” defense, all the power to them, but the next logical step would be to consider that industry based on forced labour.
- strogonoff 3y agoAn argument can be made that developing an LLM or running an LLM for personal use (or academic research, etc.) could be fair. Making money off of it by selling access to an LLM based on copyrighted works (whether the LLM is open or not) isn’t. If I use BT to download a music album I have bought off Bandcamp after Bandcamp goes out of business, that should be fair. If I set up a business that makes money by helping people download whatever albums they want, not caring whether they paid for them, that is obviously illegal. This leads to my position on derivative models, if they are based on models that themselves are trained on copyrighted works without proper licensing.