3 ms·
Not really. Maps are also mere data, but they are quite successfully copyrightable. There's even a concept of trap streets [0] used to find out if someone used
by matrix_overload 3y ago
Not really. Maps are also mere data, but they are quite successfully copyrightable. There's even a concept of trap streets [0] used to find out if someone used your data in their map without permission.
AI models don't have an established legal framework yet, but it's reasonable to assume that similar rules will apply here.
[0] https://en.wikipedia.org/wiki/Trap_street https://en.wikipedia.org/wiki/Trap_street
- modeless 3y agoMaps are not mere data. There is a lot of creativity involved in choosing the data to present and leave out, the style and colors, the arrangement of labels, etc. That's why maps are copyrightable. There is comparatively little creativity involved in feeding large fractions of the internet through a standard transformer model. Neither is there any significant creativity involved in the presentation of the raw weights. It's not at all clear to me that the weights are or should be copyrightable. People will certainly try, though, and like all regulatory regimes copyright loves to expand and never voluntarily shrinks, so they may succeed. Honestly I think the most likely outcome is that model weights will be ruled as derived works of the input dataset, and courts will try to enforce that people who train models must license their entire dataset specifically for model training. Some would cheer that but I personally think it would be a disaster.
- orbital-decay 3y agoThe training process is closely curated, bootstrapped, the bulk data is mixed with whatever manual data you have, and it generally requires tremendous amounts of human input and expertise for the model to be even remotely good. It's definitely not just "feeding the data to the model".
- modeless 3y agoCompiling and checking and selecting and filtering and sorting the phone numbers for a phone book and printing that book and distributing it is not trivial. It involves significant work and even a few creative and editorial decisions, but on balance it's not that creative of a process, and the resulting phone book is not copyrightable. As for mixing in your own data, if the model weights inherit the copyright of the data then every existing large language model is illegal.