4 ms·
The training data contains most likely insane amounts of copyrighted material. That’s why virtually none of the “open models” come with their training data
by Nesco 2y ago
The training data contains most likely insane amounts of copyrighted material. That’s why virtually none of the “open models” come with their training data
- enriquto 2y ago> The training data contains most likely insane amounts of copyrighted material. If that is the case then the weights must inherit all these copyrights. It has been shown (at least in image processing) that you can extract many training images from the weights, almost verbatim. Hiding the training data does not solve this issue. But regardless of copyright issues, people here are complaining about the malicious use of the term "open source", to signify a completely different thing (more like "open api").
- tempfile 2y ago> If that is the case then the weights must inherit all these copyrights. Not if it's a fair use (which is obviously the defence they're hoping for)
- anon373839 2y agoAlso, fair use is just one defense to a copyright infringement claim. The plaintiff first has to prove the elements of infringement; if they can't do this, no defense is needed.