4 ms·
AFAIK, using copyrighted data to train does not necessarily make the trained model "toxic". "Authors Guild, Inc. v. Google, Inc." case [1] is viewed as a key pr
by donpark 5y ago
AFAIK, using copyrighted data to train does not necessarily make the trained model "toxic". "Authors Guild, Inc. v. Google, Inc." case [1] is viewed as a key precedent for this view.
[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
- pabs3 5y agoThe phrase is "toxic candy" not "toxic", see the policy for what it means. Most data is protected by copyright, but I assume you meant proprietary rather than copyrighted. Using proprietary data might not matter under copyright law, but it does matter in terms of the Debian machine learning policy and DFSG, because the non-free data cannot be shipped in Debian main and thus cannot be used to train a model shipped in main.
- pabs3 5y agoHmm, that case doesn't appear to be about ML though, could you explain how it is considered a precedent for ML?
- donpark 5y agoSee https://towardsdatascience.com/the-most-important-supreme-court-decision-for-data-science-and-machine-learning-44cfc1c1bcaf https://towardsdatascience.com/the-most-important-supreme-co...
- pabs3 5y agoThanks. Its interesting that this only applies to countries with the concept of fair use, which is unfortunately not widespread.