3 ms·
In my opinion it was covered under the GitHub terms of service and is clearly transformative. I am optimistic that the courts will find it so and we can put the
by Permit 4y ago
In my opinion it was covered under the GitHub terms of service and is clearly transformative. I am optimistic that the courts will find it so and we can put these debates to rest similar to how we’ve done for web scraping.
- jacquesm 4y agoSorry but terms of service don't give you a blanket license to re-purpose someone else's copyrighted work at your pleasure. By that same token any hosting provider could make a small change to their terms of service and suddenly all of the data of all of the customers would be theirs. Copyright does not work that way, you need to actively sign away your rights.
- Permit 4y ago> Sorry but terms of service don't give you a blanket license to re-purpose someone else's copyrighted work at your pleasure. That’s correct. It’s also not what happened here. I don’t violate your copyright when I scrape your web page. I don’t violate your copyright when I train a model on a web page that I scraped. I violate your copyright only if I use that tool to produce code that violates your copyright. This is similar to how I don’t commit a copyright violation when I read your code, I commit the violation when I produce and publish code that violates your copyright.
- jacquesm 4y agoFirst you wrote: > In my opinion it was covered under the GitHub terms of service and is clearly transformative Now you change the argument completely and let's just say that even with that change that is not how I understand that copyright works and leave it at that. You're very welcome to your own interpretation. Best of luck.
- belorn 4y agoIf the training data discussion ends on that note than there is exist a large upside. Youtube and other video hosting services has a massive collection of the worlds images, music and sound. Anyone who who want to build an AI will have access to practically unlimited amount of training data, meaning the barrier to entry will be quite low.
- leereeves 4y agoI doubt YouTube will let anyone download more than a trivial fraction of the videos they host, but it has been used for some specialized projects like Minecraft AI (70,000 hours of video). https://openai.com/blog/vpt/ https://openai.com/blog/vpt/
- belorn 4y agoApple made that interpretation regarding their app-store. It is also the reason why GPL software is not accepted on the store. Authors would need to give apple a full copyright transfer agreement (or exemptions of the GPL terms), including any dependencies, and that was simply more costly than just banning GPL from the store.
- kmeisthax 4y agoThe GitHub ToS only applies a blanket license to copy as a fallback if you forget to stick a license file in your code - i.e. it's there so that if you post code to GitHub you can't sue GitHub for doing the thing you told it to do. Furthermore, that only covers stuff intentionally uploaded to GitHub by the author; people mirroring repos hosted elsewhere do not have the legal capacity to license attribution-free, copyleft-free AI training.