3 ms·
> Sorry but terms of service don't give you a blanket license to re-purpose someone else's copyrighted work at your pleasure. That’s correct. It’s also not wha
by Permit 4y ago
> Sorry but terms of service don't give you a blanket license to re-purpose someone else's copyrighted work at your pleasure.
That’s correct. It’s also not what happened here. I don’t violate your copyright when I scrape your web page. I don’t violate your copyright when I train a model on a web page that I scraped.
I violate your copyright only if I use that tool to produce code that violates your copyright. This is similar to how I don’t commit a copyright violation when I read your code, I commit the violation when I produce and publish code that violates your copyright.
- jacquesm 4y agoFirst you wrote: > In my opinion it was covered under the GitHub terms of service and is clearly transformative Now you change the argument completely and let's just say that even with that change that is not how I understand that copyright works and leave it at that. You're very welcome to your own interpretation. Best of luck.
- belorn 4y agoIf the training data discussion ends on that note than there is exist a large upside. Youtube and other video hosting services has a massive collection of the worlds images, music and sound. Anyone who who want to build an AI will have access to practically unlimited amount of training data, meaning the barrier to entry will be quite low.
- leereeves 4y agoI doubt YouTube will let anyone download more than a trivial fraction of the videos they host, but it has been used for some specialized projects like Minecraft AI (70,000 hours of video). https://openai.com/blog/vpt/ https://openai.com/blog/vpt/