3 ms·
I think for code it will be pretty hard, unless the writer gives explicit permission to train the model on the copyrighted data, like the case for Copilot in Gi
by serverlessmania 4y ago
I think for code it will be pretty hard, unless the writer gives explicit permission to train the model on the copyrighted data, like the case for Copilot in GitHub.
And I don't think it is fair use at all, imagine a company like OpenAI train the model on its own internal docs and code, then you'll be able to ask the model to replicate ChatGPT and copilot, or even closed software like Photoshop.
- williamcotton 4y agoI'm sorry, what don't you think is fair use, Supabase Clippy?
- serverlessmania 4y agoI'm not talking about Supabase Clippy, but more about training models on copyrighted data without asking for permission (like private copyrighted code in GitHub for example and yes, I don't call that fair use)
- williamcotton 4y agoSupebase Clippy uses the same trained model as Copilot, OpenAI’s GPT family of large-language models, including having trained on all of the code in GitHub, without having asked permission and without regard to copyright license. These tools are being released under the assumption that training large-language models will be found as fair use of any copyrighted works. Are y’all starting to see the arguments for why the model and the outputs of the model are two different issues and that the models themselves, and the products built on top of them, will be considered fair use and that the liability for copyright infringement lays completely with the person using the tool?