3 ms·
I would like to know this too. I understand that GitHub is a private company and you have to accept their T&C, but surely they aren't allowed to use source code
by bsd44 5y ago
I would like to know this too. I understand that GitHub is a private company and you have to accept their T&C, but surely they aren't allowed to use source code found elsewhere on the internet to train their ML models without asking for permission first unless it's a B2B cooperation such as with Stackoverflow.
- lacker 5y agoAccording to the discussion at this link, you do not need permission to use copyrighted data to train AI models. Copyright prevents you from copying data, it doesn't prevent you from learning from it. https://twitter.com/luis_in_brief/status/1410985742268911631?s=20 https://twitter.com/luis_in_brief/status/1410985742268911631...
- marcosdumay 5y agoTo train your model, yeah, probably ok. But I don't think anybody will see people using the duplicated code that AI insert on your codebase the same way.
- lloydatkinson 5y agoOh man I can’t imagine the consequences for certain languages and frameworks if it uses SO answers though. Imagine if it trained in all the dumb and ancient answers like “how do I get the length of a string in javascript” and took the first accepted answer of “use jquery”
- bsd44 5y agoThis raises an issue of trolling. What prevents developers to generate "inappropriate" code to feed it to this algorithm the same way they did with the Microsoft Chat bot for example? That will surely reflect on the quality of code generated by this AI system and therefore the stability and security of applications built.
- kingofclams 5y agoI’m sure this will happen, and there will definitely be instances of the bot giving users bad code, but it would be incredibly difficult to make it solely give out bad code.
- deleted 5y ago[deleted]