4 ms·
I am very happy to see the starting to think about and gather the community around this novel issue. This sheds a new light on existing IP laws. It's pretty ob
by lta 4y ago
I am very happy to see the starting to think about and gather the community around this novel issue. This sheds a new light on existing IP laws.
It's pretty obvious that without the open source code corpus, such tool would not have been possible, hence the (IMHO) justified derivative work question.
Companies and individuals have spent decades building this open corpus and in exchange they deserve to have their will (aka license) respected. For a significant fraction of those, it means sharing the derivative works under the same license.
It's pretty sad, though not surprising to many of us, to see that Microsoft isn't really playing openly and nicely with the FOSS community about those issues.
When people started saying Microsoft had changed and was a fair player now, I had my doubts, and this doesn't help.
- tzs 4y ago> It's pretty obvious that without the open source code corpus, such tool would not have been possible, hence the (IMHO) justified derivative work question. Note though that a work not being possible without your work is not sufficient to make that work a derivative work of your work. It just suggests that you need to take a closer look at the relationship between your work and the other work. For example Windows applications, even ones that make intimate use of the behavior of Windows and would take significant rewrites to port elsewhere or to run under current Windows compatible operating systems like ReactOS or under things like Wine, are not automatically derivative works of Windows. To be a derivative work the work has to include copyrighted elements from your work in a way that is not covered by fair use. That's why clean room reverse engineering works--by making sure the coders do not have access to the work being reverse engineered they cannot copy any copyrighted elements from it and so cannot produce a derivative work. I suspect that under current copyright law it is possible to do something like Copilot without the output violating copyright but it may need to be more sophisticated than the current Copilot. From the few examples of Copilot output I've seen it seems to output stuff that would probably either be covered by fair use or that doesn't have enough creativity to be copyrighted. But from what people have said it occasionally spits out longer things that seem likely to be copyrighted and not covered by fair use. What may be necessary for systems like this is to couple them with a second AI that can recognize when the first AI is making a suggestion that goes beyond fair use and stops it. I don't know if it is currently possible to make such an AI. Where would you get a good set of training data? The above was about the output of Copilot. Another question is whether Copilot itself is legal. When you train an AI on some data is there a copy of that data in the AI? If there is then Copilot may be an infringement of the copying right. In the US copyright law defines copies in 17 USC 101, where they are defined as > [...] material objects, other than phonorecords, in which a work is fixed by any method now known or later developed, and from which the work can be perceived, reproduced, or otherwise communicated, either directly or with the aid of a machine or device. The term “copies” includes the material object, other than a phonorecord, in which the work is first fixed. Is a collection of neural net weights something from which you can perceive, reproduce, or otherwise communicate the individual works the net was trained on? Or is it more like some kind of hash of the work? My guess is that both the output of AIs and the AIs themselves are sufficiently beyond what anyone was contemplating the last time there was a major update of copyright to deal with new technology that to fit AI in we are probably going to need a major update to the law.
- lta 4y agoThe points you are making are excellent. I don't have much to add, but I wanted to thank you. The second one is particularly interesting and brings load of very interesting questions. Like, is it truly learning or just reciting ? From my limited knowledge about copilot, the latter might be more likely, so the copyright law might be triggered ?
- jstummbillig 4y ago> It's pretty obvious that without the open source code corpus, such tool would not have been possible How is that obvious? I am relatively certain that MS committing their entire code base as training material would do the trick, if that's what it took. Or, additionally, maybe licensing some other huge high quality code bases (restricted to just training the ai) for a few million bucks? It's not like it would be an issue to find vendors happy to remonetize their already written code. Given what's at stake here and who sits at the helm, I don't see how Copilot would not or could not be moved forwards regardless.
- abirch 4y agoTo be honest, I wish that better code repositories we're given more weight. Microsoft's should be given more weight, especially considering the security issues that is autosuggested by copilot. https://arxiv.org/abs/2108.09293 https://arxiv.org/abs/2108.09293
- californical 4y agoIf that would’ve worked, then they should’ve done that instead though! Why create such huge possibility for legal battles for themselves (and their users)? I would feel much more comfortable with the idea of using copilot if I was sure that I wouldn’t generate copyright infringing code, which Microsoft could guarantee if they had licensed the training data. There would be huge benefits to going that route, but they didn’t. That’s what makes me think that it’s impossible because they needed a huge volume of data, and the only way to get enough was to take it without consent of the license owners.
- jstummbillig 4y ago> Why create such huge possibility for legal battles for themselves (and their users)? Simple: They disagree with you on the risk of that happening and the potential cost if it did.
- lta 4y agoI wish they were as careful of other's people's code IP as they're with their own. If this is not derivative work, how come they didn't use their own codebase instead of taking the opensource ones ? Either their codebase is low quality, either they think it would leak their IP, meaning it's derivative work. Neither solution sounds great for them
- fartcannon 4y agoEspecially when it barfs out niche code with only a few samples verbatim.
- mirntyfirty 4y agoCopilot, I’m looking to build an email client.... “How about Outlook366?”
- lobocinza 4y agoThe old "embrace, extend, extinguish".