6 ms·
Ok, my curiosity has been fired here... I have conjured up two scenarios here: Let's say I use copilot to generate a bunch of code for an app, something subst
by _Understated_ 5y ago
Ok, my curiosity has been fired here...
I have conjured up two scenarios here:
Let's say I use copilot to generate a bunch of code for an app, something substantial, and it regurgitates a load of bits and pieces from many sources it got from GitHub, I'd assume there won't be any attribution in it... it will be as if Copilot made the code itself (I know it sort of does but lets not split hairs!). I'm guessing the prevailing theory (from GiitHub anyway) is that I'm legitimately allowed to do this.
Now, let's say I generated all that code by manually copying and pasting chunks of code from a whole bunch of repos, whether they are open source, unlicensed, whatever. Would I not be ripe for legal issues? I could potentially find all the code that copilot generated and just copy and paste it from each of the sources and not mention that in my license. What if I told everyone "yeah, I just copied and pasted this from loads of Github repos and didn't put any attribution in my code". I'd assume that (morality aside) I'd be asking for trouble!
Am I missing something? Am I misunderstanding the situation, or the capabilities of copilot?
- lukeplato 5y agoCopilot is a commercial paid service that generates money for Microsoft
- _Understated_ 5y agoYeah, that bit I realise but the point I was getting at is this: if I take someone else's code, use chunks of it in my app, say that it's mine and make money from it is that not illegal? Or, at least in violation of the license? Superficially at least, Copilot (from my understanding) is "copying" code, letting me use it in my app, and making money from it. I'm just trying to wrap my head around it. Let's be clear, I am not a lawyer, but it seems... strange!
- Diggsey 5y agoAlso NAL, but I think there's far more of a case that users of Copilot might violate copyright rather than Copilot itself: - Only a very small proportion of Copilot generated code is reproduced verbatim, so if you specifically built a product just from copied-verbatim code, your act of selecting and combining those pieces of copyrighted code would be creating a derivative work. - GitHub is not selling the copyrighted code, they are selling the tool itself. Google is literally the same thing: you could theoretically create a product by googling for prefixes of copyrighted code and then copying the remainder straight out of the search results. It's you who would be violating copyright, not google.
- ghoward 5y agoI think there is an argument to be made that Copilot is producing derivative code, though. It may produce copies verbatim, and that's a violation, but far more often, it produces a mixture of things it was trained on, most of which probably have some sort of license requiring attribution at the very least.
- deleted 5y ago[deleted]
- klntsky 5y agoDoes copilot seem strange, or maybe the concept of intellectual property does?
- _Understated_ 5y agoCopilot isn't strange from a technical prespective. The strange bit is how they are allowed to use other peoples code to create derivative works (this is how I see it from my non-legal perspective anyway). Even if it's legal (to the letter of the law, not the spirit) it leaves a sour taste.
- stonemetal12 5y agoBoth the Copy machine and VCR were found to be legal because they had substantial non infringing uses. As is I don't see how Copilot does. It could, if trained on public domain or attribution free code only, unfortunately there probably isn't enough code out there to train the model adequately under such rules.
- wpietri 5y agoI think you're right. Especially given that Copilot can reproduce significant blocks of code: https://twitter.com/mitsuhiko/status/1410886329924194309 https://twitter.com/mitsuhiko/status/1410886329924194309 Famous code: https://en.wikipedia.org/wiki/Fast_inverse_square_root#Overview_of_the_code https://en.wikipedia.org/wiki/Fast_inverse_square_root#Overv...
- treesprite82 5y agoI see this held up as an example a lot, but the fast inverse square root algorithm didn't originate from Quake and is in hundreds of repositories - many with permissive licenses like WTFPL and many including the same comments. GitHub claims they didn't find any "recitations" that appeared fewer than 10 times in the training data. That doesn't mean it's a completely solved issue (some code may be repeated in many repositories but always GPL, and there are limitations to how they detect recitations), but from rare cases of generating already-common solutions people seem to be concluding that all it does it copy paste.
- wpietri 5y agoThat may be true, although even GitHub doesn't know for sure. But the problem remains: they're reproducing other people's code without regard to license status.
- schneidmaster 5y agoThere's a decent bit of caselaw indicating that computers reading and using a copyrighted work simply "don't count" in terms of copyright infringement -- only humans can infringe copyright. This article[0] does a pretty good job of summarizing the rationale that the courts have provided. My (non-lawyer) take is that GitHub is pushing this just half a step farther -- if computers can consume copyrighted material, and use it to answer questions like "was this essay plagiarized", then in GitHub's view they can also use it to train an AI model (even if it occasionally spits back out snippets of the copyrighted training data). Microsoft has enough lawyers on staff that I'm sure they have analyzed this in depth and believe they at least have a defensible position. [0]: https://slate.com/technology/2016/08/in-copyright-law-computers-and-robots-dont-count.html https://slate.com/technology/2016/08/in-copyright-law-comput...
- _Understated_ 5y agoI don't doubt that an army of lawyers has poured over this but they have size on their side: the cost of litigation vs potential revenue will be a massive factor. Edit: > There's a decent bit of caselaw indicating that computers reading and using a copyrighted work simply "don't count" in terms of copyright infringement. That means their computer can read any code it wants, do whatever it wants with the code, then they can monetise that by giving YOU the code. Would they then be indemnified by saying "no Microsoft human read or used this code"? However, if you then use the code and look at it, does that make you liable?
- schneidmaster 5y agoAgain, not a lawyer, just a guy who likes reading this stuff. The devil is usually in the details of copyright cases. The Turnitin case hinged substantially on whether Turnitin's use of copyrighted essays was "fair use". There are four factors[0] which determine fair use; the two more relevant factors here are "the purpose and character of your use" and "the effect of the use upon the potential market". The court found that Turnitin's use was highly "transformative" (meaning they didn't just e.g. republish essays; they transformed the copyrighted material into a black-box plagiarism detection service) and also found that Turnitin's use had minimal effect on the market (this is where "computers don't count" comes in -- computers reading copyrighted material don't affect the market much because a computer wasn't ever going to buy an essay). I would be shocked if GitHub's lawyers didn't argue that using copyrighted material as training data for an AI model is highly transformative. There may be snippets available from the original but they are completely divorced from their original context and virtually unrecognizable unless they happen to be famous like the Quake inverse square root algorithm. And I think GitHub's lawyers would also argue that Copilot's use does not affect the _original_ market -- e.g. it does not hurt Quake's sales if their algorithm is anonymously used in a probably totally unrelated codebase. Your counterexample would probably fail both tests -- it's not transformative use if your software hands out complete pieces of copyrighted software, and it would definitely affect the market if Copilot gave me the entire source code of Quake for my own game. [0]: https://fairuse.stanford.edu/overview/fair-use/four-factors https://fairuse.stanford.edu/overview/fair-use/four-factors
- deleted 5y ago[deleted]
- rgbrenner 5y ago"I'm guessing the prevailing theory (from GiitHub anyway) is that I'm legitimately allowed to do this." No. Copilot is a technical preview. In the final release, if it reproduces code verbatim, it'll tell you and present the correct license.
- ghoward 5y agoDoesn't matter that it's a technical preview; people are using it now, GitHub has already used it internally. So if it infringes now, there is already code out there being used that does infringe.
- rgbrenner 5y agoGitHub appears to be tracking every snippet that they're generating during their trials: https://docs.github.com/en/github/copilot/research-recitation https://docs.github.com/en/github/copilot/research-recitatio... Are you doing that? If not, then I wouldn't use GitHub's use as justification to engage in copyright infringement.
- ghoward 5y agoOh, I am not using Copilot. But other people not part of GitHub are. And those are still violations.
- x4e 5y agoHow will it find the “correct” license? Will it check the LICENSE file? Simply having a LICENSE file is not a declaration that all the code in that repo is under that LICENSE. What if specific lines/files are specified to be under different licenses? What if the publisher of the repo is publishing it under an incorrect license in bad faith? Will github be responsible if it tells me the wrong license?
- kstrauser 5y agoSuppose Copilot was Composer and it generated personalized songs for you after being trained on Spotify's library. If you started performing the resulting song and it contained recognizable clips of others, I guarantee you'd have lawyers coming after you. I don't see this as fundamentally different. It's unlikely that the Free Software Foundation is going to track you down for including some GNU code in your single-user repo. If you used their stuff in a popular commercial project and they got wind of it, you might expect to receive a cease and desist at best.
- BlueTemplar 5y agoCopilot is just a tool, legally it cannot "make code", you're the one making it. See also : Napster, including how it was condemned for facilitating copyright infringement (what Microsoft is risking here, though the offense is likely to be much milder, of course).
- alerighi 5y agoCopying/pasting code from open source projects it's considered fair use. Come on, who doesn't do that? I mean, sure you don't copy an entire file, but you tend to copy a snippet, or in the end you look at how is done and you done the exact same way (that is the same of copying it!) I would say there is not a problem in there.
- sild 5y agoIf you are copy and pasting code from open source projects into your own project, then I think that is more likely to be considered copyright infringement than fair use. Fair use is generally for things like criticism, parody, teaching etc. Obviously this kind of thing would need to be judged on a case-by-case basis, but I think you are on shaky ground here.
- devetec 5y agoCopilot isn't a retrieval model. It's a generative model. It learns the coding techniques, not retrieving snippets. Only 0.1% of code it generates is regurgitated, and even that is usually pretty common code.