15 ms·
Julia Reda's analysis depends on the factual claim in this key passage: > In a few cases, Copilot also reproduces short snippets from the training datasets, ac
by codesections 5y ago
Julia Reda's analysis depends on the factual claim in this key passage:
> In a few cases, Copilot also reproduces short snippets from the training datasets, according to GitHub’s FAQ.
> This line of reasoning is dangerous in two respects: On the one hand, it suggests that even reproducing the smallest excerpts of protected works constitutes copyright infringement. This is not the case. Such use is only relevant under copyright law if the excerpt used is in turn original and unique enough to reach the threshold of originality.
That analysis may have been reasonable when the post was first written, but subsequent examples seem to show Copilot reproducing far more than the "smallest excerpts" of existing code. For example, the excerpt from the Quake source code[0] appears to easily meet the standard of originality.
[0]: https://news.ycombinator.com/item?id=27710287 https://news.ycombinator.com/item?id=27710287
- make3 5y agoIt can, that does not mean that it will, in any case other than people actively probing it for that.
- creshal 5y agoI'm not sure if making an "analysis" without doing any research whatsoever is reasonable.
- codesections 5y agoI'm not sure either —which is why I said "may have been reasonable" instead of "was reasonable" :) I can see an argument for doing your own research, but I can also see an argument for basing an analysis on what GitHub said in the FAQ — I'm honestly a bit surprised that Microsoft's lawyers let them say that with a product that can reproduce such large blocks of verbatim code.
- creshal 5y agoMy guess is that their lawyers weren't consulted, and that the Github people just shipped it on their own.
- yunohn 5y agoThat literally cannot happen in FAANG/MS, esp not when the CEO announces the product in a public blog post.
- swiftcoder 5y agoYep. Individuals in a FAANG don't have the ability to launch a product without review. Just drafting a press release for a new product involves Comms oversight and VP-level approval.
- hnfong 5y agoIt's not obvious that Microsoft is violating copyright yet. The main concern is whether the product makes others liable. So it could be that the executives really wanted to do it, and the lawyers thought "OK, technically we're not violating anything...."
- onion2k 5y agoThe example you linked to is talking about a 16 line function from the Quake source. The Quake source is 167,594 lines in total (counting the C code only). Does that really fail to meet the standard for "smallest excerpt"?
- emodendroket 5y agoNot only that, but it is clearly someone going out of their way to make it do that. I’m not sure that that is a reasonable test of how the program typically behaves.
- hmfrh 5y ago> I’m not sure that that is a reasonable test of how the program typically behaves. That's not what people care about, people care about their copyright being blatantly violated by a massive corporation _without any consequences_.
- emodendroket 5y agoOk, but is “I can go out of my way to make it misbehave” adequate proof that the copyright is being violated?
- ghoward 5y agoNot GP. Yes, it is, because that means that the algorithm will produce that copyrighted code regardless of the intent of the person who makes it misbehave. People could both accidentally and "accidentally" make it reproduce copyrighted code. In the first case, it's unintentional. In the second, how could you prove it's intentional? Because of this whole mess, I am actually adding clauses to FOSS licenses that I am writing, just to ensure that my copyright on my code is not infringed by code laundering.
- rndgermandude 5y ago>I am actually adding clauses to FOSS licenses that I am writing Doesn't this make your new licenses incompatible to a lot of existing licenses?
- denton-scratch 5y ago> appears to easily meet the standard of originality It's an algorithm. In the olden days, you couldn't copyright an algorithm, even an original one. There's only so many ways you can express an algorithm; the best ways are using code. So is it the intention that rewriting in Python an algorithm previously expressed in C would be infringing? Suppose the algorithm is re-expressed in English? Allowing copyrights on algorithms is tantamount to allowing copyrights on thought-processes. Here come the thought-police. Take cover.
- eesmith 5y agoThere are no thought police here. In the US, copyright may include the choice of variable names, the organization of the code into modules and functions, and other aspects which where there are the creative choices that may be protected under copyright law. The relevant process is described at https://en.wikipedia.org/wiki/Abstraction-Filtration-Comparison_test https://en.wikipedia.org/wiki/Abstraction-Filtration-Compari... , which comes from the court case at https://en.wikipedia.org/wiki/Computer_Associates_International,_Inc._v._Altai,_Inc https://en.wikipedia.org/wiki/Computer_Associates_Internatio.... nearly 30 years ago: > the court presented a three-step test to determine substantial similarity, abstraction-filtration-comparison. This process is based on other previously established copyright principles of merger, scenes a faire, and the public domain.[1] In this test, the court must first determine the allegedly infringed program's constituent structural parts. Then, the parts are filtered to extract any non-protected elements. Non-protected elements include: elements made for efficiency (i.e. elements with a limited number of ways it can be expressed and thus incidental to the idea), elements dictated by external factors (i.e. standard techniques), and design elements taken from the public domain. Any of these non-protected elements are thrown out and the remaining elements are compared with the allegedly infringing program's elements to determine substantial similarity. Emphasis mine. This specifically highlights that your example ('only so many ways you can express an algorithm') is not protected under US copyright law. The originality requirement only applies to other aspects of the generated code, which in this case would include the comments that Copilot generated, and which clearly are not required for the algorithm to work. For thought police like you describe, look to patent law.
- mytailorisrich 5y ago
- eterevsky 5y agoThe excerpt from Quake code is literally one of the most famous functions out there. There is no wonder that it was reproduced verbatim. The share of such code, according to Github is really small. It would be quite straightforward to write an additional filter that would check the generated code against the training corpus to exclude exact copies.
- tovej 5y agoWould it? What would the threshold be? Twenty lines copied verbatim? Ten lines copied verbatim? What about boiler plate like ten #include statements at the beginning of a file? Or licenses in comments? What if someone has a one-liner that's unique enough to be protected by copyright?
- hobs 5y agohttps://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_Inc.#Supreme_Court https://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_... I think that was a big part of the Google Vs Oracle case - how much copying constitutes an infringement? It looks like they made a fairly complex rubric to apply in the future, it appears it would be on a case by case basis.
- e3bc54b2 5y ago> The excerpt from Quake code is literally one of the most famous functions out there. There is no wonder that it was reproduced verbatim. The question that brings is that this was found because it is so famous, but what if it is repeating Joe Schmoe's weekend library project, but we will never know because its not famous?
- 0-_-0 5y agoBecause someone already checked, and it doesn't: https://docs.github.com/en/github/copilot/research-recitation https://docs.github.com/en/github/copilot/research-recitatio... Every literally quoted part that could infringe appears at least 10 times in the training data
- riedel 5y agoI would love to try a session of clean room reverse engineering using copilot. I would bet you get reasonably far for very common libraries with not much effort. The question would be if such compression/decompression would infringe copyright.
- Joeri 5y agoBut that fast inverse square root example is particularly interesting because it is also a derivative work. Carmack did not invent it, and several variations of it had been passed around over time. Algorithms should not be subject to copyright, that way lies madness. It would prevent new generations from building on top of the work of their predecessors, because copyright lasts a very long time. The amounts of code that github copilot reproduces fall squarely into the “shouldn’t be subject to copyright” domain for me, even if they pass the bar for originality.
- klodolph 5y agoSomething which is a “derivative work” is still copyrighted. In fact, by definition, a “derivative work” is copyrightable. It’s the minimum threshold at which something, based on something else, gets its own, new copyright. The algorithm is not copyrighted, but the source code of the function is copyrighted. You could learn how the algorithm works by reading the function, and then write your own function that implements the same algorithm. Algorithms are not copyrightable, they are not subject to copyright. Source code is copyrightable. Copilot is not reproducing just the algorithm, it is spitting out large chunks of the copyrighted source code, verbatim.
- modeless 5y agoThe funny thing about the Quake function is, id Software is almost certainly not the origin of the code. They copied it from somewhere else, possibly added profane comments, then slapped GPLv2 on it. Did they even have the right to do that? From an IP absolutist standpoint, probably not. https://www.beyond3d.com/content/articles/8/ https://www.beyond3d.com/content/articles/8/
- robertlagrant 5y agoIt's not the actual copying of the idea, but the verbatim reproduction of the function, comments and all. I think people somehow thought that copilot could write code, and so verbatim reproduction was surprising to them.
- jonas21 5y agoA quick search shows that this snippet, including comments, is included in thousands of Github repos [1], so it's not surprising that the model learned to reproduce it verbatim. It's such a famous snippet that it's even included in full on Wikipedia [2]. I wouldn't be surprised if the next version of Copilot filtered these out. [1] https://github.com/search?q=0x5f3759df+what+the+fuck&type=code https://github.com/search?q=0x5f3759df+what+the+fuck&type=co... [2] https://en.wikipedia.org/wiki/Fast_inverse_square_root#Overview_of_the_code https://en.wikipedia.org/wiki/Fast_inverse_square_root#Overv...
- fridif 5y agoThey did not copy the implementation, they copied the general idea of what the algorithm should do. Do not go down this line of reasoning, otherwise we will be copyrighting the concept of for loops.
- modeless 5y ago> They did not copy the implementation, they copied the general idea of what the algorithm should do [Citation needed]
- joe_the_user 5y ago[Ianal] The thing about the situation is that "copying code you found on the Internet" certainly isn't automatically, always legal. That you engaged in copying X from the Internet doesn't make illegal either. Your source for the source code your incorporate into a product doesn't matter, what matters is whether that code is copyrighted and what the license terms (if any) are (and people saying "copyright doesn't apply to machines" are wildly misinterpreting things imo). Given what's come out, it seems plausible that you could coax the source of whatever smallish open source project you wished out of copilot. Claiming copyright on that code wouldn't be legal regardless of Copilot. Whether Microsoft/Github would be liable is another question as far as I can tell. I mean, Youtube-dl can be used to violate copyright but it isn't liable for those violations. The only way Copilot is different from youtube-dl is that it tells it's users everything is OK and "they told me it was OK" is generally not a legal defense (IE, I don't know for sure but I'd be shocked if the app shielded it's users from liability). All the open source code is certainly "free to look at" and Copilot putting that on a programmers screen isn't doing more than letting the programmer look at it until the programmer does something (incorporating it into a released work they claim as their own would be act). The question is how easily a programmer could accidentally come up with a large enough piece of a copyrighted work using Copilot. That question seems to be open. TL;DR; My entirel amateur legal opinion is that Copilot can't violate copyright but that it's users certainly can.