5 ms·
Humans do violate copyright if they use copyrighted passages directly in their work and pass it off as their own without any attribution, which is what copilot
by iepathos 4y ago
Humans do violate copyright if they use copyrighted passages directly in their work and pass it off as their own without any attribution, which is what copilot has been show to sometimes do, though not always. Copilot will sometimes offer chunks of code that can be found verbatim in open source code bases and passes it off to users without attribution. I agree it is ok to learn from copyrighted work and reproduce new different work from a human or a machine learning algorithm, but it isn't ok to pass along exact copies as your own without attribution. Microsoft will likely need to add checks to prevent copilot from offering verbatim copies of code going forward to try to avoid copyright violations here.
- deleted 4y ago[deleted]
- Vetch 4y ago> which is what copilot has been show to sometimes do In those cases it seems that humans are already copying code without also propagating licenses appropriately. LLMs are more likely to memorize things which occur a lot (and I'd bet rare things that are representative of some conceptual axis). The main examples presented so far, Davis and Carmack, have the property of having been copied a lot. The generative model is only surfacing an existing pattern of ignoring attribution. Sort of like the code-gen version of generating bigotry if appropriately prompted. I'll also note that this pattern of retrievable memorizing of copyrighted and sensitive material is present in GPT-3 too and not just for code. As the situation is equivalent, a lawsuit should address the concerns of non-programmers too.
- Vetch 4y agoOne thing I worry about is if the uncertainty around copyright violation cools down activity in open models while raising the price of commercial offerings. Commercial entities can afford devoting resources towards mitigating copyright violations such as eating the cost of maintaining a database of frequently copied code and identifying most likely origin combined with a large semantic database of code snippets. An open equivalent might be wary of being accused of contributing to copyright violations since in that scenario, there is no way to force people to respect it.
- diffeomorphism 4y agoThat seems like a weak defense: "sure, we violated copyright but only because many other people do, too". Kinda the same problem as YouTube. Lots of people copy movies on the high seas, but if you are as big as yt you cannot easily get away with it.
- Vetch 4y agoNo, this is not meant as a defense. My point is that it's an issue that is already rampant and what Copilot (or any model) does is make it more readily visible. This is not like youtube because Github is already hosting those violations and people are already inappropriately copying or including such code. It matters not whether the local inclusion was fetched by copilot or a human fetched it using more manual steps through search.
- xfer 4y agoThe difference is if you build tools that can be used to violate existing copyright laws your tool will get taken offline by the same corporations. Yet they are selling one that can be used to do so.
- Vetch 4y agoYes, I agree there's a measure of double standards to this. It's why I feel it's important that AI does not remain in the control of just a handful of corporations. The decks are stacked against though, given how data and compute intensive SOTA is. But in defense of copilot, code regurgitation is uncommon in routine use. An editor extension allowing search of github would be at least as easy to use to violate licenses but I do not think it'd be taken down since that would not be its core offering. Copilot goes far beyond mere search and provides a useful service. GPT-3 can also be prompted into generating copyrighted works of writers but I do not see people talking as if that is its primary utility nor as much clamoring in these forums to end that service.
- xfer 4y ago
- odo1242 4y ago> not just for code This is quite important, actually, and I don't think enough people realize this. I am a photographer sometimes and it would be really cool if I could share my photos online under a copyright license that forbids their use in training AI.
- frumper 4y agoThe problem seems to be that someone could just copy your photo and repost it without that license and we're back to the same spot.
- sooyoo 4y ago> Microsoft will likely need to add checks to prevent copilot from offering verbatim copies of code going forward to try to avoid copyright violations here. Or they could integrate a way to find the produced output back in the corpus if it's sufficiently close and provide a reference/attribution. Basically whatever tool a copyright lawyer would use to track down original work. And that's just the engineering solution. The AI researcher solution would be to extend AI learning algorithms to attach attribution metadata to the learned data so that the output could already come annotated with information about the source. But the latter is much harder to do, so maybe the engineering solution would suffice.
- alpaca128 4y agoA Twitter thread linked yesterday showed that a keyword and the name of the original code's author in the Copilot prompt produced an almost exact copy of that developer's code. Copilot already does know the origin sometimes. Edit: here's the related tweet: https://twitter.com/DocSparse/status/1581632706693079042 https://twitter.com/DocSparse/status/1581632706693079042
- SahAssar 4y ago> Or they could integrate a way to find the produced output back in the corpus if it's sufficiently close and provide a reference/attribution. Basically whatever tool a copyright lawyer would use to track down original work. That assumes that the licenses of your code and the original code are compatible which often isn't the case.
- sooyoo 4y agoNo, it doesn't assume that. Ensuring that they are compatible would be the next step. Either manually by the user or automatically by showing a fat warning or retracting the suggested code completion.
- bayindirh 4y ago> Microsoft will likely need to add checks to prevent copilot from offering verbatim copies of code going forward to try to avoid copyright violations here. A user once replied to one of my comment[0] about this with the following: > It's not really an issue when you're a large software corporation; you already have mechanisms in place to check for license compliance in everything that ships, including F/OSS plagiarism checks [1]. IOW, from my understanding, they don't care. Big players do their own checks anyway, and small fish won't be creating problems because it's too convenient for them. Classic Microsoft (as I know from 90s). The bigger thread can be seen in [2]. [0]: https://news.ycombinator.com/item?id=32534697 https://news.ycombinator.com/item?id=32534697 [1]: https://news.ycombinator.com/item?id=32539467 https://news.ycombinator.com/item?id=32539467 [2]: https://news.ycombinator.com/item?id=32533531 https://news.ycombinator.com/item?id=32533531
- adamhp 4y agoI'm curious about the endgame of copyright with respect to software. At some point, enough people will have written enough code that you can't write code anymore because some fragment of it violates a copyright. Where does the line get drawn? There's only so many ways to do certain algorithms, like DFS or BFS.
- salawat 4y agoAt some point it'll have to come home to roost that code is a subset of discrete mathematics first, a literary/artistic work second. There is really no way around it.
- ghaff 4y ago>you can't write code anymore because some fragment of it violates a copyright Copyright (unlike patents in general) allows for independent creation. If I sit down to write a quicksort routine, it is going to look extremely similar to a zillion other quicksort routines out there. The other question (IANAL) is whether writing a quicksort routine is even a creative act at this point.