4 ms·
> which is what copilot has been show to sometimes do In those cases it seems that humans are already copying code without also propagating licenses appropriat
by Vetch 4y ago
> which is what copilot has been show to sometimes do
In those cases it seems that humans are already copying code without also propagating licenses appropriately. LLMs are more likely to memorize things which occur a lot (and I'd bet rare things that are representative of some conceptual axis).
The main examples presented so far, Davis and Carmack, have the property of having been copied a lot. The generative model is only surfacing an existing pattern of ignoring attribution. Sort of like the code-gen version of generating bigotry if appropriately prompted.
I'll also note that this pattern of retrievable memorizing of copyrighted and sensitive material is present in GPT-3 too and not just for code. As the situation is equivalent, a lawsuit should address the concerns of non-programmers too.
- Vetch 4y agoOne thing I worry about is if the uncertainty around copyright violation cools down activity in open models while raising the price of commercial offerings. Commercial entities can afford devoting resources towards mitigating copyright violations such as eating the cost of maintaining a database of frequently copied code and identifying most likely origin combined with a large semantic database of code snippets. An open equivalent might be wary of being accused of contributing to copyright violations since in that scenario, there is no way to force people to respect it.
- diffeomorphism 4y agoThat seems like a weak defense: "sure, we violated copyright but only because many other people do, too". Kinda the same problem as YouTube. Lots of people copy movies on the high seas, but if you are as big as yt you cannot easily get away with it.
- Vetch 4y agoNo, this is not meant as a defense. My point is that it's an issue that is already rampant and what Copilot (or any model) does is make it more readily visible. This is not like youtube because Github is already hosting those violations and people are already inappropriately copying or including such code. It matters not whether the local inclusion was fetched by copilot or a human fetched it using more manual steps through search.
- xfer 4y agoThe difference is if you build tools that can be used to violate existing copyright laws your tool will get taken offline by the same corporations. Yet they are selling one that can be used to do so.
- Vetch 4y agoYes, I agree there's a measure of double standards to this. It's why I feel it's important that AI does not remain in the control of just a handful of corporations. The decks are stacked against though, given how data and compute intensive SOTA is. But in defense of copilot, code regurgitation is uncommon in routine use. An editor extension allowing search of github would be at least as easy to use to violate licenses but I do not think it'd be taken down since that would not be its core offering. Copilot goes far beyond mere search and provides a useful service. GPT-3 can also be prompted into generating copyrighted works of writers but I do not see people talking as if that is its primary utility nor as much clamoring in these forums to end that service.
- xfer 4y ago> An editor extension allowing search of github would be at least as easy to use to violate licenses Well, try doing that for music or movies or proprietary leaked codebase. If you think copilot is uniquely producing things that are not that different from humans then surely no one would have any problem with feeding it massive amounts of corporate programs? I am not aware what writers are doing but there have been plenty of uproar regarding stable diffusion. I have a feeling that if any tools like this get built for musicians/film-makers, it will look vastly different from the current situation.
- Vetch 4y agoFirst, the music and to an extent movie industry enforcement of IP are uniquely pathological. But I am not talking about music or movies. I am contending that a simple search extension being much less capable than Copilot and so even more scopeable as aiding copyright violation would not be taken down. > surely no one would have any problem with feeding it massive amounts of corporate programs? There is a similar gymnastics done by human engineers today due to the issue of patents. I don't think this is a good trend to uphold. > I am not aware what writers are doing but there have been plenty of uproar regarding stable diffusion Yes but mostly in the art community. On HN there were plenty of arguments just the other day how it is not the same for art and programmers have a stronger case. I disagree but regardless, the case is exactly equivalent for GPT-3 and writers but it wasn't an issue generating about a thousand comments on respecting IP and ceasing deployment of LLMs until copilot.
- SkyBelow 4y agoWithout a sample case I can't say for certain, but couldn't it also be a defense that some code is generic enough that it shouldn't be copyrighted?
- odo1242 4y ago> not just for code This is quite important, actually, and I don't think enough people realize this. I am a photographer sometimes and it would be really cool if I could share my photos online under a copyright license that forbids their use in training AI.
- frumper 4y agoThe problem seems to be that someone could just copy your photo and repost it without that license and we're back to the same spot.