3 ms·
Given the (AFAIK) questionable copyrightability of LLM-generated code (a number of jurisdictions do not allow assignment of copyright to non-humans nevermind th
by mathstuf 2y ago
Given the (AFAIK) questionable copyrightability of LLM-generated code (a number of jurisdictions do not allow assignment of copyright to non-humans nevermind the derivative work question of its inputs), these proclamations (Arch and Debian have had similar discussions) sound, to me, like it is mostly a clarification of already-disallowed activities (where copyright must be able to be attributed or otherwise traceable).
Basically, if you're using Copilot and making constructive[1] contributions, there's likely some human element involved where copyright can be applied. If you're just slapping together a pipeline to generate and submit patches without oversight, this just says "don't do that" more strongly than already exists for "don't contribute crap please". I see it as a way to help stem the flow of an LLM-generated
deluge of contributions that flood the review queue with sub-par work that just ends up wasting precious reviewer time. If this discourages LLM-script-kiddies from flooding FreeBSD with such things and instead doing it for other projects, it seems like a win for FreeBSD to me.
[1] https://xkcd.com/810/ https://xkcd.com/810/
- hedora 2y agoThe law hasn’t been settled yet, but I strongly suspect that the courts will rule that when an LLM memorizes an input and outputs something substantially identical, that the copyright of the original is maintained. If not, then we can just bias these models toward memorization, input harry potter books or whatever, and then declare output text that’s 99.95% identical to the input to be public domain. The obvious problem this creates for groups like NetBSD is that there will be copyright trolls that leverage the lack of provenance of LLM output in order to extort people that use these tools.
- mathstuf 2y agoI suspect that the LLM creators feel the same way, deep down. Otherwise I want to know why Microsoft, for instance, doesn't train Copilot on the Office and Windows codebases.