7 ms·
This is interesting, but many permissive licenses still require attribution at the project or file level. If Codeium doesn't produce these when producing "verb
by jamesmunns 3y ago
This is interesting, but many permissive licenses still require attribution at the project or file level.
If Codeium doesn't produce these when producing "verbatim enough" snippets, how is this actually better, besides avoiding a GPL boogeyman?
I get that there have been fewer (if any? I'm not aware of any) MIT/Apache2.0/MPL2.0 license violations that have gone to court than GPL violations, but this still feels like an "address the symptoms" and not "address the cause" difference.
- blibble 3y ago> If Codeium doesn't produce these when producing "verbatim enough" snippets, how is this actually better, besides avoiding a GPL boogeyman? it's not if they've trained on MIT/Apache 2.0/... then they're just as liable as people that have trained on GPL they would be limited to training on licenses that don't require attribution (BSD2, public domain, etc) which I suspect limits the size of the training set so much that the output would be useless Codium here is unintentionally making an argument that undermines legal confidence in their own product interesting choice!
- gus_massa 3y agoIANAL, but I expect people to use MIT/BSD to be less angry about partial reuse of the code than people that uses GPL.
- pornel 3y agoI'm not a lawyer either, but I don't think "less angry" is a legal term.
- JamesBarney 3y agoNot but it's big determinant in how you get sued. Several lawyers haven give the advice the best way to avoid a lawsuit is don't be an asshole. The second best way is to spend a bunch of money on an attorney.
- GuB-42 3y agoI wonder if, in order to deal with attribution, the system could simply build a multi-megabyte file with "this code is derived from:" followed by all the authors the system could gather from the training data set.
- dalmo3 3y agoHey, I want my name on that list. Here's my contribution: {
- throwaway290 3y agoPerhaps copyright is what being circumvented, not just GPL. What Microsoft does is take your original work, create derivative works and sell them for profit. Unless it's under creative commons zero or public domain it shouldn't be legal...
- judge2020 3y agoNot exactly a huge distinction there because the licenses themselves provide exemptions to copyright, so by definition you're both "circumventing" GPL and committing legal copyright infringement if you copy it and don't attribute it under the terms the code's license requires. Of course, the entire basis for LLMs being legal is that they use work collectively to know how code/language works and how to write it in relation to the given context. In this case, the legal defense is that the tool is like a human that learned how to code by looking at CC-BY-SA and other licensed publicly-available code and assimilating it into their own fleshy human neural network. This only becomes shaky once you add in regurgitating code verbatim, but humans do this too, so the solution there is the copilot setting that tries to detect and revert any verbatim generated code snippets.
- anileated 3y agoYou can’t claim to have an entity with human-like understanding doing the coding if you don’t grant it basic human rights. They want to have it both ways: they want you to think the LLM is like a human because it’s “learning” (which in ML is the same word but completely different idea) so that you let them ignore copyright, but it’s not like a human of course because it can’t think, no sir (so you do not make them grant it human rights, because then they can’t exploit it like they do anymore).
- visarga 3y ago> What Microsoft does is take your original work, create derivative works and sell them for profit. Unless it's under creative commons zero or public domain it shouldn't be legal... Why should it not be legal? Doesn't that make copyright equally powerful with patents? Copyright should restrict only replication of expression not replication of ideas.
- hathawsh 3y agoAs an experiment, I just asked ChatGPT to "Please identify the open source projects that contain the following code" and pasted the sample from the article. Sure enough, it pointed me at the SuiteSparse library, which is correct (but not exactly where ChatGPT thinks it is). This means Codeium (and others) could theoretically use AI to identify possible attributions that should be included in a project. Of course, if someone figures out an algorithm that does that, people could use the same algorithm to identify missing attributions and plagiarism in other projects and throw lawsuits around. (Sigh)
- mumblemumble 3y agoYou don't need AI for that, just fuzzy search.
- etherealG 3y ago[flagged]
- jacooper 3y agoThey should just attribute code, like how amazon is already doing.
- tlavoie 3y agoIf the code is being provided as output from an LLM, I don't know that they themselves could even say where the code came from. Attribution might not be possible in that case. Similarly, how might one remove say, GPL code from the model, without regenerating from all the rest of the inputs alone?
- elproxy 3y agoAnd just the same, copyleft licenses (which are not "non-permissive" in my book) such as GPL don't "mean that you cannot without consent", they just want you to share the result back under the same license (which is often an issue for some corporate projects).