28 ms·
GitHub Copilot is not infringing copyright
- tyingq 5y agoGuess round 2 will have Copilot dumping to AST, changing function and variable names, then dumping back to source.
- MattIPv4 5y agoThis seems to completely ignore the fact that we've seen Copilot regurgitating exact copies of existing code, and even with the incorrect license attached when it was asked for it. [0] [0] https://twitter.com/mitsuhiko/status/1410886329924194309 https://twitter.com/mitsuhiko/status/1410886329924194309
- sfletcher 5y agoThe Google Books case cited here allowed Google to show exact snippets (extracts) from the copyrighted books, hard to see how this is any different.
- creshal 5y agoIt also has no relevance for the discussion at hand. Yes, Github can display all of its content – that's kind of the point of it. But Copilot doesn't exist to show you random code snippets for the sole purpose of showing them. Using this copyrighted material to create derivative works is a completely different use case, and not covered at all by the Google Books ruling, or any other I'm aware of.
- mjburgess 5y agoYou wouldn't be allowed to make derivative works of those books; i.e. copy/paste into your own work. Google isnt making new books, or enabling people to make derivative copies; it is merely previewing a book. Github search is a preview. Copilot is a copy/paste.
- duckmysick 5y agoWhat about Google Books Ngram Viewer? Isn't that a derivative work based on copyrighted content? It's more than just a search or preview - it contains both novel information and snippets of existing content. Is linguistic corpus a special case? https://books.google.com/ngrams https://books.google.com/ngrams
- pessimizer 5y agohttps://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,.... The research value actually turned activity that would be infringing into activity that was not infringing. Take away things like the ngram viewer and Google Books infringes.
- timdaub 5y agoAgree. I've written a comment on her blog about this. Hoping she'll enable it. I've published an opinion piece on the subject matter myself: https://rugpullindex.com/blog#BuiltonStolenData https://rugpullindex.com/blog#BuiltonStolenData Edit: My comment was enabled: https://juliareda.eu/2021/07/github-copilot-is-not-infringing-your-copyright/#comment-38373 https://juliareda.eu/2021/07/github-copilot-is-not-infringin...
- bootlooped 5y agoThat function exists in hundreds, if not thousands, of GitHub repositories. The function is so well known it has it's own Wikipedia page. If there is a more famous function in computer science, I don't know what it is. The fact that a machine trained on GitHub repositories might reproduce such common code is not alarming or surprising to me. I think people are using this as an example and implying it's happening all over the place, but I've yet to see another example like it.
- user5994461 5y ago>>> If there is a more famous function in computer science Maybe fizzbuzz? Let's try to auto generate some fizzbuzz code, see what we get :D
- chrisseaton 5y agoBut doesn’t Copilot generate verbatim copies of entire copyright methods that implement non-trivial novel algorithms, including comments? The article doesn’t seem to address this?
- creshal 5y agoYes. The author apparently did no research of her own and just assumed Github's FAQ was trustworthy.
- chrisseaton 5y agoData mined and regenerated it.
- ramraj07 5y agoYes and you did no further research on your own since GitHub already said it's going to fix that (and a competent engineer would know it's trivial to fix that as well).
- creshal 5y agoBlindly trusting PR announcements isn't "research" the second time around either.
- phoe-krk 5y ago> On the other hand, the argument that the outputs of GitHub Copilot are derivative works of the training data is based on the assumption that a machine can produce works. This assumption is wrong and counterproductive. Copyright law has only ever applied to intellectual creations – where there is no creator, there is no work. This means that machine-generated code like that of GitHub Copilot is not a work under copyright law at all, so it is not a derivative work either. The output of a machine simply does not qualify for copyright protection – it is in the public domain. That is good news for the open movement and not something that needs fixing. This is very good news. This line of thought implies that we can legally feed all proprietary code into GitHub Copilot in order to teach it all the patented and secret tricks of the companies we can see (since data mining is not copyright infrigement) in order to have it print those secrets back when we ask it to (so they become public domain). /s
- tgsovlerkhgsel 5y ago"Patented and secret tricks" are not protected by copyright, if the output was an actual reimplementation of an idea instead of Copilot regurgitating existing code (https://news.ycombinator.com/item?id=27710287 https://news.ycombinator.com/item?id=27710287). The specific implementations are protected by copyright, and the ideas may be protected by patents. In the case of "secret" tricks, they may be protected by trade secret laws, but not if it's in a public GitHub repo.
- jakobdabo 5y ago> The output of a machine simply does not qualify for copyright protection Good, does the `cp` or `cat` command qualify for the "output of a machine"? Now I can uncopyright everything, hooray. What about converting a video or an image to another format? Again, it's just output of a machine. Added: Really, I would've been happy if this was the situation, as I'm, in general, against patents and copyright (in the form that they are now being used).
- bsza 5y agoExcept Copilot itself is not open source, so your only way to feed that proprietary code into it would be to upload it to github, which would make you an infringer.
- creshal 5y ago> On the other hand, the argument that the outputs of GitHub Copilot are derivative works of the training data is based on the assumption that a machine can produce works. This assumption is wrong and counterproductive. Copyright law has only ever applied to intellectual creations – where there is no creator, there is no work. Cool. I'll just train my new AI on 20 different copies of the same Disney movie and have it generate a new movie. Checkmate, lawyers!
- alpaca128 5y ago"The model might be slightly overfitted, but no creator, no work"
- arcturus17 5y agoYou do understand that's not how laws work in general, right? The court of law would probably have you unveil and tear apart your process and find that you were trying to plagiarize in a roundabout way.
- hmfrh 5y ago> You do understand that's not how laws work in general, right? The only reason the law doesn't work this way for Microsoft Copilot is because the copyright holders are individuals who do not have the capital or expertise to file suit. If Microsoft instead released a video editor addon that was trained on Disney movies and which would sometimes insert scenes of _any_ Disney movie you can bet your ass we wouldn't be having the same discussion.
- visarga 5y agoComparing code to movies - in code even a single char difference can change the meaning of everything, in movies - you can skip whole scenes and still get the meaning. I don't think the two are compatible, they are judged by different standards.
- creshal 5y ago> The court of law would probably have you unveil and tear apart your process and find that you were trying to plagiarize in a roundabout way. Well, yes, that's the point.
- codesections 5y agoJulia Reda's analysis depends on the factual claim in this key passage: > In a few cases, Copilot also reproduces short snippets from the training datasets, according to GitHub’s FAQ. > This line of reasoning is dangerous in two respects: On the one hand, it suggests that even reproducing the smallest excerpts of protected works constitutes copyright infringement. This is not the case. Such use is only relevant under copyright law if the excerpt used is in turn original and unique enough to reach the threshold of originality. That analysis may have been reasonable when the post was first written, but subsequent examples seem to show Copilot reproducing far more than the "smallest excerpts" of existing code. For example, the excerpt from the Quake source code[0] appears to easily meet the standard of originality. [0]: https://news.ycombinator.com/item?id=27710287 https://news.ycombinator.com/item?id=27710287
- make3 5y agoIt can, that does not mean that it will, in any case other than people actively probing it for that.
- creshal 5y agoI'm not sure if making an "analysis" without doing any research whatsoever is reasonable.
- codesections 5y agoI'm not sure either —which is why I said "may have been reasonable" instead of "was reasonable" :) I can see an argument for doing your own research, but I can also see an argument for basing an analysis on what GitHub said in the FAQ — I'm honestly a bit surprised that Microsoft's lawyers let them say that with a product that can reproduce such large blocks of verbatim code.
- creshal 5y agoMy guess is that their lawyers weren't consulted, and that the Github people just shipped it on their own.
- sascha_sl 5y agoI frankly think that the "free culture" label and extremely permissive licenses of many open source project are nothing but a redistribution of wealth upwards. Those with existing capital can make profitable unfree derivative works without any benefit to original authors. This relationship must go both ways if you want actual free culture. Stop producing MIT/BSD code in your non-work time. This is not a research project, this is a commercial work that produces verbatim copies of code without disclosing its license (or having a license grant in many cases). It doesn't matter how it manages to reproduce it either. It does.
- wccrawford 5y agoThis is the closest anyone has ever come to convincing me to use GPL instead of MIT license. But I still want to support small developers with anything I produce for fun, and I'm not willing to give that up to spite the big developers. For instance, I wrote a small class to load OBJ files in Unity because I needed it for an idea. I went ahead and put it on Github for others that need it, too. I could easily see someone having an idea similar to mine that needed that and couldn't find it out there. (I think there are more libraries like that now, though.) I wanted them to feel comfortable using it, even if they eventually make money with their game. If a big corp uses that code, too, that sucks. But there's no good way to draw that line in a license, so I didn't. Having said that, in the future I could see releasing some software that I don't think anyone should profit from, and in that case I'd GPL it. Previously, I'd have just defaulted to the same MIT license. I'm just not sure what that'd be yet.
- enriquto 5y ago> This is the closest anyone has ever come to convincing me to use GPL instead of MIT license. > But I still want to support small developers with anything I produce for fun, and I'm not willing to give that up to spite the big developers. Another line of thought that may help you choose a license: Do not think only about other developers. Think about the final users of your code that will be running your algorithms on their computers. The GPL protects the right of these users to see and modify the code they run (your code). So-called permissive licenses, on the other hand, let middlemen to strip this right from your users. Users of your code are in fact freer thanks to copyleft licenses.
- zxcb1 5y agoOpen source developers deserve the same rights as corporations. As a side note, in a not so distant future there may be decompilers enhanced by artificial intelligence.
- uCantCauseUCant 5y agoI felt a great disturbance in the AI-community, as if millions of voices suddenly cried out in terror, of GPL Code in there output, and were suddenly silenced. I fear something terrible has happened.
- ClumsyPilot 5y agoJulia is one the few MEPs that properly engages with issues of copyright and is active in IT. I really appreciate it, even if I dont always agree with her
- chrisseaton 5y agoBut she doesn’t seem to have engaged - she seems ignorant of basic facts of what the technology is doing in practice if you read the other comments here which give specific examples.
- elcapitan 5y agoFormer MEP, btw.
- toyg 5y agoThis is actually why I was so disappointed by her analysis having some very glaring errors. With friends like these...
- CyberRabbi 5y agoPoliticians are very rarely trustworthy. Who funds her?
- ocdtrekkie 5y agoShe meets with tech company lobbyists pretty regularly according to her meeting log.
- mabbo 5y ago> Copyright law has only ever applied to intellectual creations – where there is no creator, there is no work. This means that machine-generated code like that of GitHub Copilot is not a work under copyright law at all, so it is not a derivative work either. The output of a machine simply does not qualify for copyright protection – it is in the public domain This is fantastic news. I'm going to create a bot that crawls sites like GitHub searching for popular libraries. Then it will copy them- sans any license- to it's own website where it will sell these libraries under a new name. Since there is no creator here, just a piece of software, then there is no copyright violation. My system simply is "inspired" by the original source code using a proprietary algorithm that I call "Copy and paste". I'm open to accepting venture capital for this project.
- jdright 5y agoYou know what, I love this idea! We can do the same with music with very few adaptations to the algorithm. This idea is worth gold!
- verelo 5y agoId expect that since that’s your intention, as described above, then it is copyright and you would be the actor. In the GitHub case, it isn’t the primary intent but rather a byproduct of the goal of helping another developer? Not a lawyer, but spent enough time with lawyers to know that what you’re describing won’t fly. I don’t even know what GitHub is doing will fly, maybe they’re hoping it gets tested.
- 5y ago
- SXX 5y agoI think it's time for someone to train AI on leaked proprietary code and source-available code like Unreal Engine. It's cool that we have so much of it right now. Then we'll see how fast Microsoft and others will shut it down.
- flazx 5y ago"This is a slightly modified version of my original German-language article first published on heise.de under a CC-by 4.0 license." Heise appears to be quite $bigcorp friendly recently.
- bennyp101 5y agoCountdown to Oracle lawsuit in 3, 2 ...
- captaincaveman 5y agoIf I understand what is being stated correctly; even if I assert a prohibition in my licence for my creative work (code) not to be used by Copilot (or any other machine learning model as training data), it wouldn't matter as its not covered by Copyright?
- denton-scratch 5y ago> The output of a machine simply does not qualify for copyright protection "Simply"? If it were that simple, surely that would mean that the output of the Unix "cp" program would not qualify? What about a DVD copier? I'm OK with copyright as it used to be, back when I was a teenager; the right expired with the author's life. Corporations couldn't own copyrights. There was no burden on the author to register their rights. And copyright was a civil matter; you sued for actual damages. Infringement wasn't a crime. I'm not OK with modern copyright law, with criminal penalties, rights that can be transferred to entities that are essentially immortal, and copyright terms that keep getting extended, just before Mickey Mouse and Elvis Presley become public domain.
- tzs 5y ago>> The output of a machine simply does not qualify for copyright protection > "Simply"? If it were that simple, surely that would mean that the output of the Unix "cp" program would not qualify? What about a DVD copier? It means that the output of "cp" does not qualify for copyright protection as a derivative work. The output is still a copy of the input and would be subject to the same copyright as that input. Roughly, a derivative work is a new work that incorporates some copyrightable elements from a previous work. The derivative work gets its own copyright separate from the copyrights of those incorporated elements.
- adriancr 5y ago> Roughly, a derivative work is a new work that incorporates some copyrightable elements from a previous work. By this logic: - someone could go and copy functions and/or entire files from GPL code bases and use them with a different license. - someone could use copilot or similar to learn from all available GPL code. Is resulting code GPL? - someone could use copilot or similar to learn from open source code of their competitors that license doesn't allow them to use. Are the results legal?
- mthoms 5y agoThe definition of a "derivative work" as stated is correct. The copyright status of a derivative work is a separate issue: A derivative work can be considered infringing, and a derivative work can be considered non-infringing (ie. due to Fair Use).
- yunohn 5y agoHere we go again, a legal expert weighs in with a long and detailed post about Copilot; And HN rallies to criticize it because Copilot can reproduce some snippets when forced to.
- joshuaissac 5y ago> a legal expert weighs in [...] And HN rallies to criticize it That is an appeal to authority. Being a legal expert does not excuse one's writing from critical analysis. In this case, the post does not address Copilot reproducing large segments of copyrighted code verbatim. That is valid criticism.
- yunohn 5y agoIt is not an appeal to authority. I'm saying the expert is providing a legal explanation, and HN is throwing anecdotes around. There is no logical fallacy since HN refuses to even have a logical discussion about Copilot.
- IncRnd 5y agoThe GP is correct. It is a logical fallacy that by definition is an appeal to authority, and this is the logical discussion.
- floatingatoll 5y agoThis discussion is heavily biased and prioritizes people's emotional need to be credited and/or paid for their work over a discussion of the legal and ethical concerns at play here. It disregards the comments of an expert in the field and focuses instead on demands that may well be unsupported by copyright law. For example, GitHub license section D.4 specifically grants GitHub the right to display your content, analyze your content, and reproduce it in full to other users of the service. Yet no one seems particularly interested in discussing that here today, because it isn't compatible with the outrage that people are prioritizing on HN when discussion Copilot. I would have expected HN to be better than Reddit in this regard, but I'm not seeing it yet. I don't know if the expert is right or wrong here, but nothing in today's comments suggests anything new or curious that hasn't already been ranted about in every prior thread about this topic. I specifically care about copyright law and it's disappointing to see HN having a group tantrum instead of a discussion. https://docs.github.com/en/github/site-policy/github-terms-of-service#d-user-generated-content https://docs.github.com/en/github/site-policy/github-terms-o...
- treffer 5y agoWell, I have a hard time drawing a line between GitHub Copilot and a compression algorithm. If you can reproduce a verbatim copy of Quake source code after taking that source code as input before then that's compression. A really fancy, but still. And given that it reproduces the source code: it has to hold that somewhere. It would be very interesting if someone could reproduce the Quake example with AGPL code, then request the whole model + code because it clearly contains the AGPL code in some encoded form.
- abriosi 5y agoSome purists may say learning is compressing
- Syzygies 5y agoYes! In every form, lossy compression is distilling meaningful information from noise. This is a great legal question as it concerns our use of machine agents. We can learn from copyrighted literature or code that we read. Why can't our agents?
- AlotOfReading 5y agoBecause the process is different. You and any computer agent are allowed to learn the functional, non-copyrightable elements of fast inverse sqrt. When you need that functionality, you can write code that implements your understanding of those non-copyrightable elements and gain copyright over the resulting creative expression. What you can't do is copy all of the creative expression in the original (such as comments) without complying with the terms of the license. Moreover, reproducing the magic constants is a strong indication that your process didn't independently derive your code because the constants used in the original are unique and non-optimal.
- anticensor 5y agoI should include a term in my licenses that licensees explicitly waive their rights to fair use and/or fair dealing.
- truffdog 5y agoIf Microsoft is confident that Copilot is not a parrot, they should include their proprietary codebases in the training database.
- BuildTheRobots 5y agoDoes anyone know which codebases got included? I get the impression copilot scraped github - but as it's an internal tool, did it only scrape public repos or has private repos also been slurped?
- xdennis 5y agoThere are torrents of leaked Windows source code. Someone with access to Microsoft Copilot could try to see if reproduces the code there.
- dleslie 5y agoIf they truly believe copilot does not produce derivative works, then there is no downside to indexing their own code in its entirety; it would probably improve copilot's behaviour. Well, Microsoft, show us you believe your own arguments!
- hnfong 5y agoA counter argument that Microsoft can use is: "The code we write at Microsoft is so bad that it will decrease the quality of the generated output". :)
- rektide 5y agoI've seen way too many screenshots of a dozen-line complete XHR wrappers being suggested[1] to complete a function to imagine Copilot as a generative machine. It's a somewhat fancy copy paste engine, with phenomenal search. But it's smuggled through enough complexity & machinery to obfuscate any legal obligations that might be attached to the original source material. The article does not set itself up to address this at all: > Since Copilot also uses the numerous GitHub repositories under copyleft licences such as the GPL as training material, some commentators accuse GitHub of copyright infringement, because Copilot itself is not released under a copyleft licence, but is to be offered as a paid service after a test phase. I'm all for discussion of whether Copilot itself has to be copyleft. But to me, the immediate concern is that Copilot seems like a way to take copyleft works and remove the copyleft license from those works. [1] https://mastodon.social/@cjd/106513694972486353 https://mastodon.social/@cjd/106513694972486353
- hnfong 5y agoIsn't that what GPL allows, and is what the AGPL is for if you don't want people to take your code and host it as an online service?
- rektide 5y agothe service itself is a source-code-copier. GPL does not permit you to copy source code without attribution. this copier does not provide attribution. as i just said, i'm not so interested in debating the source-code-copier's licensing. i think it could go either way but i don't really care. the copied source code that the source code copier copies is interesting to me, and feel like the stochiastic parrot act bullshit they are pulling is massive massive sinfully evil bullshit without attribution. the stochiastic parrot can't just ignore all the licensing of what it parrots out.
- jordigh 5y ago> The output of a machine simply does not qualify for copyright protection Wolfram disagrees, and he's got lawyers and money too. Whom do we believe? http://www.groklaw.net/article.php?story=20090518204959409 http://www.groklaw.net/article.php?story=20090518204959409
- Mindwipe 5y agoWolfram is discussing American law, Reda European. (I'm still not sure I agree with Reda, but the point is at least arguable under European law and depends on the circumstances).
- jordigh 5y agoBut there's a Berne convention that kind of unifies copyright around the world, right? It's not like something can be copyrighted in one country but not another.
- Mindwipe 5y agoThe Berne convention sets down some basic principles, but there are an awful lot of edge cases and it is very much the case that things can be copyrighted in one country but not another. Heck, the duration of copyright isn't even uniform around the world.
- jfmc 5y agoModern AI seems more like machine-assisted collage (or pictures, code, text, etc.) than anything else. Someone (of some other algorithm) needs to be added to ensure that the whole thing makes sense. The big problem here is that when an artist creates a collage he/she knows the sources. Here provenance is lost. [1] Collage (/kəˈlɑːʒ/, from the French: coller, "to glue" or "to stick together";[1]) is a technique of art creation, primarily used in the visual arts, but in music too, by which art results from an assemblage of different forms, thus creating a new whole.
- temac 5y ago> What is astonishing about the current debate is that the calls for the broadest possible interpretation of copyright are now coming from within the Free Software community. It is not astonishing at all given: * proprietary codebase have not been indexed by copilot (at least a public version of it) * arguably derived code will be used in proprietary programs
- dleslie 5y agoYah, not sure what is astonishing about outrage in response to what appears to be a method for laundering GPL'd software. Copilot ought only to have indexed public domain, wtf, and other wide-open licensed software. They should remove all GPL'd software from their model, even if that means retraining from scratch.
- TeMPOraL 5y agoIt's not just GPL, they arguably should remove MIT, BSD and most other Open Source software too, as it's hard to tell when any given snippet crosses a threshold where the original license demands attribution or other things. People seem to forget that even MIT license has actual conditions in it.
- dento 5y agoNot just GPL, even works with MIT/Apache/BSD license require attribution
- makecheck 5y agoOf course derivative works are being produced!! Whether you blame Copilot or the developer using it, the result is something that required the original developer of the code in order to be constructed. Have we reached the point where every “class X” must become “class X_GPL2_CopyrightJohnQSmith_AllRightsReserved” in every code base out there? Do we need to go from header comments at the top of a file to reminder comments at the end of every line?
- softwaredoug 5y agoThis is kind of beside the point. Something can still be unethical and perfectly legal. The issue is that machine learning can whitewash a developers intended license. Or put differently, as a GitHub customer, are you comfortable with your code being used this way? Instead of a passive host, your code is now being used to create tremendous value for GitHub and Microsoft. Do you feel your trust has been violated? (regardless of legality).
- kalium-xyz 5y ago“ If it looks like a duck, swims like a duck, and quacks like a duck, then it probably is a duck.” I don’t see my license respected for code it regurgitates that I wrote, there is nothing more to this.
- varispeed 5y agoMass processing, repackaging and then selling the data is an exploitative business these multi-billion companies run without paying anything to the people who produced the data. This is wrong and should be stamped out.
- swiley 5y agoSo copyright is dead then? Can we merge all the leaked driver source into Linux and have decent OSes on handhelds yet? If I train an "ML autocomplete" on the "OpenNT" source can I share it legally?
- Causality1 5y agoThe output of a machine simply does not qualify for copyright protection – it is in the public domain. Is it just me or is that a patently ridiculous statement? The output of a machine belongs to the person owning/using the machine. If I use a digital camera to take a picture of a copyrighted image I'm still committing copyright infringement despite the output being created by a machine and a bunch of image processing software.
- dominicjj 5y ago"(of course, free software licenses would still fulfil the important function of contractually requiring the publication of modified source code)" No no no. Licenses are NOT contracts. Someone who copies or makes derivative works of copylefted software which they then distribute is obliged to remain within the bounds of the license not because they voluntarily promised, but because they don't have any right to act at all except as the license permits. https://www.gnu.org/philosophy/enforcing-gpl.en.html https://www.gnu.org/philosophy/enforcing-gpl.en.html
- fredgrott 5y agoand what pre-tell makes it a NON contract? License all by themselves are forms of contracts In fact the bill of rights is one
- roywiggins 5y agoYou need to actively assent to a contract. Some software has contracts ("EULAs") but you are bound by the license whether you agree or not. https://en.wikipedia.org/wiki/Meeting_of_the_minds?wprov=sfla1 https://en.wikipedia.org/wiki/Meeting_of_the_minds?wprov=sfl...
- adrusi 5y agoA license isn't a contract that binds the licensee, it's a contract that only binds the rightsholder. Since you, the licensee, are not relinquishing any rights in the contract, there's no need for you to agree to anything. The only rights being relinquished are the rightsholder's right to pursue legal retribution for some uses of their work that would otherwise be violations of copyright. You dont have to call it a contract, but it is a legal document in which one or more parties legally bind themselves, which seems like an adequate definition of a contract to me, and has more etymological fidelity to the word "contract" than other possible definitions that would exclude licenses. A contract is a legal instrument by which the breadth of your rights contract — as in become smaller.
- robbedpeter 5y ago
- Syzygies 5y agoWhatever the law, when does learning from what we read devolve into plagiarism? The poster child for this category would be those programs that generate nonsense English text that recognizably resembles a known author. They choose the next character at random, conditionally based on the previous characters. Too short a context, and the results are gibberish. Too long a context, and the results are plagiarism.
- Sr_developer 5y agoThis is a supposedly progressive politician, young, in an advanced country, her personal platform runs almost entirely on copyright issues and yet she gets almost everything wrong, what can you expect from your usual dinosaurs?
- marcosdumay 5y agoDepends on who is founding the dinosaurs. She had to work really hard to get to that conclusion she stated.
- maweki 5y agoThe output of Copilot may be not a derivative work, but the trained model surely is, right?
- boleary-gl 5y agoI’d agree with this conclusion if it wasn’t clear that it is very possible - if not common - for Copilot to just completely copy code. That isn’t fair use - that’s a clear violation of copyright regardless of license.
- dr_kiszonka 5y agoMy less lofty personal gripe with Copilot is as follows. I worked hard to produce quality code. GitHub will make money off my code. Copilot users will make money using my code. I - the creator - will make nothing. At the very least, I should have been asked whether my code can used by Copilot and I should get at least a share of the profit Copilot generates every month, where the share equals to my code / all training code used by Copilot. The latter part could be gamed by other developers in the future, but it's the best I could come up with.
- FeepingCreature 5y agoIf you didn't want your code to be reused or even commercialized by others, you really shouldn't have made it opensource.
- swiftcoder 5y agoIf my code isn't released under a permissive license, then I might have the expectation that those wishing to use my code for commercial purposes will contact me and pay for a commercial license. This is sort of the whole point of non-commercial licensing (and often, of the GPL itself, since many potential licensors don't wish to deal with GPL restrictions).
- IshKebab 5y agoSure but did you have the expectation that people wouldn't read your code and learn from it? I think even non-commercial licensing can't prevent that. If your code is so super-special that you don't want people to read it and go "ah that's a neat linked list reversal algorithm" or whatever then your only options are software patents or keeping it entirely closed source. Maybe trade secrets, but they tend to apply in very very limited circumstances. I doubt any software would qualify.
- drran 5y agoYes, people can read my open-sourced code and learn from it, like they can do with paints, movies, sculpts, and books. No, I don't allow to copy my code freely. Can you explain, what point you are trying to defend here?
- yakubin 5y ago> The output of a machine simply does not qualify for copyright protection – it is in the public domain. Does it mean that compiler output does not qualify for copyright protection and I may legally share copies of MS Word via torrent?
- betwixthewires 5y ago> ...some commentators accuse GitHub of copyright infringement, because Copilot itself is not released under a copyleft licence... This is not why. The issue at hand as I understand it is that people using copilot will potentially have code snippets in their work that are already licensed they do not know the license for and that they will not license properly as a result. That's in the first paragraph. If you enter this discussion with an incorrect presumption from the outset I don't see how you can form a valid defense. > However, by doing so, the copyleft scene is essentially demanding an extension of copyright to actions that have for good reason not been covered by copyright. No. Nobody is asking for an extension of copyright protection, we are asking for the existing reach of copyright to be respected. We built our licenses based on a ruleset that we were told is fair. You don't get to violate rules you made and then claim that copyleft people only made their licenses because as a workaround to copyright and so are being hypocrites. > Others focus on Copilot’s ability to generate outputs based on the training data. One may find both ethically reprehensible, but copyright is not violated in the process. The arguments I've heard are not that Microsoft is using publicly available information to train it's AI. The argument is that people are potentially (and in some current cases demonstrably) getting copy pasted code snippets from licensed software. If you can't see the plainly obvious problem here it's because you're trying not to. Also a point made in the article, that machine generated things cannot be copyright because copyright requires a creator, brings up an interesting question as to whether works by people who used copilot can be licensed at all.
- deleted 5y ago[deleted]
- kube-system 5y ago> If it were not possible to prohibit the use and modification of software code by means of copyright, then there would be no need for licences that prevent developers from making use of those prohibition rights (of course, free software licenses would still fulfil the important function of contractually requiring the publication of modified source code). The parenthetical backpedaling here is the entire point of copyleft. If it wasn't, copyleft wouldn't exist -- people would just release their software as public domain. The opposite of "copyleft" isn't "copyright". The opposite of "copyleft" is "never published", in which case, copyright is irrelevant. There is plenty of commercial closed-source software based on software released under permissive licenses like BSD, MIT or Apache, because they are not copyleft.
- sombremesa 5y agoI’d argue that copyright is still relevant when the source code isn’t published. It’s not too difficult to copy an algorithm from a binary even if you don’t have the source.
- kube-system 5y agoFair. When I wrote that, I was thinking "not published" as in server software.
- marcodiego 5y agoSimple way to fix this mess: allow to user to choose training data samples licenses.
- stabbles 5y agoAnd then what license do you choose? Many licenses require you to copy the original license verbatim, which may include the author's name and the date.
- scotty79 5y agoDon't you think that our world would be way more relaxed and flourishing place if lawers kept their noses out of software like they are keeping them out of math?
- CyberRabbi 5y ago> Works licensed under copyleft may be copied, modified and distributed by all, as long as any copies or derivative works may in turn be re-used under the same license conditions. This creates a virtuous circle, thanks to which more and more innovations are open to the general public. She claims that Copilot advanced the goals of copyleft but copilot does not create a “virtuous cycle” of generating more public IP. The customers of Copilot use Copilot extract public work through Copilot for themselves and are not compelled to contribute back. Copilot is anti-FOSS plain and simple.
- ksec 5y agoSo may be it is best to have a separate license for Machine Learning? Let's call it copilot licences. ( May be it is better to call it an exemption ? ) You will need AGPL / GPL / LGPL / MIT / Apache / BSD + Copilot licences before it can be used for training? Knowing there are a very small possibility that some code snippet will be the output? I mean we could endless debate this with no solution unless this is put into court.
- vharuck 5y ago>What would then stop a music label from training an AI with its music catalogue to automatically generate every tune imaginable and prohibit its use by third parties? What would stop publishers from generating millions of sentences and privatising language in the process? The existing barrier we have is that, unless the music label can prove a human artist has listened to the specific song matching the artist's, there's no copyright violation. A copyright protects creators from having their work copied. It doesn't give them ownership over matching works. I'm sure there are plenty of pairs of novels with the same first sentence despite each author never having read the other's work.
- rjmunro 5y agoNote that Patents and Trademarks are not like this. You can innocently recreate an invention or a similar trademark and you are still infringing. This often causes confusion - people apply the rules of one type of IP to the others, but they have almost nothing in common.
- mrh0057 5y agoWhy is everyone ignoring the fact what neural networks do? It is being used as a search context aware pattern matching and use that to predict what you will write next. Of course it's going to return copyrighted works based on what you right. It's a pattern matching algorithm what exactly did they think it was going to do?
- ShamelessC 5y agoAs best I can tell, people on Hacker News largely think of machine learning as some sort of statistical trick that they don't actually need to apply any further understanding towards. They see it repeat Doom code verbatim and assume it is capable of repeating any and all code it's ever seen verbatim - hence "laundering". What they maybe aren't considering is that specific snippet is famous. It has likely been pasted thousands of times with and without attribution on public GitHub repositories. Yes, it has seen code before. No, it didn't memorize the entirety of the dataset it was trained on. If it did - it has explicitly overfit, won't generalize to downstream tasks and ultimately failed at being useful in the general case. Unfortunately, "we don't know" still, but what may have happened is that their transformer architecture creates a more efficient representation of the byte pair encoding representing the code. In doing so, it is able to learn about context, structure, and logic of the language it is trained on. Anyways, I think this whole thing is absurd. So far - every "atrocity" I have seen committed by copilot is easily achievable with GitHub advanced search using "code contains text".
- glitchc 5y agoI disagree with this article. GitHub Copilot is indeed infringing copyright and not only in a grey zone, but in a very clear black and white fashion that our corporate taskmasters (Microsoft included) have defended as infringement. The legal debate around copyright infringement has always centered around the rights granted by the owner vs the rights appropriated by the user, with the owner's wants superseding user needs/wants. Any open-source code available on Github is controlled by the copyright notice of the owner granting specific rights to users. Copilot is a commercial product, therefore, Github can only use code that the owners make available for commercial use. Every other instance of code used is a case of copyright infringement, a clear case by Microsoft's own definition of copyright infringement [1][2]. Github (and by extension Microsoft) is gambling on the fact that their license agreement granting them a license to the code in exchange for access to the platform supersedes the individual copyright notices attached to each repo. This is a fine line to walk and will likely not survive in a court of law. They are betting on deep lawyer pockets to see them through this, but are more likely than not to lose this battle. I suspect we will see how this plays out in the coming months. [1] https://www.microsoft.com/info/Cloud.html https://www.microsoft.com/info/Cloud.html [2] https://github.com/contact/dmca https://github.com/contact/dmca
- lubujackson 5y ago"Copilot is a commercial product, therefore, Github can only use code that the owners make available for commercial use." IANAL, but this doesn't sound quite right. There is a difference between "using" code (running it in a commercial product) and manipulating it as arbitrary data within a commercial product. It definitely can be a gray area, but let's say I use Amazon's service where I email a PDF to my Kindle - is it Amazon's responsibility to know the copyright status of the PDF, or mine? In both cases a commercial product is manipulating copywritten data for the benefit of a user.
- emrah 5y agoEven if it's legal for Copilot to do what it does, does it not violate GPL to take pieces of GPL'ed code and use them in a commercial product?
- malwrar 5y agoWho cares if they're infringing copyright? Microsoft bought the place that has a lot of our code and now is going to try and sell us a tool that will regurgitate it back on demand. The entire software industry is already largely based and advanced by the unpaid labor of open-source software project developers, GitHub as a popular open source ally could at least pretend to honor the gentleman's agreement of at least agreeing to respect the open-source origins of a ton of its stack. If the tool was also open we probably wouldn't have nearly as big a problem, but I guess Microsoft has to recoup the cost of their completely unnecessary purchase.
- deleted 5y ago[deleted]
- oolonthegreat 5y agoSuch a weird argument: "Copyleft people should not argue for better copyright". What does that even mean?
- pessimizer 5y agoShe's arguing that copyleft people are arguing for an effective extension of copyright into places IP lobbyists are currently fighting for. It's not a good framing. She's saying that we shouldn't argue for copyright to be consistent if we're against copyright - arguing that we should make a moral argument against a legal situation. It's as if we couldn't argue against drug companies being allowed to sell heroin if we were anti-drug war and drugs would remain illegal. It's a strategy argument that leads nowhere. If the result of making machine written works also subject to copyright results in all possible songs being copyrighted by a machine, that's a good outcome. It's obviously absurd and weakens the entire concept. We should demand consistency. If this is fine, we might as well stop enforcing the GPL, too. It's a trick of copyright to further the cause of anti-copyright. I'm sure somebody can write an "auto-fork" that will digest GPL'd code and rearrange and rephrase it in order to spit out a clone.
- moralestapia 5y ago>it suggests that even reproducing the smallest excerpts of protected works constitutes copyright infringement Actually, it is. It has to do with whether the small excerpt is copying what could be called "the heart of the work"; which in the case of code I would argue is almost always what you are after. No one's gonna copy the indentation style, boilerplate around functions/blocks, punctuation, etc. You always go for the "functional" part of the code, which is definitely "the heart of the work". The heart of Carmack's fast inverse square root lies in its selection of a particular set of constants and operations that happen (i.e. were designed) to approximate the square root without taking an expensive path. Copyright law would look at this novelty; I don't think it would argue around "the use of subtraction and multiplication in a computer program", as that would be plainly stupid. I am surprised that someone who is supposedly an expert in copyright law does not (or pretends not to know) about this, not only that, but to actually suggest the opposite. This is copyright 101, come on.
- emrah 5y agoCopilot itself may not be infringing copyright or GPL, but its users will be if they incorporate its suggestions into their commercial products.
- sprafa 5y agoAmazing how this was never an issue when other “AI” systems use other people’s data to learn how to drive cars/write text. But man you start messing with developer data and suddenly there are ethical issues! Amazing turnaround. Face it - AI as we currently call it is just a very sophisticated data sorting algo in most cases (let’s ignore the AlphaZero non supervised learning type). Everyone was getting celebrating when Common Man was destroyed by devs commoditising their knowledge through data capture. But now suddenly it’s a problem! Mess with a man's pocket.
- dj_mc_merlin 5y agoI think a good deal of engineers here should familiarize themselves with Julia Reda and her work and ask themselves if they have the legal knowledge to debate on this matter. Common knowledge is not acceptable to determine truth. Would you really respect the opinion of some dude who's only used Excel about your profession?
- josefx 5y agoShe cites githubs smallest excerpt claim for her reasoning when we already know that the tool happily reproduces entire functions with comments verbatim. Also her claims about machine generated code have a really funny interaction with the cp command. Clearly cp MicrosoftWindows11Source.zip FreeWindows.zip is not a creative process, cp is a command executed by a machine hence the contents of FreeWindows.zip are now entirely public domain. Man were was she when people where sued over creating entire libraries of public domain movies using BitTorrent?
- ramraj07 5y agoJust like you accuse her of being out of date with recent findings, y'all seem conveniently out of date with githubs assurance that they will be adding checks to not regurgitate full chunks of code. So what exactly is your point then?
- josefx 5y agoTheir own paper notes that they would at most inform you of the source. If they can even detect it, as their detection isn't perfect either[1]. [1]https://docs.github.com/en/github/copilot/research-recitation https://docs.github.com/en/github/copilot/research-recitatio...
- ramraj07 5y agoWon't work. This tool is attacking them like the presence of a vegan attacks some hardcore meat eaters. They might realize deep down that this is not an argument they can win but it offends their core existence in some ways so they can't help but die defending their incoherent arguments. Ethical or not, it's clear Microsoft isn't going to get into real legal trouble due to this, and if the tool is genuinely useful, it's going to "allow the laundering of GPL code" into companies, whatever that means. If that offends people then they better learn the lesson and not produce open source any more. I'm not happy but if thats the direction the natural progression of things take whatever let's see where that goes.
- hu3 5y agoPlease copy my code. Reality is I'll be gone in 100 years tops and I'd be more than glad if my crappy code actually helps someone. As for attribution, we all learn by looking at code from all kinds of licenses. Between Stack Overflow, projects hosted in GitHub, libraries that sit on our vendor directories and even closed source projects there's a lot that is carried over to new projects without attribution. We're heading to a world were most projects are basically libraries glued together anyway. Standing on the shoulders of giants and all that. The dream of an omniscient pair programming buddy is slowly coming to fruition and I for one welcome. Copilot is just a tool, fancy search engine for the code that's available online. Projects should be judged by the way they use Copilot just like I'm judged if I misuse my car. I couldn't care less whether my name is shoved in some ever increasing CONTRIBUTOR.md file that no one but machines will read. I'm actually going to start documenting blocks of code more thoroughly so Copilot can better infer what each block does.
- COMMENT___ 5y agoTLDR; GitHub will eventually add some kind of "data usage reporting" utility that could show which parts of final your code made with help of this CuckPilot could potentially infringe copyright with links to other known sources of these parts of code. Then they will tell you that it is your responsibility to ensure that your final code does not have copyright issues.
- alkonaut 5y agoWhether Copilot infringes copyright is a muddy area. I personally would like to think that the world where machines can be trained on any data is easier to live in than one where trained machines are tainted by the licens of input. The interesting question however isn't whether Copilot infringes copyrights, but whether those that use copilot do.
- Rapzid 5y agoOne of the points being made is that in the worst-case scenario of getting Copilot to repeat back verbatim chunks of code from projects, something that's not its primary use case, it would be a situation similar to a copy machine. You can copy a page out of a book, or the whole book, and be covered under fair use. But you can't sell your copy on Amazon. And if you did, the copy machine nor Xerox ran afoul of copyright law. You could also use a copy machine to copy fragments of the Linux kernel source out of a book about the Linux source and use them to construct an entirely original work that's not considered derivative. The devil's in the details, but GitHub talks at some length about the plagiarization issue and their plans to detect and link back to where verbatim chunks exist in the training data to let the operator decide what to do soo.. IDK.
- turtletontine 5y agoThe idea that the debate actually does a disservice to copyleft by relying on the strictest interpretations of copyright is an interesting perspective to me, but the rest of this seems pretty weak. (Caveat that I'm no lawyer.) Copilot can regurgitate verbatim chunks of other codebases: it seems absurd to me that that wouldn't count as derivative work.
- reilly3000 5y agoHow does one address the fact that 95% of software is based on the same basic tropes? At a certain level of density, all code trying to achieve a similar function to legally-protected code will convene on an implementation that is almost indistinguishable. With LOC accreting exponentially, only time will determine when we reach that threshold. The Copilots of the world serve to accelerate and monetize this reality.
- deleted 5y ago[deleted]
- Cort3z 5y agoI wonder how long it will take for the licenses to start explicitly disallowing this sort of usage. It is clearly something that many open source writers dislike, and in my opinion, rightly so.
- chx 5y ago> The short code snippets that Copilot reproduces from training data are unlikely to reach the threshold of originality. I can only repeat myself: In light of Google v. Oracle going as far as the Supreme Court I find your confidence in this quite astonishing.
- visarga 5y agoCopilot is the moment when simple functions have been commoditized, you can have as many as you like almost for free, and adapted to any project. Just spend a moment to admire the transition, it's a new stage of post-scarcity. AI can recreate photos, paintings, sounds, voice, music, human faces, text, dialogue, math, proteins, and now code. It does all this while allowing humans to control and direct the whole process, and create original combinations. They all have no economic value to own and are free to use now, like words in a language. Enjoy! Remember Karpathy's Char-RNN? How long we've come. http://karpathy.github.io/2015/05/21/rnn-effectiveness/ http://karpathy.github.io/2015/05/21/rnn-effectiveness/
- noobermin 5y agoI've said this before, but I hope the issue isn't infringement per se, but that the produced code isn't automatically GPL'ed. The author argues that machine generated code isn't copyrighted and this is good because it essentially fits the "data wants to be free" mentality, but I'd say tell that to the people who use it. Will they, after using something derived from open source, have to open source their code? No, they won't. If anything, this finally provides closed source developers with what they've always wanted, a means to rip open source code without having to return contributions. Julia Reda hints at that last bit as being an issue but only in a parenthetical. To the author, that literally is the whole point. Do people not remember the Free Software vs. Open Source debate? Or GPL vs BSD? The requirement that derived works also be free is literally the important bit in Free Software. This only fits the mentality of "data wanting to be free" if your model of that idea includes the permissive sensibility and doesn't care about actually changing the state of things, which is making free software more widely used in the world over proprietary software.
- deleted 5y ago[deleted]
- orthoxerox 5y agoWhether Copilot itself violates GPL or not is one issue. Whether the code produced by Copilot violates GPL or not is a whole different independent issue. If I am walking down the street, find a piece of paper with code on it, pick it up and add the code to my program and this code turns out to be licensed under the GPL then my program becomes a derivative work. It doesn't matter who wrote it on that piece of paper, whether it's a 100% correct copy of the GPLed code or not or if there are mistakes in it.
- alfiedotwtf 5y agoHas anyone tried dumping the debugging symbols from a Microsoft binary e.g explorer.exe and tried to autocomplete^Wcopilot its functions? Would be interesting how far Microsoft could be pushed before they ate their own hat.
- deleted 5y ago[deleted]
- dragonwriter 5y ago> Copyleft does not benefit from tighter copyright laws Of course it does, at least the goal copyleft serves for RMS style Free Software ideologues. While copyleft may be motivated by an ideology that prefers no copyright protections, at least for software, it relies on copyright maximalism to avoid nonfree derivatives. From advocates viewpoint, the worst situation is a copyright regime that is strong enough that it allows nonfree software to exist but is also weak enough that it prevents creating an iron wall that prevents the use of software built by ideolgoical opponents of nonfree software from being used to advance nonfree software.
- ZoomZoomZoom 5y ago> On the other hand, the argument that the outputs of GitHub Copilot are derivative works of the training data is based on the assumption that a machine can produce works. This assumption is wrong and counterproductive. Wow, what a bunch of, ahem, logic of questionable quality. "On the other hand, the argument that the outputs of Copilot are works is based on the assumption that a machine can produce. This assumption is wrong and counterproductive. It just moves electrons and hums a bit." This is a reductio ad absurdum. The argument is bogus, because what matters is the result of a person/other legal entity using the machine and its software.
- robbrown451 5y agoThis is not the black and white issue that the article implies it is, with statements such as "Machine-generated code is not a derivative work". Imagine a web scraping robot that just grabbed textual news articles and spit them out verbatim to searchers (without giving credit or linking to the original). That is obviously copyright infringement, even though it is done by a robot. Now imagine it does slight modifications to the text, using a thesaurus and maybe a bit of AI. It might substitute "is able to" for "can", or "frequently" for "often", but otherwise everything is left as is. Is that "machine generated"? Same goes for a hypothetical bot that scrapes existing music, and after listening to "He's So Fine", comes up with the melody for "My Sweet Lord." (as per the famous George Harrison copyright case from the 70s) It isn't off the hook simply because a machine was involved. If it truly "learns" what makes a good melody, and uses that to generate a very different melody (that might be equally similar to a dozen different songs), that's different. There is a full spectrum between simple bots that copy verbatim, and something that "deeply learns" and then writes a new article, or generates new source code, or writes a new melody, or whatever. I don't have a strong opinion on GitHub Copilot since I haven't really studied what it does and therefore I don't know where it lies on that spectrum, but this article is not useful if the author doesn't really explore the nuance, and treats everything as absolutes. (and I should say, I am very much of the opinion that copyright law as it is, is hopelessly broken, and I am always glad to see things like Copilot just so we can see it demonstrated why. But that is veering off topic...)
- wheelie_boy 5y agoYes, similarly I could definitely create a simple ML algorithm and feed it a single codebase to learn. Then it would be possible to predictively output code, such that it reproduces that entire codebase verbatim. There's nothing magical about a complex copying method that makes that less of a copyright infringement. There may be some threshold where it becomes fair use, but I agree with you that it's not cut and dry at all. The argument that someone could create all possible works of art or music or whatever and copyright them all is a ridiculous idea, from someone who doesn't understand exponentials.
- wruza 5y agoThese sorts of discussions always puzzle me. Copyright is not an objective thing, it’s a contract that supports the cash flow of a creator-distributor-consumer chain, given the former two assumed it will work to cover their expenses and return profits they expected. If your AI produces Beatles-like or even better music based on Beatles albums and/or some more, an AI-aware judge may (or may not, depending on the lobbying activity) decide that it is a direct derivative work, and in case it is automated, all copyright rules apply as in “copy” “right”. There is no need for technical objectivity to exist in between, because this law is not about technicalities. What seems like a loophole may be closed easily by a court decision based on much higher matters than “similarity” or “reconstruction”. If anyone can take your album at the release date and “reshuffle” a free version not worse than the original with few clicks, it is obvious damage to the copyright holder and it demotivates creating it. AI couldn’t do any of that back then, and they didn’t include right terms to cover that, but now it can, and they will just add that, unless someone (MS in this case) has better lawyers, who are ready to create a wide-enough precedent and drag it through all instances.
- aeturnum 5y agoI'm not a copyright expert but I wonder about an implication of two of this author's points: - Reading and remembering-about (like reading a book yourself) things does not infringe on copyright. - Copyright does not apply to the output of mechanistic code generation (as opposed to the human-written code that generates the code). So where does that leave the quake snippet (setting aside its own release as open source)? Assuming this technical description is correct, Copilot does not contain the code, just the correct weights to contextually reproduce it perfectly. Copyright does not apply to the chunk that Copilot produces, so does the code simply exist as Copilot created it without license? If that is correct, what are the limits? Could I train a ML algorithm to reproduce binaries from context and, if those produced binaries happen to be identical to other copyrighted products, then it's fine?
- monetus 5y agoI like this train of thought.
- deleted 5y ago[deleted]
- nhumrich 5y ago>> machine generated code is not derivative work Even if true, that doesn't indemnify copilot here. There is no way to _prove_ whether the code was generated by copilot vs yourself. Copilot is just autocomplete, so its still a human checking in the code. While it might not be illegal for copilot to generate those things, its illegal for a human to check it in and claim it as their own.
- flippinburgers 5y agoEmbrace, extend, extinguish.
- madrox 5y agoCopilot, to me, feels like a faster Stack Overflow. We already copy code snippets from all kinds of places across the web without thinking about how it's licensed. Sometimes, we copy whole functions and files. We're responsible for understanding what's going into our project. We don't blame NPM when it allows us to import a package into a project that subsequently violates the license. I'm absolutely sure this happens more than anyone cares to admit.
- nixpulvis 5y agoClaiming that generated work isn't work seems completely wrong. The hard to argue fact is, looking at the result, it doesn't really matter who wrote it, just how it reads, and what it does. What is lost in so much of the arguments about Copilot is that someone still needs to actually verify the code does the right thing. I have a feeling this tool does little but increase little bugs like off-by-one errors or all kinds of havoc; primarily because of false confidence in the autocomplete.
- zeptonix 5y agoReally having trouble getting the unrelenting hatred here on this site for something that's fundamentally new, clearly represents progress, and is obviously a trend that's here to stay. <CrazyIdea>Maybe the laws, licensing, etc. need to, you know, adapt, change, and evolve a little bit with time also -- just as everything else changes with time.</CrazyIdea>
- justshowpost 5y agoHuh? The article is full of waffle, but in the condensed form it shills for: (a) treating an adopted code as «trivial» as i++, which is pure demagogy because what we've seen already on that CoPilot video is NOT trivial (b) dismissal of (let me put it straight) piracy as somewhat special case of fair-use, which is valid only when code in question stays on video as prop, the real code ISN'T fair-use (c) accepting (a) and (b) above as ultimate truth just because bogey stricter copyright laws hurts FOSS. And the water is wet. This is absolutely meaningless filler because we know what stricter copyright laws hurts everyone since Napster days. So my overall impression from this reading is just... Huh? > My name is Julia, I'm the Pirate in the European Parliament. Didn't she split with Piratenpartei?
- codelord 5y agoTo understand if Copilot is infringing the licenses of codes used in its training data we have to get into the details of what it does and how it works. We can't make a general statement for any code generation software that was trained on open source code. It is a possible that at some point maybe even in not so distant future we will have ML models so good they can understand abstract concepts, learn and invent new algorithms and implementations by reading code. Such a ML model can be argued is learning similar to a human and hence it's not infringing any copyrights because it's not copying implementations it's learning concepts and ideas. But we are not there yet. When we get there we will know. Because at that point Siri would be able to have seamless conversations with you. At least half the jobs would disappear in favor of robots in a short time. The world would be a different place. Let's talk about what Copilot actually can do. It can copy snippets of code from Github while changing variable names. It can autocomplete trivial boilerplate code. If it's automagically generating a function for you that actually does something useful like sorting an array, you can be absolutely sure that it's just copy pasting it from an existing repo with some cosmetic changes.
- sizt 5y agoCopyright infringing or not, this genie will never go back into the bottle.
- tzahifadida 5y agoI believe it is true that in most cases it won't be infringing. Since even though it can sometimes output some trivial code verbatim to the original, that original work won't run or compile and therefor it is not even a work, just gibrish. Simply limiting the copilot scraping to software with at least a few files will probably resolve that issue. Moreover, if I where github I would simply change that statement regarding the legality and let the developer make the choice if to use or not. More often then not, it would probably make this academic talk that no one cares about. a few functions here and there is not work. Almost never people try to copy stuff like a micro kernel or something so small as to constitute a work. Personally I would rather treat this as a search engine. I don't copy paste code but I may be old fashion and probably the exception here.