9 ms·
It shouldn't do that, and we are taking steps to avoid reciting training data in the output: https://copilot.github.com/#faq-does-github-copilot-recite-code-fro
by natfriedman 5y ago
It shouldn't do that, and we are taking steps to avoid reciting training data in the output:
https://copilot.github.com/#faq-does-github-copilot-recite-code-from-the-training-set https://copilot.github.com/#faq-does-github-copilot-recite-c...
https://docs.github.com/en/early-access/github/copilot/research-recitation https://docs.github.com/en/early-access/github/copilot/resea...
In terms of the permissibility of training on public code, the jurisprudence here – broadly relied upon by the machine learning community – is that training ML models is fair use. We are certain this will be an area of discussion in the US and around the world and we're eager to participate.
- Hamuko 5y ago>training ML models is fair use How does that apply to countries where Fair Use is not a thing? As in, if you train a model on a fair use basis in the US and I start using the model somewhere else?
- KMnO4 5y agoI don’t think it’s fair to ask a US company to comment on legalities outside of the US.
- detaro 5y agoIt's fair to expect a international company pushing its products all over the world to be prepared to comment on non-US jurisdictions. (I have some sympathy for "we have a local market, and that's what we are solely targeting and preparing for" in companies where that is actually the case, but that's really not what we are dealing with in the case of Microsoft/GitHub)
- toomuchtodo 5y agoOne would expect GitHub (owned by Microsoft) to have engaged corporate counsel for an opinion (backed by statue and case law), and to be prepared to disable the functionality in jurisdictions where it’s incompatible with local IP law.
- deleted 5y ago[deleted]
- Asmod4n 5y agoFair use doesn’t exist in Germany.
- SCLeo 5y ago> ...the jurisprudence here – broadly relied upon by the machine learning community – is that training ML models is fair use. To be honest, I doubt that. Maybe I am special, but if I am releasing some code under GPL, I really don't want it to be used in training a closed source model, which will be used in a closed source software generating code for closed source projects.
- yjftsjthsd-h 5y agoIs it any different than training a human? What if a person learned programming by hacking on GPL public code and then went to build proprietary software?
- woodruffw 5y agoA human being who has learned from reading GPL'd code can make the informed, intelligent decision to not copy that code. My understanding of the open problem here is whether the ML model is intelligently recommending entire fragments that are explicitly licensed under the GPL. That would be a licensing violation, if a human did it.
- 10000truths 5y ago> A human being who has learned from reading GPL'd code can make the informed, intelligent decision to not copy that code. A model can do this as well. Getting the length of a substring match isn’t rocket science.
- akavel 5y agoActually, I believe it's tricky to say if even human can actually do that safely. There's the whole concept of "cleanroom rewrite" - meaning, if you want to rewrite some GPL or closed-source project into a different license, you should make sure you never ever seen even a glimpse of the original code. If you look on GPL or closed-source code (or, actually, code governed by any other license), it's hard to prove you didn't accidentally/subconsciously remember parts of this code, and copy them into your "rewrite" project even if "you made a decision to not copy". The border between "inspired by" and "blatant copyright infringement" is blurry and messy. If that was already so tricky and troublesome legal-wise before, my first instinct is that with the Copilot it could be even more legally murky territory. IANAL, yet I'd feel better if they made some [legally binding] promises that their model is based only on code carefully verified to have one of an explicit (and published) whitelist of permissive licenses. (Even this could be tricky, with MIT etc. actually requiring some mention in your advertising materials [which is often forgotten], but now that's a completely different level of trouble than not knowing if I'm infringing GPL or some closed-source code, or other weird license.)
- npteljes 5y ago> ...the jurisprudence here – broadly relied upon by the machine learning community – is that training ML models is fair use. If you train az ML model on GPL code, and then make it output some code, would that not make the result a derivative of the GPL licensed inputs? But I guess this could be similar to musical composition. If the output doesn't resemble any of the inputs, or contains significant continous portions of them, then it's not a derivative.
- IncRnd 5y ago> If the output doesn't resemble any of the inputs, or contains significant continous portions of them, then it's not a derivative. In this particular case, the output resembles the inputs, or there is no reason to use Github Copilot.
- sicromoft 5y agoYou just shared a URL that says "Please do not share this URL publicly".
- jamie_ca 5y agoWell, he's also GitHub's CEO so it's probably just fine.
- deleted 5y ago[deleted]
- eqtn 5y agoWould i be able to use something like this in the near future to produce a proprietary linux kernel?
- jazzyjackson 5y ago> It shouldn't do that, and we are taking steps to avoid reciting training data in the output This just gives me a flashback to copying homework in school, “make sure you change some of the words around so it’s not obvious” I’m sure you’re right Re: jurisprudence, but it never sat right with me that AI engineers get to produce these big, impressive models but the people who created the training data will never be compensated, let alone asked. So I posted my face on Flickr, how should I know I’m consenting to benefit someone’s killer robot facial recognition?
- ramraj07 5y agoWait I thought y'all argued Google didn't copy Java for Android, now that big tech is copying your code you're crying wolf?
- InspiredIdiot 5y agoThe whole point of that case begins with the admission "yes of course Google copied." They copied the API. The argument was that copying an API to enable interoperability was fair use. It went to the Supreme Court because no law explicitly said that was fair use and no previous case had settled the point definitively. And the reason Google could be confident they copied only the API is because they made sure the humans who did it understood both the difference and the importance of the difference between API and implementation. I don't think there is a credible argument that any AI existing today can make such a distinction.
- CyberRabbi 5y ago> training ML models is fair use In what context? You are planning on commercializing Copilot and in that case the calculus on whether or not using copyright protected material for your own benefit changes drastically.
- josourcing 5y agoIt isn't. US copyright law says brief excerpts of copyright material may, under certain circumstances, be quoted verbatim ----> for purposes such as criticism, news reporting, teaching, and research <----, without the need for permission from or payment to the copyright holder. Copilot is not criticizing, reporting, teaching, or researching anything. So claiming fair use is the result of total ignorance or disregard.