19 ms·
Howdy, folks. Ryan here from the GitHub Copilot product team. I don’t know how the original poster’s machine was set-up, but I’m gonna throw out a few theories
by _ryanjsalva 4y ago
Howdy, folks. Ryan here from the GitHub Copilot product team. I don’t know how the original poster’s machine was set-up, but I’m gonna throw out a few theories about what could be happening.
If similar code is open in your VS Code project, Copilot can draw context from those adjacent files. This can make it appear that the public model was trained on your private code, when in fact the context is drawn from local files. For example, this is how Copilot includes variable and method names relevant to your project in suggestions.
It’s also possible that your code – or very similar code – appears many times over in public repositories. While Copilot doesn’t suggest code from specific repositories, it does repeat patterns. The OpenAI codex model (from which Copilot is derived) works a lot like a translation tool. When you use Google to translate from English to Spanish, it’s not like the service has ever seen that particular sentence before. Instead, the translation service understands language patterns (i.e. syntax, semantics, common phrases). In the same way, Copilot translates from English to Python, Rust, JavaScript, etc. The model learns language patterns based on vast amounts of public data. Especially when a code fragment appears hundreds or thousands of times, the model can interpret it as a pattern. We’ve found this happens in <1% of suggestions. To ensure every suggestion is unique, Copilot offers a filter to block suggestions >150 characters that match public data. If you’re not already using the filter, I recommend turning it on by visiting the Copilot tab in user settings.
This is a new area of development, and we’re all learning. I’m personally spending a lot of time chatting with developers, copyright experts, and community stakeholders to understand the most responsible way to leverage LLMs. My biggest take-away: LLM maintainers (like GitHub) must transparently discuss the way models are built and implemented. There’s a lot of reverse-engineering happening in the community which leads to skepticism and the occasional misunderstanding. We’ll be working to improve on that front with more blog posts from our engineers and data scientists over the coming months.
- binarymax 4y agoLater in the thread he stated the code was not on the machine he tested copilot with. Copilot training data should have been sanitized better. In addition: any code that is produced by copilot that uses a source that is licensed, MUST follow the practices of that license, including copyright headers.
- danielheath 4y agoRight - but if someone pushes the same code to github and changes the licence file to say "public domain", what's the legally correct way to proceed? What's the morally correct way to proceed?
- Marazan 4y agoIf you want a massive corpus of training data theb you can create it by hand like grandpappy used to do rather than just thieving it whilst telling yourself it is fine.
- lmm 4y agoLegally, if you're publishing a derived work without legitimate permission then you're civilly liable for statutory + actual damages, the only thing you're avoiding is the treble damages for wilful infringement. Morally I'd say you should make a reasonable good faith effort to verify that you have a real license for everything you're using. When you're importing something on the scale of "all of Github" that means a bit more effort than just blindly trusting the file in the repository. When I worked with an F500 we would have a human explicitly review the license of each dependency; the review was pretty cursory, but it would've been enough to catch someone blatantly ripping off a popular repo.
- d1sxeyes 4y agoHow do you know GH didn't? Maybe they only included repos with LICENSE.MD files which followed a known permissive licence? What if a particular piece of code is licensed restrictively, and then (assuming without malice) accidentally included in a piece of software with a permissive license? What if a particular piece of code is licensed permissively (in a way that allows relicensing, for example), but then included in a software package with a more restrictive licence. How could you tell if the original code is licensed permissively or not? At what point do Github have to become absolute arbiters of the original authorship of the code in order to determine who is authorised to issue licenses for the code? How would they do so? How could you prove ownership to Github? What consequences could there be if you were unable to prove ownership? That's before we even get to more nuanced ethical questions like a human learning to code will inevitably learn from reading code, even if the code they read is not permissively licensed. Why then, would an AI learning to code not be allowed to do the same?
- soulofmischief 4y agoYour response is much appreciated. The dust will settle eventually, it seems distrust is the de facto way which people interpret new technologies.
- binarymax 4y agoIMO someone should be responding to the Twitter thread itself. I’m not sure why the official response is here when the copyright owner raised the complaint on Twitter.
- _ryanjsalva 4y agoGood call. Done. https://twitter.com/docsparse/status/1581461734665367554?s=46&t=yF5zaMKk62GD32GddpVd2Q https://twitter.com/docsparse/status/1581461734665367554?s=4...
- svnt 4y agoThis doesn’t at all address the primary issue, which is one of licensing. Is it a valid defense against copyright infringement to say “we don’t know where we got it, maybe someone else copied it from you first?” If someone violated the copyright of a song by sampling too much of it and released it in the public domain (or failed to claim it at all), and you take the entire sample from them, would that hold up in a legal setting? I doubt it.
- minhazm 4y agoIt does address it, although not that clearly. This happens all the time with news media. They will post a picture and say they got permission from X person, but X person actually didn't even own the copyright in the first place. That doesn't make any of it okay, but it does mean that the organization has legal cover in this case and the worst that will happen is that they'll have to take the content down. In GitHub's case if that same code snippet is found in other repo's that have different licensing then it's difficult to really prove who owns the copyright, it's a legal issue between the original copyright owner and the person that re-distributed the work. They can submit a DCMA takedown notice for the other repo's. But it's pretty unlikely Github gets into any legal trouble as long as they can prove that they got the snippet from someone else.
- leni536 4y agoIf that's true, than Github is just "washing its hands". Not at all reassuring for copyright holders and users of copilot.
- helsinkiandrew 4y agoThat code seems to appear in thousands of repositories on GitHub, I’m sure some of them haven’t copied the license. The vast majority of people who would use a matrix transform function they got from code completion (or from a GitHub or stack overflow search) probably don’t care what the license is. They’ll just paste in the code. To many developers publicly viewable code is in the public domain. Code pilot just shortens the search by a few seconds. Microsoft should try todo better (I’m not sure how), but the sad fact is that trying to enforce a license on a code fragment is like dropping dollar bills on the sidewalk with a note pinned to them saying “do not buy candy with this dollar”
- water-your-self 4y agoCan you clarify what you mean when you describe Github as an LLM maintainer? What LLM does github maintain?
- _ryanjsalva 4y agoProbably more accurate to say, "LLM service provider." Ultimately, GitHub distributes a derivative of OpenAI's Codex – though, the version in production has been tuned considerably.
- A4ET8a8uTh0 4y agoThank you for the response ( especially since it does not read like a corporate damage control response ). I will admit that I am conflicted, because I can see some really cool potential applications of Copilot, but I can't say I am not concerned if what Tim maintains is accurate for several different reasons. Lets say Copilot becomes the way of the future. Does it mean we will be able to trust the code more or less? We already have people, who copy paste stack overflow without trying to understand what the code does. This is a different level, where machine learning seems to suggest a snippet. If it works 70% of time, we will have a new generation of programmers management always wanted.
- _ryanjsalva 4y agoI firmly believe that GitHub Copilot isn't a replacement for thinking, breathing, reasoning developers on the other side of the keyboard. Nor is it a replacement for best practices that ensure proper code quality like linting, code reviews, testing, security audits, etc. All the research suggests that AI-assisted auto-complete merely helps developers go faster with more focus/flow. For example, there's an NYU study that compared security vulnerabilities produced by developers with and without AI-assistend auto-complete. The study found that developers produced the same number of potential vulnerabilities whether they used AI auto-complete or not. In other words, the judgement of the developer was the stronger indicator of code quality. The bottom line is that your expertise matters. Copilot just frees you up to focus on the more creative work rather than fussing over syntax, boilerplate, etc.
- trasz 4y agoThis is very much a standard damage control. Notice how the drone completely ignored actual problems and instead derailed the whole thread with fake ones and broken analogies.
- didibus 4y ago> The OpenAI codex model (from which Copilot is derived) works a lot like a translation tool. When you use Google to translate from English to Spanish, it’s not like the service has ever seen that particular sentence before. Instead, the translation service understands language patterns (i.e. syntax, semantics, common phrases). In the same way, Copilot translates from English to Python, Rust, JavaScript, etc. The model learns language patterns based on vast amounts of public data As I understand, this isn't proven is it? We don't know that the model isn't simply stitching and approximating back to the closest combination of all the data it saw, versus actually understanding the concepts and logic. Or is my understanding already behind times?
- mahogany 4y ago> This is a new area of development, and we’re all learning. I’m personally spending a lot of time chatting with developers, copyright experts, and community stakeholders to understand the most responsible way to leverage LLMs. Given that there have been major concerns about copyright infringements and license violations since the announcement of Copilot, wouldn't it have been better to do some more of this "learning", and determine what responsibilities may be expected of you by the broader community, before unleashing the product into the wild? For example, why not train it on opt-in repositories for a few years first, and iron out the kinks?
- throwaway1851 4y ago> why not train it on opt-in repositories for a few years first, and iron out the kinks? Ha ha. Because then the product couldn’t be built. Better to steal now and ask forgiveness later, or better yet, deny the theft ever occurred.
- codalan 4y agoIf Copilot was designed with any ethics in mind, it would have been an opt-in model. Instead, they scoured and plagiarized everyone's source code without their consent.
- ovi256 4y agoBecause the ethical opt-in model builders are still working on putting together their cleanly sourced dataset.
- concordDance 4y agoCopyright infringement is not theft in the most important sense that matters. Theft is normally negative sum, copyright infringement is almost always positive sum.
- tauwauwau 4y agoHad to find this after a long time IT Crowd Piracy Warning https://www.youtube.com/watch?v=ALZZx1xmAzg https://www.youtube.com/watch?v=ALZZx1xmAzg
- pabs3 4y agoI read that the Amazon equivalent of GitHub Copilot does respect licensing properly, maybe you can talk to them about adopting their approach.
- mike_d 4y ago> When you use Google to translate from English to Spanish, it’s not like the service has ever seen that particular sentence before. But that is exactly how it works. Translation companies license (or produce) huge corpuses of common sentences across multiple languages that are either used directly or fed into a model. Third party human translators are asked to assign rights to the translation company. https://support.google.com/translate/answer/2534530 https://support.google.com/translate/answer/2534530
- briffle 4y agoTheir are lots of sheets properly licensed that show the notes to play ‘stairway to heaven’. Many intro to guitar books, etc. If I publish myself playing that song without the copyright owners permission (and typically attribution) I am looking at some very, very negative outcomes. The fact that there are many copies correctly licensed (or not) does not obsolve me of anything. Curious how this any different?
- zimpenfish 4y ago> If I publish myself playing that song without the copyright owners permission Music licensing is bonkers but AFAIR (at least in the UK) I think you're allowed to do covers without explicit permission[1] - you'll have to give the original writers/composers the appropriate share of any money you make. [1] Which is why you (used to?) get, e.g., supermarkets playing covers of songs rather than the originals because it's cheaper.
- Timwi 4y agoWhat's the appropriate share?
- zimpenfish 4y agoIn the UK, at least, it seemed to depend on several decades of accumulated rules and whatnots that only the PRS understood[1] (but I haven't been involved in anything related to music licensing for a few years and even then it was baffling.) [1] Things like "was it on the radio or a TV show or a live performance or a recording? who was the composer? which licensing region was it in?" etc.
- nabakin 4y agoAre variable names randomized before being trained on? If so, that could prove something else is going on because all of the variable names Copilot outputted were the same.
- speedgoose 4y agoCopilot is good at naming variables, I don’t think you should randomise them.
- nabakin 4y agoI'm thinking Copilot may be good at naming variables and still use randomized variable naming in their training set
- carom 4y agoCopilot is the largest disrespect to open source software I have ever seen. It is a derivative work of open source code and it is not released under the same license. It is also capable of laundering open source code. Congratulations for working on the "extinguish" phase of embrace, extend, extinguish for open source.
- sngz 4y agoI really wonder what all those people who said Microsoft acquisition of Github was a good thing for opensource think now. I'm sure there will still be mental gymnastics involved.
- efdee 4y agoClaiming Github's Copilot is Microsoft's "Extinguish" step against open source _is_ the mental gymnastics.
- carom 4y agoIt really isn't mental gymnastic. Copilot, the model, is a program that is a derivative work of open source code. It should be open source.
- efdee 4y agoThat is another discussion. He is not claiming that it should be open source - he is claiming it is created to destroy open source.
- carom 4y agoHe is me. Allowing large companies to ignore licenses and giving them a tool to launder licensed code at scale is a significant threat to the integrity of open source licenses.
- deleted 4y ago[deleted]
- spookyuser 4y agoThe way that copilot seems to understand not just variables from the file you’re working with but functions classes and variables from all the files in your folder is incredible!
- tomcam 4y agoThat’s a whole lot of words that don’t address TFA at all.
- cycomanic 4y ago2 things: 1. you make it out like a translation from e.g. English to Spanish wouldn't fall under copyright. That's incorrect, in most juristictions I am aware of, it actually fall under the copyright of the original work and fall under its own copyright. 2. When will copilot be released open source, it is pretty clear by now that it is a derivative of all the OSS code so how about following the licensing?
- choppaface 4y agoYour long long paragraph about neighboring code editors is disproven: https://news.ycombinator.com/item?id=33227395 https://news.ycombinator.com/item?id=33227395 You’re really not going to solve this problem with marketing (“blog posts”) or some pro-Github story from data scientists. You need a DMCA / removal request feature akin to Google image search and you need work on understanding product problems from the customer perspective.
- deleted 4y ago[deleted]
- socialismisok 4y agoHey Ryan! Have you ever done any reading on the Luddites? They weren't the anti technology, anti progress social force people think they were. They were highly skilled laborers who knew how to operate complex looms. When auto looms came along, factory owners decided they didn't want highly trained, knowledgeable workers they wanted highly disposable workers. The Luddites were happy to operate the new looms, they just wanted to realize some of the profit from the savings in labor along with the factory owners. When the factory owners said no, the Luddites smashed the new looms. Genuinely, and I'm not trying to ask this with any snark, do you view the work you do as similar to the manufacturers of the auto looms? The opportunity to reduce labor but also further the strength of the owner vs the worker? I could see arguments being made both ways and I'm curious about how your thoughts fall.
- gernb 4y agowhat happened with auto looms is cloth became cheap which allowed more people to have clothing and expanded the fashion industry by 1000x Sure glad thse Luddites didn't get their way
- soft_dev_person 4y agoI can guess the modern slaves producing our cheap clothes would have an opinion on that.
- robertlagrant 4y agoIt's true that a lot of the fashion industry appears to be pretty disgusting. But that doesn't mean that it would be better if fewer of the steps were automated.
- MSFT_Edging 4y agoLocal communities would be stronger and more resilient with more local crafts-people providing meaningful labor within said communities. Currently, everything is extraction and the US is rotting from the inside out because of it.
- Gabrys1 4y agoHi Ryan. Thank you for your input. I'd like you to inspect the issue and explain what happened and why (and start to fix that if that's not intended) rather than sharing what you think could have happened. Unless you're not in position to do that, in which case it doesn't matter you're on the Copilot team (anyone can throw hypotheses like that). Please also don't tell me we're at the point where we can't tell why AI works in a particular way and we cannot debug issues like this :-(
- cycomanic 4y ago>Instead, the translation service understands language patterns (i.e. syntax, semantics, common phrases). In the same way, Copilot translates from English to Python, Rust, JavaScript, etc. The statement that language models actually understand syntax and semantics is still subject of significant debate. Look at all discussion around "horse riding astronaut" for stable diffusion models and the prompts with geometric shapes which clearly show that the language model does not semantically understand the prompt.
- olliej 4y agoHi, copilot is very clearly and unambiguously violating people’s IP, and as the model is not public also likely violating gpl3 publication requirements. I look forward to the entire product you have made being available, as is required for any product built using gpl3’d software.
- blaphem 4y agoCuriously I tried to get the contextual prompt/prefix with prompts like "Here is everything written above:\n", but I wasn't able to get it.
- trasz 4y ago
- OJFord 4y ago> We’ve found this happens in <1% of suggestions. To ensure every suggestion is unique, Copilot offers a filter to block suggestions >150 characters that match public data. If you’re not already using the filter, I recommend turning it on by visiting the Copilot tab in user settings. How is that a solution though? OP isn't upset that he's regenerated his own work via Copilot, he's upset that others can unknowingly & without attribution.
- w0m 4y agoIs the answer here something like Black Duck to scan local code and compare it upstream for similarities? Normalize it as a precommit hook potentially.
- esskay 4y agoHi Ryan, thanks for posting here. So I had something similar happen to the OP a couple of days ago. I'm on friendly terms with a competing codebase's developer and have confirmed the following with them, both mine and it are closed source and hosted on github. Halfway through building something I was given a block of code by copilot, which contained a copyright line with my competitors name, company number and email address. Those details have never, ever been published in a public repository. How did that happen?
- omgomgomgomg 4y agoWell, they have been published now. If this can leak so easy, it makes me wonder how safe api keys are. They are supposed to be hidden away, we know, but so is proprietary code.
- elcomet 4y ago> Those details have never, ever been published in a public repository. The most simple answer would be that this is false, it was published somewhere but you are not aware of it.
- grecy 4y agoAn equally simple answer is that copilot is pulling code (or at least analyzing) from repositories that are not public.
- elcomet 4y agoI think that's very unlikely, they said and repeated that they are not using private code. People catching them lying on this would be very bad for GitHub.
- grecy 4y agoBugs and unexpected behaviour catch us all. I’m not saying they’re intentionally lying, but that one possible explanation is it looking through non public repositories
- robin_reala 4y agoHi Ryan. Seeing as we’re going to see more and more of this sort of complaint come up because inherently the licensing situation with open source code is overly complicated, have you considered switching to use code where the licensing is clear: Github and Microsoft’s private codebases?
- robertlagrant 4y ago> If similar code is open in your VS Code project, Copilot can draw context from those adjacent files I'm concerned that "draw context from" is a euphemism. Does it mean it uses code that's only on your laptop to train its AI?
- kaetemi 4y agoIt's personalised to you. In general, it mimics the project's coding style. (And if the project code is terrible, good luck.)
- dustedcodes 4y agoDoes Copilot keep attribution of public code if it reuses it in its suggestions? For example, if I copy pasted code from someone in my open source project, and the copied code was subjected to required attribution will Copilot keep that attribution when it copies my code again?
- jraph 4y ago> My biggest take-away: LLM maintainers (like GitHub) must transparently discuss the way models are built and implemented The best way to be transparent about a software implementation is to open source the thing. If that's your take away, this is the only logical thing to do. Blogs posts would be appreciated but are not enough. We can only trust what you say, we cannot verify anything.
- klqpl 4y agoThere is a very simple solution: Use Microsoft proprietary code for training the model. Keep your hands off open source code. There you have the "most responsible way". The GPL should be updated to prohibit code to be used for "learning" (i.e., regurgitating copyrighted fragments).
- KennyBlanken 4y ago> This is a new area of development, and we’re all learning. Being a new area of development doesn't release you from your obligation to make sure what you're doing is ethical and legal FIRST. > I’m personally spending a lot of time chatting with developers, copyright experts, and community stakeholders to understand the most responsible way to leverage LLMs. And yet oddly nowhere did the phrase "I reached out to OP to discuss with with them" appear anywhere in your response." Nope. Being part of GitHub's infamous Social Media "incident response" team was more important than actually figuring out what was going on. You don't even say that you will look into the situation with OP, or speak to them. waves to all the github employees who will be reading this comment because someone on Github's marketing team links to it
- zelphirkalt 4y agoThis response expresses some of the things, that are to criticize about MS' Copilot project. But also I don't like the instant attempt to subtle discredit the report by dropping something like "I don’t know how the original poster’s machine was set-up" in the first (or second, if you want to be technical) phrase. First consider that you made a mistake yourself, _then_ ask, whether the fault could be on the other side. I really dislike this high-horse down-talking tone. Maybe it was not meant to sound like that, maybe this kind of talk has become a habit without noticing. Lets assume that, giving a benefit of a doubt. Onto the actual matter: > If similar code is open in your VS Code project, Copilot can draw context from those adjacent files. This can make it appear that the public model was trained on your private code, when in fact the context is drawn from local files. For example, this is how Copilot includes variable and method names relevant to your project in suggestions. How comes, that Copilot hasn't indicated, where the code came from? How can it ever seem, like the code came from elsewhere? That is the actual question. We still need Copilot to point us to repositories or snippets on Github, when it suggests copies of code (including just renaming variables). Otherwise the human is taken out of the loop and no one is checking copyright infringements and license violations. This has been requested for a long time. Time for Copilot to actually respect rights of developers and users of software. > It’s also possible that your code – or very similar code – appears many times over in public repositories. So basically it propagates license violations. Great. Like I said, the human needs to be kept in the loop and Copilot needs to empower the user to check where the code came from. > This is a new area of development, and we’re all learning. The problem is not, that this is a new development or that we are all learning. That is fine. Sure, we all need to learn. However, when there is clearly a problem with how Copilot works, it is the responsibility of the Copilot development team to halt any further violations and first fix that problem, before letting the train roll on and violating more people's rights. The way this is being handled, by just shrugging and rolling on, maybe at some point fixing things, is simply not acceptable.
- CapsAdmin 4y ago> How comes, that Copilot hasn't indicated, where the code came from? I can't say for sure about copilot but in general you don't have that kind of information. The problem is a bit like trying to add debug symbols back to some highly optimized binary program.
- Timwi 4y ago> I’m personally spending a lot of time chatting with developers, copyright experts, and community stakeholders to understand the most responsible way to leverage LLMs. This claim rings extremely hollow when your team refuses to do any of the obvious things that developers, experts and community stakeholders in this very thread (and the rest of this website) are telling you. You still haven't open-sourced Copilot. You still haven't trained it on Microsoft internal code such as Windows and Office. You still haven't made the model freely available for anyone to run locally. Until you do any of these things, you are not acting in the interest of the community and you are just exploiting people and their code for your own profit.
- pretendscholar 4y ago> This is a new area of development, and we’re all learning. Is this the tact your organization would take if someone else’s code completion software was generating proprietary Microsoft’s proprietary code?
- dogleash 4y agoThe topic of code generation that's near-identical or overwhelmingly-similar to existing code has come up a number of times. While the problem is obvious, it's a bit opaque and poorly communicated. How has your team defined, specified and clearly articulated these issues with generation? How do you test your generation to distinguish between fixing a problem vs reducing obvious true positives (i.e. unintentionally making the problem less visible without eliminating it)? Without some communication on those fronts (which maybe I've just not seen yet), I'm not surprised that you get pushback against your product from people feel like you're taking a cover-our-ass-and-YOLO approach.
- an1sotropy 4y agoRyan - one word you avoided using in your reply is "license". The fact that Copilot is reproducing code without attribution is not the legal show-stopper here has much as the fact Copilot is reproducing code without documenting what license it falls under. If your hope is that saying "it came out of our ML model" somehow removes Copilot from the well-established legal framework of licensing, I think you're wrong, and you are creating a minefield that I and others choose to stay well clear of. The revenue from Copilot, and the rest of MS, can probably pay your legal bills, but certainly not mine.