22 ms·
Lots of questions: - the generated code by AI belongs to me or GitHub? - under what license the generated code falls under? - if generated code becomes t
by stephen82 5y ago
Lots of questions:
- the generated code by AI belongs to me or GitHub?
- under what license the generated code falls under?
- if generated code becomes the reason for infringment, who gets the blame or legal action?
- how can anyone prove the code was actually generated by Copilot and not the project owner?
- if a project member does not agree with the usage of Copilot, what should we do as a team?
- can Copilot copy code from other projects and use that excerpt code?
- if yes, *WHY* ?!
- who is going to deal with legalese for something he or she was not responsible in the first place?
- what about conflicts of interest?
- can GitHub guarantee that Copilot won't use proprietary code excerpts in FOSS-ed projects that could lead to new "Google vs Oracle" API cases?
- natfriedman 5y agoYou should read the FAQ at the bottom of the page; I think it answers all of your questions: https://copilot.github.com/#faqs https://copilot.github.com/#faqs
- rozab 5y agoThis page has a looping back button hijack for me
- samtheprogram 5y agoThe most important question, whether you own the code, is sort of maybe vaguely answered under “How will GitHub Copilot get better over time?” > You can use the code anywhere, but you do so at your own risk. Something more explicit than this would be nice. Is there a specific license? EDIT: also, there’s multiple sections to a FAQ, notice the drop down... under “Do I need to credit GitHub Copilot for helping me write code?”, the answer is also no. Until a specific license (or explicit lack there-of) is provided, I can’t use this except to mess around.
- netcraft 5y agoI dont see the answer to a single one of their questions on that page - did you link to where you intended? Edit: you have to click the things on the left, I didn't realize they were tabs.
- viccuad 5y ago> You should read the FAQ at the bottom of the page; I think it answers all of your questions: https://copilot.github.com/#faqs https://copilot.github.com/#faqs Read it all, and the questions still stand. Could you, or any on your team, point me on where the questions are answered? In particular, the FAQ doesn't assure that the "training set from publicly available data" doesn't contain license or patent violations, nor if that code is considered tainted for a particular use.
- res0nat0r 5y agoFrom the faq: > GitHub Copilot is a code synthesizer, not a search engine: the vast majority of the code that it suggests is uniquely generated and has never been seen before. We found that about 0.1% of the time, the suggestion may contain some snippets that are verbatim from the training set. I'm guessing this covers it. I'm not sure if someone posting their code online, but explicitly saying you're not allowed to look at it, getting ingested into this system with billions of other inputs could somehow make you liable in court for some kind of infringement.
- viccuad 5y agothat doesn't include patent violations nor license violations or compatibility between licenses. Which would be the most numerous and non-trivial cases.
- res0nat0r 5y agoHow is it possible to determine if you've violated a random patent from somewhere on the internet via a small snippet of customized auto-generated code? Does everyone in this thread contact their lawyers after cutting and pasting a mergesort example from Stackoverflow that they've modified to fit their needs? Seems folks are reaching a bit.
- IncRnd 5y agoFor that very reason, many companies have policies that forbid copying code from online (especially from StackOverflow).
- dvaun 5y agoNone of the questions and answers in this section hold information about how the generated code affects licensing. None of the links in this section contain information about licensing, either.
- kitsune_ 5y agoSorry Nat, but I don't think it really answers anything. I would argue that using GPL code during training falls under Copilot being a derivative work of said code. I mean if you look at how a language model works, than it's pretty straightforward. The word "code synthesizer" alone insinuates as much. I think this will probably ultimately tested in court.
- gpm 5y ago> - under what license the generated code falls under? Is it even copyrighted? Generally my understand is that to be copyrightable it has to be the output of a human creative process, this doesn't seem to qualify (I am not a lawyer). See also, monkeys can't hold copyright: https://en.wikipedia.org/wiki/Monkey_selfie_copyright_dispute#Naruto_et_al_v._David_Slater https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...
- lawtalkinghuman 5y agoIn the US, yes. Elsewhere, not necessarily.
- tlamponi 5y ago> Is it even copyrighted? Isn't it subject to the licenses the model was created from, as the learning is basically just an automated transformation of the code, which would be still the original license - as else I could just run some minifier, or some other, more elaborate, code transformation, on some FOSS project, for example the Linux kernel, and relicense it under whatever? Does not sound right to me, but IANAL and I also did not really look at how this specific model/s is/are generated. If I did some AI on existing code I'd be quite cautious and group by compatible licences classes, asking the user what their projects licence is and then only use the compatible parts of the models.-Anything else seems not really ethical and rather uncharted territory in law to me, which may not mean much as IANAL and just some random voice on the internet, but FWIW at least I tried to understand quite a few FOSS licences to decide what I can use in projects and what not. Anybody knows of some relevant cases of AI and their input data the model was from, ideally in jurisdictions being the US or any European Country ones?
- gpm 5y ago(Not a lawyer, and only at all familiar with US law, definitely uncharted territory) No, I don't believe it is, at least to the extent that the model isn't just copy and pasting code directly. Creating the model implicates copyright law, that's creating a derivative work. It's probably fair use (transformative, not competing in the market place, etc), but whether or not it is fair use is github's problem and liability, and only if they didn't have a valid license (which they should have for any open source inputs, since they're not distributing the model). I think the output of the model is just straight up not copyrighted though. A license is a grant of rights, you don't need to be granted rights to use code that is not copyrighted. Remember you don't sue for a license violation (that's not illegal), you sue for copyright infringement. You can't violate a copyright that doesn't exist in the first place. Sometimes a "license" is interpreted as a contract rather than a license, in which you agreed to terms and conditions to use the code. But that didn't happen here, you didn't agree to terms and conditions, you weren't even told them, there was no meeting of minds, so that can't be held against you. The "worst case" here (which I doubt is the case - since I doubt this AI implicates any contract-like licenses), is that github violated a contract they agreed to, but I don't think that implicates you, you aren't a party to the contract, there was no meeting of minds, you have a code snippet free of copyright received from github...
- chuinard 5y agoSome of your questions aren't easy to answer. Maybe the first two were OK to ask. Others would probably require lawyers and maybe even courts to decide. This is a pretty cool new product just being shared on an online discussion forum. If you are serious about using it for a company, talk to your lawyers, get in touch with Github's people, and maybe hash out these very specific details on the side. Your comment came off as super negative to me.
- Tainnor 5y ago> This is a pretty cool new product just being shared on an online discussion forum. This is not one lone developer with a passion promoting their cool side-project. It's GitHub, which is an established brand and therefore already has a leg up, promoting their new project for active use. I think in this case, it's very relevant to post these kinds of questions here, since other people will very probably have similar questions.
- ericbarrett 5y agoRegardless of tone, I thought it was chock full of great questions that raised all kinds of important issues, and I’m really curious to hear the answers.
- peddling-brink 5y agoI think these are very important questions. The commenter isn't interrogating some indy programmer. This is a product of a subsidiary of Microsoft, who I guarantee has already had a lawyer, or several, consider these questions.
- king_magic 5y agoNo, they are all entirely reasonable questions. Yeah, they might require lawyers to answer - tough shit. Understanding the legal landscape that ones' product lives in is part of a company's responsibility.
- amelius 5y agoDoes Copilot phone home?
- gpm 5y agoWhen you sign up for the waitlist it asks permission for additional telemetry, so yes. Also the "how it works" image seems to show the actual model is on github's servers.
- heavyset_go 5y agoYes, and with the code you're writing/generating.
- amelius 5y agoThis obviously sucks. Can't companies write code that runs on customer's premises these days? Are they too afraid somebody will extract their deep learning model? I have no other explanation. And the irony is that these companies are effectively transferring their own fears to their customers.
- jmmcd 5y agoIt's a large and gpu-hungry model.
- natfriedman 5y agoIn general: (1) training ML systems on public data is fair use (2) the output belongs to the operator, just like with a compiler. On the training question specifically, you can find OpenAI's position, as submitted to the USPTO here: https://www.uspto.gov/sites/default/files/documents/OpenAI_RFC-84-FR-58141.pdf https://www.uspto.gov/sites/default/files/documents/OpenAI_R... We expect that IP and AI will be an interesting policy discussion around the world in the coming years, and we're eager to participate!