14 ms·
An open source lawyer’s view on the copilot class action lawsuit
- nomilk 4y agoIf organic neural networks are allowed to read and learn from open source code, why should an artificial one be any different?
- geysersam 4y ago1. Humans are not neural networks. 2. Humans are not allowed to directly copy even rather short snippets of licenced code. 3. Humans do not have the capacity to memorize the entirity GitHub.
- fhd2 4y agoI can't shake the feeling that a lot of the logic around ML models having more or less the same "rights" as humans comes from misleading marketing that they, in any shape or form, resemble human intelligence. AI is a buzzword applied to any kind of algorithm for an activity that people previously thought couldn't be automated. Back when I was young, graph pathfinding algorithms where called AI. A few decades later they are a well understood commodity and I haven't seen anyone call them AI for a while. Maybe that'll happen to LLMs too, given a few years?
- nomilk 4y agoAn argument in favour of legality of web scraping is if a human can look at websites and collect data, then why shouldn't they be allowed to do the same programatically? This is the same but for use of open source code: if humans are allowed to use one specific (organic) neural network to read, process, and use open source code, then why shouldn't they be allowed to use some other neural network, artificial or otherwise.
- melagonster 4y agoBut the analogue on code is not machine learning, it should be automatically download code.
- nomilk 4y agoThe specifics don't matter so much as the general idea that if a human can do it (anything), then why can't the human make a tool that can do it from them, thus saving them the work.
- jimktrains2 4y agoNo, scraping stems from a service not placing any limits on its access cannot complain that it was accessed. With code, that is denoted via the license, which when supplied with the code and especially as metadata before downloading (as is the case with GitHub) is the common means with which those limits are placed. Humans and neural networks process information very differently and it's disingenuous to imply otherwise.
- geysersam 4y agoThis is the slippery slope argument. It's not inconsistent to allow human "webscraping" while disallowing massive machine web scraping. Most important, it's about what the owner of the website considers to be appropriate. A neural network is closer to a database than a human brain. So this is akin to saying: I can store your personal data in my human brain (without your consent), why am I not allowed to do it in PostgreSQL?
- deleted 4y ago[deleted]
- throwaway290 4y agoFor one, an organic network (for the sake of the argument I'll play along if you want to reduce a human to this) has rights, freedoms and ethical values and is not controlled by a single entity and has not specifically been instantiated to generate profit for such.
- visarga 4y agoI think copyright itself might be on its way out. What meaning does a copyright have when I can click "Variations" on anything and get 4 suggestions in 10 seconds? Imagine how good they will be by 2030.
- izacus 4y agoThere has never been more support for tightening and enforcing copyright than there is today. This is very unlikely to change due to megacorps like Microsoft, Disney, Apple et.al. having a massive vested interest to use it to extract maximum profits.
- classified 4y agoCopyright protection for the rich and powerful, while those who cannot afford armies of lawyers get their stuff stolen by machine learning models. Sounds credible to me.
- 6stringmerc 4y agoI’m not rich and powerful but I’m glad Disney is going to destroy MystickInk on principle. Check out my essay I’ve submitted recently.
- guntars 4y agoI find Copilot most useful for filling out debug statements such as this: println(“foo at {:x} is {:?}”, &foo as *const _ as usize, foo); It almost always writes what I would have. How DARE I steal from open source contributors like that?!
- 3836293648 4y agoGiven how rust-y that looks, are you aware of `dbg!`? Debug printing + file name + line number and it does it properly to stderr
- classified 4y ago
- LesZedCB 4y agoout of curiosity, would anybody else cease to have an issue copilot if it was an open source model? i'm not paying for copilot right now because i'm waiting for this to shake out. but i'd be happy to pay (even their current asking price) if i knew the model was also open source and could be self hosted. maybe this is the wrong way to ask the question, but hopefully it makes sense
- runnerup 4y agoIf it was GPL it could use GPL code and legally there would be no debate.
- deleted 4y ago[deleted]
- Jweb_Guru 4y agoThe project could, yes. It wouldn't necessarily change the legality of using it in non-GPL projects, though. If people were only using it in license-compatible projects and it was license-compatible with GPL, I doubt anyone would have any complaints (even though in theory it could also be picking up stuff from other incompatible licenses).
- comice 4y agoOne of the requirements of the GPL is that credit is given (and indeed this is needed for enforcement to work because the GPL leverages copyright).
- throwaway290 4y agoIf it was a true OSS project, first it would not clearly benefit a single near-monopoly by using my code (as in, that wouldn't be its purpose), and second I'm sure its contributors would be well placed to understand the issue and from the start bake in a reliable, transparent mechanism for opting out. As is, it's EEE applied to open source-- Microsoft's ultimate play against the ethic that brought us Linux among other things. When your brainchild gets gobbled up faster than you can blink, pushed to people who never learn about your existence, and a megacorp that you are ethically opposed to profits from the process, the need for self-actualization is no longer addressed. The fundamental incentive that pushes us to publish in the open, to have other humans acknowledge you and your work and feel pride in it, is being eliminated.
- steve_gh 4y agoHmmm. I'm interested in the GitHub ToS, which (if I understand correctly) basically says that GitHub and it's affiliates (MS) can use anything you post on GitHub to improve their service. What if I build an AGPL licenced service, using GitHub to coordinate development. According to the ToS MS could offer a version my service because I posted the code on GitHub, and they are using it to improve their service to me. According to my AGPL licence, they would need to share their source. So which takes precedence. The licence or the ToS?
- david_allison 4y agoAs a follow-on, what if you're mirroring code which is under an AGPL license? Are you allowed to post it on GitHub if you can't grant those rights under the ToS due to the license of the code?
- NicoJuicy 4y agoTheir service is hosting code, not writing code. That's why it's GitHub, not CodeScribe ( or something)
- Xylakant 4y agoIt’s definitely more than just hosting code - GitHub offers issue/PR management, light weight project management, an online IDE for collaborative editing and CI services at least. Arguing that GitHub provides services that aim to improve developer/development team productivity is not a stretch. And arguing that ML-assisted development support is part of that definition isn’t particularly far out either.
- frumper 4y agoThey're also allowed to add new services anytime they'd like.
- amarant 4y agoProbably the ToS. You've granted GitHub specifically license to use your code under the terms of the ToS, they effectively have 2 licenses. They can therefore choose under which licence they want to use your code, and will choose the most permissive one, or the one they have the best understanding of: in this case the ToS. Other parties are not granted license under the ToS, and so will have to abide by the AGPL.
- Havoc 4y agoWhat’s the point of licenses if TOS overrides it?
- amarant 4y agoThe ToS only applies to GitHub(which includes Microsoft, apparently) Other parties will still have to abide by your license.
- junon 4y agoGithub's TOS doesn't infringe on any licenses. https://docs.github.com/en/site-policy/github-terms/github-terms-of-service#d-user-generated-content https://docs.github.com/en/site-policy/github-terms/github-t... I'm actually surprised they allowed Copilot to happen, given this section: > This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service, except that as part of the right to archive Your Content, GitHub may permit our partners to store and archive Your Content in public repositories in connection with the GitHub Arctic Code Vault and GitHub Archive Program. One could make the argument they had no intrinsic right to use the software for Copilot except under the terms laid out under the respective softwares' licenses. This means any GPL code they copied by error is now in violation of the GPL by default. But IANAL.
- World177 4y agoIn my memory, when GitHub released it, they were explicit that using data like this “is common practice in machine learning.” Though, I tried to find the quote and couldn’t, so maybe my memory is wrong and I am remembering a blog post from another organization. edit: The exact quote was “Training machine learning models on publicly available data is considered fair use across the machine learning community” if you want to search for it. edit 2: https://web.archive.org/web/20210629142841/http://copilot.github.com/ https://web.archive.org/web/20210629142841/http://copilot.gi... > Frequently Asked Questions -> Training Set -> Why was GitHub Copilot trained on data from publicly available sources? > Training machine learning models on publicly available data is now common practice across the machine learning community. The models gain insight and accuracy from the public collective intelligence. But this is a new space, and we are keen to engage in a discussion with developers on these topics and lead the industry in setting appropriate standards for training AI models.
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- terminal_d 4y agoIf this isn't enough incentive to move away from github, then I don't know what is.
- baby 4y agoThis is why we can’t have nice things. Copilot is the future
- MattPalmer1086 4y agoHas anyone produced a legally watertight license or clause for other licenses that prevents code being used for training of copilot-like services?
- rwmj 4y agoIt would be a Field of Endeavor restriction so the resulting license wouldn't be open source, and I don't think (?) Copilot is trained on proprietary code. (Section 6 here: https://opensource.org/osd https://opensource.org/osd)
- MattPalmer1086 4y agoI don't really care if a license meets some arbitrary definition. Let's say I added a clause to my BSD license that prohibits the copying of this code to train ML models. Would that not immediately make GitHub in violation of this license? Or do they only train it where the license is explicitly one of the ones it knows about?
- frumper 4y agoIf you upload it to GitHub, you've already granted them a license to use it to improve their service. You aren't sharing it with GitHub under your custom BSD license, you're sharing it with GitHub users under that.
- MattPalmer1086 4y agoThanks. I finally understood its a separate license grant to GitHub.
- frumper 4y agoAlso, apologies for responding to you twice with the same thing. I think I mixed something up and didn't intend to do that
- 4y ago
- Andrew_nenakhov 4y agoA hypothetical question: imagine a filmmaker, who had studied a lot of obviously copyrighted movies by famous renowned directors. This means he has trained his neural network using their copyrighted licensed content. Does he breach copyright when he composes and films a scene? Are visual quotes copyright theft? Homages? Did George Lucas infringe copyright when he was borrowing compositions from "Triumph of the will"?
- jackdaniel 4y agoI see this argument over and over again, and it is so flawed that it is hard to bear. There is no equal sign between a person and a program. There is also that thing called "scale" that is critical to the interpretation of the action. Is eating meat fine? - maybe. Is eating all animals OK? - Hmm...
- Andrew_nenakhov 4y ago> Is eating meat fine? - maybe. Is eating all animals OK? - Hmm... This argument is hardly less flawed than the one you are criticizing. And you statement that 'there is no equal sign ...' is also unconvincing, as we're not equating these two, but the process of learning, which is quite similar.
- jimktrains2 4y ago> but the process of learning, which is quite similar. Thats the thing, there is no reason to think that they are similar.
- 6stringmerc 4y agoI have a companion piece talking about music and training AI/ML: https://medium.com/@6StringMerc/artificial-intelligence-machine-learning-ai-ml-in-music-generation-the-infinite-monkey-of-be42d4d63e0a https://medium.com/@6StringMerc/artificial-intelligence-mach...
- iLoveOncall 4y agoThis lawsuit is open-source developers destroying open-source.
- insanitybit 4y agoHN is so insanely frustrating, so many comments demonstrate that the user didn't read this article at all. Just immediately jumping into a "but what about this argument that I made?".
- robocat 4y agoPlease don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that." https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- insanitybit 4y agoYeah, I'm aware, this is just so extreme at this point it feels worth pointing out.
- hnbad 4y ago> It looks a lot more like trolling if an otherwise incredibly useful and productivity-boosting technology is being stymied by people who want to receive payouts for a lack of meaningless attributions. This one sentence threw off my entire opinion of the article as it demonstrates the author's clear bias in favor of Copilot, not just specifically in this case but in principle. Legal opinion on Copilot and generative AI in general hinges entirely on metaphors. If the AI is understood to behave like a human being building knowledge and drawing from it for inspiration, Copilot is just another way to write code. But we've already established legal precedent that machines can not hold copyright, which suggests that they can not be deemed to be creative, which could be used to argue that they are therefore just creating an inventory of copyright works and creating mechanical mashups. The author's dismissal also ignores that this would not JUST result in attribution. If Copilot indexed copyleft code and were required to provide attribution when using this code, the output might also be affected and this could in turn affect the entire code base. Worse yet, Copilot may output code with conflicting licenses. The author considers only the possibility that Copilot itself might have to inherit the license (and the dismissal that it would "help noone" because it runs on a server ignores both the existence of a (presumably self-hosted) enterprise service and the existence of licenses like AGPL, which would still apply) but it seems most people's concerns are with the output instead. I also fail to understand how the argument that it doesn't reproduce the code exactly 99% of the time is helpful. If I copy code and rename the variables and run an autoformatter on it, it's still a copy of the code. It's odd to see a lawyer use what is essentially obfuscation as a defense against copyright claims. Also 1% is an incredibly large number given how Copilot is supposed to be used and how large the potential customer base is. Given the direction GitHub is heading with "Hello GitHub" (demoed at GitHub Universe yesterday) it's not unlikely that Copilot would in some cases be used to generate hundreds, thousands or tens of thousands of lines of code in a single project. The question isn't just whether Copilot is violating the law or not, the question is why it is or isn't because that could have wide implications outside GitHub itself. But as the author points out, sadly the lawsuit doesn't try to settle this for copyright, which might be the most impactful question.
- hfglanx 4y ago
- tallanvor 4y agoWhether or not other countries give you the right to enforce your copyright even if you haven't registered it with the government is not relevant for a class action lawsuit filed in the US.
- muraiki 4y agoOr it could be that she is experienced with both software and law, and that her assessment is different than yours. > Kate’s passion for open source began in law school, under the tutelage of Eben Moglen, long-time attorney for the Free Software Foundation, founder of the Software Freedom Law Center, and author of the GPL 3. She interned at the Electronic Frontier Foundation and helped write the first complaint against the NSA for warrantless wiretapping. > At VMware and ServiceNow, she dedicated her time to designing, building, and testing internal compliance tools in collaboration with their respective internal tools teams. She is no stranger to writing specs, creating wireframes, and massive amounts of QA. So much so, that Kate and her husband, Steve Downing, co-founded Critterdom LLC, a software company whose Open Sorcerer product substantially cuts down the time it takes to manually review source code for licenses and create a customer-facing disclosure of that source code. https://katedowninglaw.com/about/ https://katedowninglaw.com/about/
- belorn 4y agoA very interesting interpretation of the github TOS. Kate Downin is saying that users of github is giving a special license to GitHub, one that bypasses the original license. However if that is true then any upload of code that users do not have 100% copyright control of is then a copyright violation since the user would not have the authority to grant github that special license. It would be similar to a user uploading a copyrighted movie to youtube, and google using that as a license to use the movie in an advertisement. I wonder if a court would think that microsoft in this case has done their due diligent to verify that the license grant that they got from users are correct and in order.
- dathinab 4y agoIt also falls under the aspect of "hidden surprises" which could mean that this part of the TOS wrt. this specific aspect might not be legally binding/valid. At least in the EU. Or it might.
- TazeTSchnitzel 4y ago> if that is true then any upload of code that users do not have 100% copyright control of is then a copyright violation since the user would not have the authority to grant github that special license That doesn't sound right. Licences can allow sublicensing, and I think all the popular open-source ones do.
- belorn 4y agoSublicensing can only create additional restrictions on top of the existing conditions inside the license. All open source licenses require at minimum that distribution provides attribution and the original copyright notice. License like GPL has additional conditions. There is also additional problems specific to sublicenses. In the United States, only exclusive licensees are assumed by statute to have a right to sublicense. The theory is that licensees of exclusive licensees are assumed to have the control/authority similar to that of the author. Nonexclusive licensees are not assumed to be granted such a monopoly by the licensor.
- hyperman1 4y ago
- mjw1007 4y agoI think this is the most interesting part: > [Github's Terms of Service] specifically identifies “GitHub” to include all of its affiliates (like Microsoft) and users of GitHub grant GitHub the right to use their content to perform and improve the “Service.” Diligent product counsel will not be surprised to learn that “Service” is defined as any services provided by “GitHub,” i.e. including all of GitHub’s affiliates.
- tryre 4y agoNo, the misinterpretation of the ToS is not the most interesting part. The part that clearly shows her colors is: "It looks a lot more like trolling if an otherwise incredibly useful and productivity-boosting technology is being stymied by people who want to receive payouts for a lack of meaningless attributions."
- 1MachineElf 4y agoAh, so she is an "open source lawyer" in an OSI Foundation sense...