9 ms·
There was many personally identifying information leaks when it initially launched. Which suggests that they snooped through OSS code and violated any licenses.
by Gentil 4y ago
There was many personally identifying information leaks when it initially launched. Which suggests that they snooped through OSS code and violated any licenses. Here are 2 answers to the main two things copilot sympathizers ask.
1. How do you know Co-pilot violated OSS licenses?
> The aforementioned PII and secrets emitted. And IIRC a support personnel said they even copied GPL licensed code. So there is nothing stopping them from copying MIT licensed ones. Which is still a violation without attribution. So no reason for the benefit of the doubt.
2. Humans copy code all the time. So why can't an AI? Or something along the lines of it...
The answer is 2 parted.
1. A human can only copy so much. Most of them are not in bad faith. And those who are in bad faith are not a humongous portion. An AI just scrapes the code in mass scale which is impossible for a human even with automation. It's also easier for it to learn from this training data compared to someone scrapping the data and maintaining a DB. This is such a common sense thing. None of the laws, policies or systems are made for AI in mind. So it is just a careless thing to say.
2. If are not adhering to OSS licenses then YOU ARE in violation. So if you are not called out or punished, DOESN'T MEAN you are right. Just means you are getting away with it.
This is just a PR exercise for MS. They already have the data and the product. A nobody has the balls to question them. The normalization of these bad behaviors with their good PR being "the guardians of the OSS" is literally taunting the entire OSS community.
This is where I hoped FSF, EFF, SFConservancy or any other group would step up. Nothing so far. Just yelling into the void hoping MS will suddenly do the right thing.
- nojito 4y agoCoPilot using code to train a model doesn’t violate any software licenses.
- ClumsyPilot 4y agothe model includes chunks of that code, they copied it into their model
- SXX 4y agoLet's train AI on leaked source code of Microsoft, Nvidia and other commercial companies and find out how soon such project going to be sued and shutdown via DMCA. Might also include source-available code like Unreal Engine, etc.
- orf 4y agoNot having the resources to defend yourself against large and litigious companies does not mean parent’s point is wrong. Training CoPilot on code doesn’t violate any software licenses.
- tyingq 4y ago>Training CoPilot on code doesn’t violate any software licenses. Even if it regurgitates large unchanged passages of code, with the advertised selling point of using anything it produces in an unrestricted fashion?
- synu 4y agoI think they are saying the training part itself doesn’t. It seems to be kind of a fine point, otherwise you could implement a license laundering AI by training it on a single codebase and asking it questions about it. If you end up with a byte by byte replica sans a LICENSE file presumably you wouldn’t get away with actually _using_ the model to clone licensed software even if training it was technically ok.
- they4kman 4y agoReading and remembering code is allowed under all the OSS licenses. It's the reproduction of the code that's restricted. The blurry question is always: how much does an expression have to change between it being classified as an exact reproduction, a derivative work, and a novel work? CoPilot would definitely fail the clean room test, though
- numpad0 4y agoThis is something I don't understand, not just about Copilot but for many NN generators: the outputs sometimes seem like an obvious ripoff of something, that no way it qualify as a novel work, although I can't prove it. Yet the outputs are treated as if all copyright issues are discussed and cleared. Just how come?
- formerly_proven 4y ago> Yet the outputs are treated as if all copyright issues are discussed and cleared. Just how come? Anyone developing and trying to sell ML models will try hard to act like this is a settled question. This is obviously a huge uncertainty factor for using various commercial ML models.
- nojito 4y agoThat doesn’t have anything to do with my comment. The model is 100% compliant with open source licenses.
- synu 4y agoAre you drawing a distinction between creating the model, which depends on viewing open source code and is presumably fine, and using it selling the model which may output licensed code and get you in trouble?
- ipaddr 4y agoThe copyright issues are downloaded to the user/developer. If co-pilot suggests anything under copyright and you publish it you get sued.
- Imnimo 4y agoWhat if I train a model on one snippet of code, and it always produces that snippet of code for all inputs? If I did not have a license to use that code, would laundering it through a model absolve me of violating the license?
- tessierashpool 4y ago> CoPilot using code to train a model doesn’t violate any software licenses. The argument GitHub advanced for this was absolutely ridiculous: that code stops being code when CoPilot analyzes it, qualifies only as text while CoPilot analyzes or suggests it, but magically turns back into code if/when the user incorporates the suggested text into their code base. The real argument GitHub made was much more credible: the CEO came on here and said something like "the legality of code reuse is an interesting new area of law and we welcome the debate," which in lawyer terms means "we have smart, well-informed lawyers, and we can easily get these cases in front of dipshit judges who don't understand technology." It's an interesting product, and I hear good things about its utility in practice. But the legal argument is utter nonsense. CoPilot definitely violates many software licenses, and knowingly so.
- mysore 4y agowho cares if it copies licensed code? if it increases developer productivity by 10-20%, its a huge net win for global prosperity.
- teakettle42 4y agoWho cares? Anyone who believes that plagiarism is wrong. Not to mention anyone that believes in respecting other people’s licensing terms when they’ve provided their code to the world for free.
- Gentil 4y agoThis view will stand only until you are effected directly. Which is a very careless way to look at it.
- treesprite82 4y ago> There was many personally identifying information leaks when it initially launched > The aforementioned PII and secrets emitted It's not trained on private GitHub repos. Any secrets it generated would have been already public and compromised. > Which suggests that they snooped through OSS code and violated any licenses That it was trained on FOSS code is already known. Whether doing so violates the licenses is probably up for courts to decide. > An AI just scrapes the code in mass scale which is impossible for a human even with automation Human programmers can do this with search engines.
- Gentil 4y ago> Any secrets it generated would have been already public and compromised. Two things.. 1. This is not some vigilante hacking groups you are talking about. This is MS - a humongous corp. They know doing this is wrong. So it shouldn't matter. It is still illegal and making a proprietary software product is exactly the reason why these things exist. Don't you think? 2. Already public unintentionally. AFAIK, any code that I produce is automatically copyrighted to me. This means if I write something in public and not provide a license, IT IS LEGALLY under the copyright protection provided to me by my country. At least that is the case in US and India which are home to a huge portion of OSS. Do correct me if I am wrong. Putting it like that in public would be just plain stupid for sure. But legally it is still mine. Reproducing it and remixing my work would be illegal. It is just my PITA now to prove that in court is all. > Whether doing so violates the licenses is probably up for courts to decide. I am yet to see any attribution to ANY OSS code that it is trained on. The PII and secrets is enough to find out the license of a repo which would make it easier to prove whether they violated it or not. Don't tell me all OSS that copilot has trained on is only public domain stuff. Even ISC license needs attribution. > Human programmers can do this with search engines. OK. CAN. And that is wrong too. So when they do, we should incriminate them as well if evidence is out there or if a matter such as this comes to light. How does that change anything I mentioned? Very curious.
- treesprite82 4y ago> So it shouldn't matter. It definitely does matter that any secrets generated were already public: * Emitting secrets from private repos would be a huge confidentiality issue (though really you shouldn't commit code secrets to git at all), as it'd be taking something that's private + exploitable and making it public * Emitting secrets that are already public doesn't cause the confidentiality issue. Once a secret is out, it's out, and should be changed immediately. By the time it's in Copilot's training set, it'll have already been on search engines/archive sites/black-hat forums/etc. Tangentially, GitHub do also do some scanning to alert of accidentally committed secrets in repos: https://docs.github.com/en/code-security/secret-scanning/about-secret-scanning https://docs.github.com/en/code-security/secret-scanning/abo... > 2. Already public unintentionally. Right, but therefore already compromised and no longer confidential. Copilot isn't leaking any secrets, someone else did by making them public. > AFAIK, any code that I produce is automatically copyrighted to me. This means if I write something in public and not provide a license, IT IS LEGALLY under the copyright protection provided to me by my country. At least that is the case in US and India which are home to a huge portion of OSS. Essentially correct, to my understanding. If you're making it public, you'll generally also give some hosting/publishing/distribution rights to the services involved - as specified by their T&C. > Reproducing it and remixing my work would be illegal The US has the concept of fair use which provides exceptions for "transformative” purposes. For example: copying and downscaling your image to use as a thumbnail, caching the webpage your work is on, or creating a parody of your work. Consider Google Books for example, where Google scanned millions of copyrighted books and made them searchable (showing snippets). This was ruled fair use due to being transformative. Question would be whether code generated by Copilot that falls under this. Ultimately it's up to the courts to decide, but I'd lean in favor of "yes". > The PII and secrets is enough to find out the license of a repo which would make it easier to prove whether they violated it or not. Don't tell me all OSS that copilot has trained on is only public domain stuff. Even ISC license needs attribution. Fair use is about unlicensed usage, so if it's fair use then it doesn't need to abide by the terms of the licenses. Even if it's ruled not to be fair use, I think they could still train it on GitHub-hosted code due to the mentioned rights you give them by agreeing to GitHub's T&C. > How does that change anything I mentioned? Very curious. Changes your claim of impossibility, so now it's just about whether there's a violation.