4 ms·
Of course it is cherry picked. The idea is that it allows you to INTENTIONALLY void any copyright you want. So let's say I obtain an illegal copy of microsoft
by cowtools 4y ago
Of course it is cherry picked. The idea is that it allows you to INTENTIONALLY void any copyright you want.
So let's say I obtain an illegal copy of microsoft windows' source code. Under this precedent, what stops me from just (overfitting) training a neural network to produce the source code verbatim, sans any license notice?
But it doesn't end there. What stops me from making a neural network that exactly reproduces the bytes of Illegally_Ripped_Disney_Movie.mp4 that I obtain from the pirate bay? Copyright need not apply.
At what point can the neural network I've described (which is intentionally designed to violate copyright) distinguishable from a neural network like Copilot and others which violate copyright extrinsically?
- jeremyjh 4y agoIt doesn't void anything. If you use Copilot to copy some licensed code illegally, you are the person in breach, not the tool. People using Copilot are possibly littering their code-base with future copyright liabilities, and they'd have no idea about it. Until someone writes an AI to find infringing software and automatically sue them...
- cowtools 4y agoYes, that's the other outcome, which is the industry getting GPL'd to death for using this. But I don't think there is an established precedent when it comes to this. I don't know much about AI/SL but I think the logical outcome is that using a transformer to generate source code is going to result in an over-trained system that produces sections of code verbatim because of the relative size of the space of all valid programs vs the space of all possible programs. It's not like art where you get soft failure if a single pixel or word is wrong: the inclusion, exclusion, or replacement of a single instruction or symbol is enough to introduce fatal bugs into a computer program. If the system doesn't have the capacity to understand programs generally (which it probably won't if it's just a transformer), then you're going to end up with a system that spits out samples from training data that do work.
- nickstinemates 4y agoA lot of M&A activity use tools like Blackduck software that does exactly this. It flags partials.
- Aeolun 4y agoIt’s fairly obvious if Copilot is regurgitating entire blocks of code that someone else wrote. In 999 out of a 1000 cases it’s just spitting out boilerplate though.
- gfodor 4y agoYep - it would be useful if more people had literacy of using the tool for these conversations. I don't blame them, that shouldn't be expected or required, but there is a large gap between how bad this looks and how materially bad it is when you take into account the actual way Copilot is usually used.
- jeremyjh 4y agoYes but if you have a large team you'll be using it hundreds of times a day. I will not be surprised if Copilot indemnity insurance is a thing in M&A in a five years.
- BeefWellington 4y agoIf you told the legal team at any mid-sized or larger company "we're pretty sure only 1 in 1000 lines of code our developers write breaches someone else's copyright" there'd be some serious hell to pay.
- Aeolun 4y agoNo, no. 1 out of every 1000 lines of code has the potential to breach some form of license. If someone is motivated to search through our entire (proprietary, private) codebase. They match it with repositories that are freely available. They’re properly motivated to make a problem out of it (some twitter randos?), and most importantly they gain some benefit out of spending hundreds of thousands of dollars engaging with our legal team. By the time you satisfy all the conditions required for it to be an issue you are talking nation-state actors.
- mmgutz 4y ago> to copy some licensed [anything] illegally, you are the person in breach, not the tool ahem napster, pirate bay ...
- rockemsockem 4y agoI don't think there will be a ruling like "anything from a neural net is yours", that'd be a bit ridiculous for very obvious reasons. I'm no copyright law expert and I'm certainly not a lawyer, but it seems to me that in your examples you're setting out to violate copyright as a goal, which seems like it would be a factor in a court case. To answer your last question, your examples are pretty clearly distinguishable from Copilot in their final states that you describe. IDK exactly *when* during overfitting that line is crossed, maybe it's crossed the moment you personally decide to knowingly publish copyrighted content and has nothing to do with the neural network itself?
- lamontcg 4y ago> The idea is that it allows you to INTENTIONALLY void any copyright you want. It doesn't void copyright. Anyone that uses code that Copilot spits out which infringes on someone else's copyright is still liable. There's no requirement for intent. That may be a mitigating factor in terms of remediation, but it cannot void the copyright itself. It can, however, produce a plague of completely ignorant copyright infringement, and since the user of Copilot has no idea where the code is coming from, there's no way to check if it was trained on infringing code. If I used Copilot I would be real worried about the legal implications for me and that I could easily be accused of copyright infringement or plagiarism[*]. Of course people seem to not really care these days if they can cheat to get ahead so this is probably a feature, and 99.9% of user won't have their reputations ruined by using it. [*] Although I'm personally more worried about the fact that most of the code will be wikipedia/blogs-quality and filled with bugs, edge cases and performance issues.
- anon2020dot00 4y agoBuggy code can easily come from copying Stack Overflow or even from Github open-source repos so I doubt CoPilot would contribute to it anymore than it already is.
- int_19h 4y agoLarge companies don't worry about this because they already have tooling that scans code and matches it against other public code out there (to avoid legal troubles if some dev quietly copies some GPL'd code or something like that). Everybody else basically has to manually check every snippet produced, which makes all the supposed convenience moot. Between that and the typical quality of Copilot snippets, it's pretty obvious who the real beneficiary of this technology is: large sweatshops like Infosys.
- abigail95 4y agoThis doesn't make sense, a byte identical copy of other work is obviously not transformed. So anyone claiming fair use relying heavily on transformation would fail. But the only thing that does is make step 3 of fair use harder to clear. Not impossible. There are fair uses of copyright that use the entire identical work as is.
- elikoga 4y agoLike this? https://github.com/ZoloZiak/WinNT4 https://github.com/ZoloZiak/WinNT4