4 ms·
Where it gets ethnically dubious is that: 1. The copilot team rushed to slap a copyright filter on top to keep these verbatim examples from showing up, and now
by ADeerAppeared 2y ago
Where it gets ethnically dubious is that:
1. The copilot team rushed to slap a copyright filter on top to keep these verbatim examples from showing up, and now claims they never happen.
2. LLMs are prone to paraphrasing. Just because you filter out verbatim copies doesn't mean there isn't still copyright infringement/plagiarism/whatever you want to call it. The copyright filter is only a legal protection, not a practical protection against the issue of copyright infringement.
Everyone who knows how these systems work understand this. The copilot FAQ to this day claims that you should run copyright scanning tools on your codebase because your developers might "copy code from an online source or library".
Github has it's own research from 2021 showing that these tools do indeed copy their training data occasionally: https://github.blog/2021-06-30-github-copilot-research-recitation/ https://github.blog/2021-06-30-github-copilot-research-recit...
They clearly know the problem is real. Their own research agreed, their FAQs and legal documents are carefully phrased to avoid admitting it. But rather than owning up to the problem, it's "Ner ner ner ner ner, you can't prove it to a boomer judge".
- squarefoot 2y ago> 1. Isn't that akin to destruction of evidence?
- ADeerAppeared 2y agoLegally? No. In spirit? ... Probably? Unlike most LLMs, Github copilot can trivially solve their copyright problem by just using only code they have the right to reproduce. They have a giant corpus of code tagged with license, SELECT BY license MIT/Equivalent and you're done, problem solved because those licenses explicitly grant permission for this kind of reuse. (It's still not very cash money to take open source work for commercial gain without paying the original authors, and there's a humorous question if MIT-copilot would need to come with a multi-gigabyte attribution file, but everyone widely agrees it's legal and permitted.) The only reason you'd hack a filter on top rather than doing the above is if you'd want to hide the copyright problem. It's an objectively worse solution.
- deleted 2y ago[deleted]
- Spivak 2y ago> Unlike most LLMs, Github copilot can trivially solve their copyright problem by just using only code they have the right to reproduce. Absolutely not trivial, in fact completely impossible by computer alone. You can't determine if you have the right to reproduce a piece of code just by looking at the code and tags themselves. *Taps the color-of-your-bits sign.* * I can fork a GPL project on Github and replace the license file with MIT. Okay to reproduce? * If I license my project as MIT but it includes code I copied inappropriately and don't have the right to reproduce myself, can Github? (No) This one is why indemnity clauses exist on contracted works. * I create a git repo for work and select the MIT license but I don't actually own the copyright on that code and so that license is worthless.
- gkbrk 2y agoThere is no difference when it comes to MIT and GPL here. If your model outputs my MIT licensed code, you still need to provide attribution in the form of a copyright notice as required by the MIT license.
- sleepybrett 2y agoHave the copyleft people, or anyone else, produced some boilerplate licenses that explicitly deny use in training models?
- abigail95 2y agoNot in any way I'm aware of - and would be required if they were served a DMCA notification/Cease and Desist against a specific prompt. The people that think Copilot is infringng their copyright would be happy with that I would think? Unless they take a much stricter definition of fair use than current courts do.
- tpmoney 2y agoNo more so than scanner/printer manufacturers adding tech to prevent you from scanning and printing currency is destruction of evidence that they are in fact producing illegal machines for counterfeiting.
- bawolff 2y agoI would think it is pretty obviously not. Is taking away a drunk driver's keys (before they get in the car) destruction of the evidence of their drunk driving?
- squarefoot 2y agoThis is not what I meant. By placing a copyright filter and claiming it never happened (please read the line I was replying to) before the system can be audited, they're indeed taking away the drunk driver's keys, which is a good thing, but also removing the offending car before Police arrives.
- bawolff 2y agoIn this metaphor, removing the car of someone who was going to drink and drive but didn't, is certainly not a crime. Presumably though you mean removing the car after drunk driving actually took place - which might be, but probably depends a lot on if the person knew, and what the intent of the action was. In the current case - its unclear if any crime took place at all, it seems clear that the primary intent was to prevent future crime not hide evidence of past ones. Most importantly the past version of the app is not destroyed (presumably). Github still has the version of the software without the copyright filter. If relavent and appropriate, the court could order them to produce the original version. It can't be destroying evidence if the evidence was not destroyed.
- squarefoot 2y agoYes, sorta. We're talking about software, therefore a piece of code that does something programmatically isn't like the drunk driver in a car that may cause more accidents, and although we aren't sure about that we prevent him/her to drive anyway just to be safe. The software would most certainly repeat its routine because it has be written to do so, that's why I wondered about destruction of evidence; by removing/modifying it, or placing filters, they would prevent it from repeating the wrongdoing, but also take away any means of auditing the software to find what happened and why.
- bawolff 2y ago> The copilot team rushed to slap a copyright filter on top to keep these verbatim examples from showing up, and now claims they never happen. Well if the copyright filter is working they indeed aren't happening. Putting in safe gaurds to prevent something from happening doesn't mean you're guilty of it. Putting a railing on a balcony doesn't imply the balcony with railing is unsafe. > LLMs are prone to paraphrasing. Just because you filter out verbatim copies doesn't mean there isn't still copyright infringement/plagiarism/whatever you want to call it Copyright infringement and plagerism are different things. Stuff can be copyright infringement without being plagerized, and can be plagerized without being copyright infringement. The two concepts are similar but should not be conflated, especially in a legal context. Courts decide based on laws, not on gut feeling about what is "fair". > They clearly know the problem is real They know the risk is real. That is not the same thing as saying that they actually comitted copyright infringement. A risk of something happening is not the same as actually doing the thing. > "Ner ner ner ner ner, you can't prove it to a boomer judge". Its always a cop-out to assume that they lost the argument because the judge didn't understand. I suspect the judge understood just fine but the law and the evidence simply wasn't on their side.
- FireBeyond 2y ago> Well if the copyright filter is working they indeed aren't happening. Putting in safe gaurds to prevent something from happening doesn't mean you're guilty of it. Putting a railing on a balcony doesn't imply the balcony with railing is unsafe. Doesn't mean you weren't, at some point, guilty of it, either. It doesn't retcon things.
- Dylan16807 2y agoYeah but I think the main concern in this situation is copilot moving forward, not their past mistakes.
- bawolff 2y agoSure, which is why we require evidence of wrong doing. Otherwise its just a witch hunt. After all, you yourself probably cannot prove that you didn't commit the same offense at some point in time in the past. Like Russel's teapot, its almost always impossible to disprove something like that.
- nl 2y ago> Just because you filter out verbatim copies doesn't mean there isn't still copyright infringement/plagiarism/whatever you want to call it. Actually, it does. The production of the output is what matters here.
- kelnos 2y agoIf you copy someone else's copyrighted work and then rearrange a few lines and rename a few things, you're probably still infringing.
- Spivak 2y agoFor a book or a song, for sure, although that isn't really punished. Search the drama surrounding a popular YA author in the 10's, Cassandra Claire. For code since you can only copy the form and not the function that might actually be enough. People do clean room implementations because of paranoia, not because it's actually a necessary requirement.
- Retric 2y agoMoving a few things around means your internal process already had copywrite infringement.
- Spivak 2y agoProbably not. Copyright infringement in the manner we're talking about presumes you already have license to access the code (like how Github does). What you don't have license to do is distribute the code -- entirely or not without meeting certain conditions. You're perfectly free to do whatever naughty things you want with the code, sans run it, in private. The literal act of making modifications isn't infringement until you distribute those modifications -- and we're talking about a situation where you've changed the code enough that it isn't considered a derivative work anymore (apparently) so that's kosher.
- 2y ago
- dspillett 2y ago> The copilot team rushed to slap a copyright filter on top to keep these verbatim examples from showing up, and now claims they never happen. More than that: the fact that they claimed it wasn't possible before adding the filter, to filter out the thing that said wasn't possible. This doesn't help me trust anything else they might say or have already said. My take on that was always: if it isn't possible, then why are MS not training the AIs on their internal code (like that for Office, in the case of MS with their copilot product) as well as public code? There must be good examples for it to learn from in there, unless of course they thing public code is massively better than their internal works.
- Aeolun 2y agoHow do you know they aren’t training it on their internal code? Since you really need to work hard to make the AI spit out anything verbatim, and you have no knowledge of their internal code, how could you ever prove or deny it?
- dspillett 2y ago> How do you know they aren’t training it on their internal code? Because if they were, they would have said. It would be an excellent answer to the concerns being discussed here: “we are so sure that there is nothing to worry about in this regard, that we are using our own code as well as the stuff we've schlepped from github and other public sources”.