7 ms·
> “Your honor, we needed so many works that it was simply not practical to ask permission of the creators.” I don’t find this argument convincing given the abil
by distcs 4y ago
> “Your honor, we needed so many works that it was simply not practical to ask permission of the creators.” I don’t find this argument convincing given the ability today to license many content types at scale for TDM, including images, music and yes, journal articles (See “Full disclosure” above), but it is an argument often offered by infringers.
Why is this type of argument even valid? Isn't this fundamentally saying, "The cost of not infringing copyright is massive, so we will glibly infringe!"
So it is not okay to infringe copyright at a small scale but okay to do it in a large scale? How can such a line of argument be sensible in court? But apparently infringers are using this line of argument. So how? Is it not absurd?
- intelVISA 4y ago>So how? Is it not absurd? it is. "Your honor, I've been drink-driving so many times I honestly can't tell you an accurate estimate any more so in both our interests let's agree it was basically uncountable, or 0." wait a minute...
- JamesSwift 4y agoOr after a crypto heist: "Your honor, it would have been onerous to ask each account owner if we could have their tokens so we just took all of them at once"
- judge2020 4y agoOn the other side, they could argue that it's like a human learning how to code over a decade of looking at the internet, and that human doesn't need to DM every code author to ask if they can learn from their content (and the risk for the author is similar given the human might one day recall some author's code verbatim and not give attribution).
- distcs 4y ago> On the other side, they could argue that it's like a human learning how to code over a decade of looking at the internet And that would make sense and it would be argued on its own merit. The judge/jury will decide if this argument is correct and legal. But the article implies that there are lawyers and infringers out there who are arguing that they could not have possibly afforded the cost of not infringing, so they were justified in their infringement. Since when did the massive cost of avoiding infringement become a valid reason to carry on with infringement? This seems just plain absurd by common sense. How do lawyers and infringers make this argument? How is it even entertained in court? What am I missing?
- panzi 4y agoSo if I make a script that automatically downloads every torrent in existence it's suddenly ok, since it is infeasible to check the copyright of them all?
- SamoyedFurFluff 4y agoBut the thing is that we explicitly allow humans to learn and develop their own skills learning from other humans, but we have our own taboos around directly copying peoples work without permission and passing it off as your own. The debate is that copilot isn’t a human, it’s a machine that outputs copied work on a statistical basis. Humans are allowed to be unoriginal, uncreative, boring, mediocre, and all sorts of things. But they’re not copying whole cloth the way copilot is.
- judge2020 4y ago> But they’re not copying whole cloth the way copilot is. Stack Overflow content is CC-BY-SA 4.0 yet I can bet most corporate codebases include tons of code snippets without a link or citation to the original answer
- throwaway4aday 4y agoWhenever I've used Copilot it never seems to copy whole sections of code. Can you provide examples of this? From what I've seen it is producing fairly generic boilerplate that has been modified based on the rest of the code in my repo so that it works with the other functions and even incorporates other pieces of my code in the same style that I'm using. The boilerplate aspect makes sense because this would be the most common sequence of tokens that it observed during training. It's somewhat miraculous that it can incorporate code on the fly from my repo. I've never seen anything that looks like a direct copy paste from elsewhere though. If you have a different observation I'd love to see it.
- EamonnMR 4y agoBehold: https://twitter.com/StefanKarpinski/status/1410971061181681674 https://twitter.com/StefanKarpinski/status/14109710611816816... Probably helps that this is from a codebase that's been forked quite a bit.
- judge2020 4y agoYou can't even code search in forked repos so maybe forks were excluded (besides commits on top of the fork)?
- simion314 4y ago>On the other side, they could argue that it's like a human learning how to code over a decade of looking at the internet, and that human doesn't need to DM every code author to ask if they can learn from their content (and the risk for the author is similar given the human might one day recall some author's code verbatim and not give attribution). It might not be that easy, I think Wine developers are not allowed to read code related to Windows, even if this code is published on GitHub. The fact you looked at the code was decided to be a risk. You also have cases with a NN producing an identical output, so you either prove your NN NEVER produces copyrighted code or you have to have a second process that is 100% correct and double checks the NN output and check for plagiarism. I am against Microsoft in this case because they decided not to put their proprietary code in the NN , would have been funny to have the AI write an open source Windows re-implementation when you feed it the Win APi documentation.
- b3morales 4y agoThe Wine developers are allowed to read whatever they want. They may choose to have a policy not to, because it makes it easier for them to prove that they didn't make unlawful use of proprietary code: they can't have copied something that they never read. If you reproduce something independent of knowledge of the original then that is a defense against copyright infringement. This is essentially the "Clean Room" tactic: https://en.wikipedia.org/wiki/Clean_room_design https://en.wikipedia.org/wiki/Clean_room_design
- stackbutterflow 4y agoWe allow humans to do what copilot does because we take into account that the human brain is very limited in this regard. If we could scan all of GitHub in under a week and recall perfectly what we saw we would already have different laws. Now that machines are able to somewhat learn like humans but 1,000,000 faster we need new laws. That's why I don't believe "but that's like humans doing X" is a strong argument.
- ClumsyPilot 4y agoIf copilot is so advanced that we need to grant it right human has, then it has right to freedom and owning copilot is a crime. I dony think Microsoft wants to go down this path
- CJefferson 4y agoIf enough code is recalled verbatim, I can sue the author of that code. That seems to fit entirely with this case -- they are suing the owner of Copilot, partially because it reproduces chunks of code.
- carom 4y agoIt is not a human though. It is a function approximated from inputs and outputs. The laws are different and the licenses call out derivative works.
- chlorion 4y agoWhy are wine developers and similar required to do clean room implementations to not be sued then? Simply reading the leaked source code of Windows makes you not eligible to contribute to wine. Why is Windows source code so much more important than mine? The other thing is that copilot is not a human, so it doesn't matter anyways. Humans are a special exception with laws, because they are intended to protect and benefit humans while also being fair. I don't think you can just substitute something in and assume that the same rules apply.
- EMIRELADERO 4y ago> Why are wine developers and similar required to do clean room implementations to not be sued then? That's the neat part, they don't. It's essentially a self-imposed limitation which contradicts actual court rulings on the matter such as Sony v. Connectix, in which the court commented on clean-room being "inefficient" and the kind of inefficiency that fair use was "designed to prevent".
- malfist 4y agoIt's the same argument people make about why crypto doesn't have to follow the laws on Know Your Customer. Because someone designed the crypto to break that law, so their hands are tied, it's too technically hard to comply.
- JimDabell 4y ago> But apparently infringers are using this line of argument. So how? Is it not absurd? You realise that they haven’t actually used that line of argument, right? The article author speculated that it might be part of the defence and then said they didn’t find it compelling. Set up a straw man and then knocked it down in virtually the same breath. Don’t waste your time complaining about legal arguments that have not been made except in the imagination of one author.
- turmeric_root 4y ago[flagged]
- SketchySeaBeast 4y ago[flagged]
- crazygringo 4y ago> So it is not okay to infringe copyright at a small scale but okay to do it in a large scale? No, I think you're missing the "transformative" part. The line of argument isn't "we're going to resell millions of codebases as-is for pure profit", which would be undisputed copyright infringement. The argument is that something highly transformative (e.g. training models) isn't infringement at all, because transformative works are covered by fair use. And that, if we still wanted to explore interpreting/changing the law to force opt-in for highly transformative things, it's logistically unreasonable, to such an extent that the transformative thing couldn't occur at all. So that it's a waste of time to even be discussing asking for permission as some kind of potential compromise or requirement. If it's transformative and therefore fair use, asking for permission is an irrelevant distraction. That's why this type of argument is valid. I'm not saying whether the argument will/should win in this particular case, but I'm definitely saying there's nothing absurd whatsoever about it.
- jfk13 4y agoYes, transformative works may be allowed. So I'd guess that creating a model is probably OK (speaking as a non-lawyer!). But using output generated by that model is another matter. The "model" is fundamentally a machine that produces output that is derived from the input it was given. And that output might not be sufficiently transformative to "escape" copyright/licensing restrictions. In the extreme case, the model's output might be a verbatim copy of a large portion of the original input ("training materials"); but even if it has been extensively modified, e.g. to conform to the coding style of a target repository or to follow a different language standard, this might not be "transformative". (Compare: A translation of Harry Potter to French looks superficially quite different from the English original, yet it is still a derivative work; and if you're planning to publish one, Ms Rowling (or her publisher) may want a word with you. And that would apply whether you translated it "manually" or pushed it through Google Translate.)
- deleted 4y ago[deleted]
- distcs 4y ago
- toyg 4y ago> Isn't this fundamentally saying, "The cost of not infringing copyright is massive, so we will glibly infringe!" Copyright is not a natural human right; it's a construct invented and conferred by governments in order to achieve certain objectives. (It's more like a state license than a right, to be honest; using "right" was a historical masterstroke from the original inventors). As such, if those objectives can be provably achieved in a better way without copyright (or rather cutting copyright a bit smaller), there might well be a case for foregoing punishment. It's an exceptionally-hard argument to make, but it's not illogical.
- icambron 4y agoThis seems like a good argument for adjusting copyright law, but seems unhelpful in interpreting it. "This law isn't a good way to achieve the government's objectives" is not the same as "this law wasn't broken". Judges do have some discretionary power in interpretation and that can take into account congress's intent, but here that would be a massive stretch. A judge would simply say it's congress's job to fix copyright if it's not the best way to achieve certain policy goals.
- toyg 4y agoAs I said, exceptionally hard in practical terms - just not as baffling as the parent poster painted it. With the right judge anything is possible, and US history is full of controversial "overreaching" judgements. (and with the deep pockets GH/MS have, it doesn't really matter if the case eventually loses on a big principle - it's just a case of dragging it long enough that, some time through the whole process of appeals, the plaintiff will get broke enough to give up or accept a deal. This line is likely just one of many that defendants will employ.)
- feoren 4y ago> Copyright is not a natural human right; it's a construct invented and conferred by governments in order to achieve certain objectives. (It's more like a state license than a right, to be honest; using "right" was a historical masterstroke from the original inventors). I don't know of a better definition for "natural human right" than "a right/privilege/protection given to everyone automatically, even if they don't know about it or claim it, unless they specifically opt out of it." We get to decide what our "natural human rights" are, and we've decided that you automatically get copyright on your creative works even if you don't know what copyright is. Seems like a good thing, and a natural human right.