4 ms·
I've heard something similar in response to Copilot in another thread (something like offering a sum of money to Github if they train their model exclusively on
by invokestatic 5y ago
I've heard something similar in response to Copilot in another thread (something like offering a sum of money to Github if they train their model exclusively on the Windows NT source code). But I think the legal theory here is that Copilot is trained on many thousands of sources. If Copilot was trained on a single source, or even a small handful of sources, the derivative work claim becomes much stronger. When trained on many sources, it becomes much harder to claim that its a derivative of another work.
Take for example a human. If I studied a bunch of different open-source projects, learned techniques from them, and implemented them in my own projects, is that a derivative work? Probably not. But if I were to reverse engineer Windows and implement the techniques I saw in ReactOS, that's where it seems issues start to arise.
- pedrocr 5y agoSo I just need to decompile Oracle's database and a few other commercial products as well and I'm good? Is Microsoft legal happy if I do Windows+Office+OpenSourceCorpus? I'd take that statement as well. Or even if they just do that themselves and train Copilot on their internal source code just as they do with the public open-source corpus. That would be a strong statement as well.
- klyrs 5y ago> When trained on many sources, it becomes much harder to claim that its a derivative of another work. Sounds good in theory, until it starts producing snippets verbatim from uniquely-identifiable sources.
- breischl 5y ago>If I studied a bunch of different open-source projects, learned techniques from them, and implemented them in my own projects, is that a derivative work? Probably not. That's pretty unclear actually. If it's quite close to the original work, it is derivative. Even though you have probably been "trained" on quite a few different codebases over the years. Hence the existence of clean-room implementations, wherein the people building a new implementation have never seen the original. Also, given that code that has been passed through a biological network (ie, brain) can constitute infringement, it seems obvious that code passed through a mechanical one could too. Maybe not in every case, but it certainly seems plausible.
- heavyset_go 5y agoYou're taking the machine learning metaphor literally. Training an ML model is not the same thing as a human being learning off of material. A human being can understand abstract concepts and reason about them based on material they learn from. An ML model is a statistical model that is closer to compilation or lossy encoding or compression. Often, ML models can encode their training data verbatim in the model itself, which is exactly what happened with Copilot and this example[1]. [1] https://twitter.com/mitsuhiko/status/1410886329924194309 https://twitter.com/mitsuhiko/status/1410886329924194309
- erhk 5y agoWell i could certainly construct a data set wjere windows NT is an atomic outlier and muddy the water with arbitrary inputs to satisy your irrelevant requirement. Perhaps ill jam some pictures of cows in, or any nubmber or animal photos. Maybe even some classical literature. Hell, maybe i even just jam a shit ton of Javascript in. Thats code right?