5 ms·
Well, I have a hard time drawing a line between GitHub Copilot and a compression algorithm. If you can reproduce a verbatim copy of Quake source code after tak
by treffer 5y ago
Well, I have a hard time drawing a line between GitHub Copilot and a compression algorithm.
If you can reproduce a verbatim copy of Quake source code after taking that source code as input before then that's compression. A really fancy, but still.
And given that it reproduces the source code: it has to hold that somewhere.
It would be very interesting if someone could reproduce the Quake example with AGPL code, then request the whole model + code because it clearly contains the AGPL code in some encoded form.
- abriosi 5y agoSome purists may say learning is compressing
- Syzygies 5y agoYes! In every form, lossy compression is distilling meaningful information from noise. This is a great legal question as it concerns our use of machine agents. We can learn from copyrighted literature or code that we read. Why can't our agents?
- AlotOfReading 5y agoBecause the process is different. You and any computer agent are allowed to learn the functional, non-copyrightable elements of fast inverse sqrt. When you need that functionality, you can write code that implements your understanding of those non-copyrightable elements and gain copyright over the resulting creative expression. What you can't do is copy all of the creative expression in the original (such as comments) without complying with the terms of the license. Moreover, reproducing the magic constants is a strong indication that your process didn't independently derive your code because the constants used in the original are unique and non-optimal.
- anticensor 5y agoI should include a term in my licenses that licensees explicitly waive their rights to fair use and/or fair dealing.
- Zababa 5y ago> We can learn from copyrighted literature or code that we read. Not everywhere. Emulators communities often prohibit people from contributing if they've read the original code to protect themselves from copyright claims.
- burnte 5y agoIf your model can't reproduce the Quake source without my input, you haven't really compressed it, especially if the dataset to recreate it is larger than the original. If I have to tell the program exactly what I want in detail to get the Quake source, that's more of a storage database. If I have to guide it intently to get it to output the Quake source, I'm heavily guiding it.
- dleslie 5y agoAll decompression requires input: the compressed artifact. In this case, the compressed artifact is the semantic queues necessary to extract the Quake inverse square root function.
- swiftcoder 5y ago> especially if the dataset to recreate it is larger than the original Many types of compression produce a compressed file larger than the original for input data that is not easily compressed. Just because a compressor is bad at compressing (some) inputs, doesn't exclude it from being a compression algorithm.
- dleslie 5y agoThis is an interesting perspective; it does, indeed, seem like Copilot is a lossy compression algorithm wrapped in a semantic search interface.
- Spivak 5y agoI mean that’s essentially what all ML is if you want to think about it that way. Training is the process of creating a space where searching for the right thing within it gives you the answer to some problem you have.
- vharuck 5y agoA compressed file containing Quake's source code would be covered by the copyright on Quake's source code. The compression algorithm would not. The algorithm cannot produce the plain-text copyrighted material without the compressed copyrighted material. Copilot has the ability to produce Quake's source code nearly by itself. And it's a work (not a person), so it can be seen as a derived work. Like a compression algorithm that sometimes tacks on the first paragraph of "50 Shades of Grey" at the end of files. I'm not a lawyer, but that's my opinion (admittedly, my opinion is softening each day). Plus, the purpose of the tool is to create code for inclusion in projects somebody will hold a copyright over, and they likely won't be the original authors. So it's output should be held to a higher standard than a compression algorithm or keyboard.
- leereeves 5y ago> Copilot has the ability to produce Quake's source code nearly by itself. Was it fed the Quake source code while training? Then it's not producing that code, it's just reproducing it, like a fancy (but imperfect) copy machine. I'm not sure it's accurate to say that the training source code is "compressed" in the parameters of the model, but certainly some approximation of the training source code is stored in the parameters.
- treffer 5y agoIt is probably a stretch, but I think less of a stretch than saying "it just a machine that learned to code and randomly reproduced these 10+ lines of code". That has IMHO a probability of 0. So if I rule that out, where does it end up? What if we put this as the grown up ML brother of the chain of LZW, PPM, dictionary assisted compression (e.g. zstd) and various attempts at using neural networks for compression? I would not want to judge this - that's why I put up the AGPL idea. Or even unlicensed code. It would be a very interesting case to watch.
- madsbuch 5y ago> A compressed file containing Quake's source code would be covered by the copyright on Quake's source code. The compression algorithm would not. What? Where does the distinction between data and algorithm go with compression algorithms? In its most abstract form a compression algorithm is function `{0, 1}^n -> {0, 1}^m` such that n < m and the output string is the result of something previously encoded. Why can't the input string be the seed used to make the machine learnt model generate the Quake source code?
- elliekelly 5y agoI know absolutely nothing about IP and even less about compression but aren’t compression algorithms usually run on copyright protected material with the consent of the rights holder or authorized licensee?
- belorn 5y agoYou would not need to produce a perfect copy. A fansub of a movie is considered an derivative of the movie, while being a far cry from being an actually copy of the movie. As a subtitle is to a movie, the quake "output" might be much smaller than quake itself.