7 ms·
> Rather than conceal the licenses of the underlying open-source code it relies on, it could in principle keep this information attached to each chunk of co
by mcbuilder 4y ago
> Rather than conceal the licenses of the underlying open-source code it relies on, it could in principle keep this information attached to each chunk of code as it wends its way through the model.
I don't know if the author understands how these transformer models work, but this would a impossible task in Byzantine complexity. The way these models work is by outputting a probably distribution of likely "token" embeddings given an input prompt. The output involves basically inverting a word2vec (probably beefed up with programming language keywords and other bells and plausible search techniques I don't have access to the details of).
This model was of course trained with real code, but you can't attach to the output any meaningful information from the gradient you get from the sample. It's a very messy computation to even think to write down (attaching a percentage that 1 training example affected 1 particular output), much less come up with a simple interpretation of.
- mnd999 4y agoThe model itself is a derivative work of GPLed code and thus should itself be open sourced. Granted by the time it’s spat out it’s probably an unreadable binary blob but it would be nice to have access to it for GPLed projects.
- mbreese 4y agoI’m not sure about the vitality of GPL here. Assuming the generated code does not exist and is not part of the training set (these are two giant caveats), how would this be different from me reading GPL code to learn how to do something, and then rewriting it in my own words? If I have not copied, but learned from GPL code, is my new code GPL? No, it’s not a derived work in terms of copyright. Am I standing on the shoulders of giants? Yes. But so long as that concept is expressed in a novel way (not just changing variable names), then it isn’t a derivative work. I haven’t worked with copilot at all to know how verbatim the results are, but it is theoretically possible to train a model with GPL code and not get GPL code generated out the other side. (Again, here be dragons).
- jen20 4y ago> how would this be different from me reading GPL code to learn how to do something, and then rewriting it in my own words? No different - that is also not allowed. The idea of clean room reverse engineering exists exactly for this reason.
- nathanielarmer 4y agoDo you have a source for that claim? My understanding is the GPL based on copyright - and that copyright only protects a specific expression, not a general idea or concept. To suggest that if something is copyrighted, humans cannot learn from it and generate thier own material seems to be absurd.
- withinboredom 4y agoUnless the model is truly coming up with something novel, there is something in it's training set that is truly similar, if not precisely the output. It should say so when that is the case, and should also say whether or not the output is novel. I'm sure GitHub could provide an API for searching code snippets, if they don't already.
- moyix 4y agoThey're currently doing this automatically for catching exact matches, and you can turn on a setting that will suppress any suggestion that was found verbatim in the training data. But of course this wouldn't catch copies where variable names have been changed, etc. One thing that I think would be really interesting (but hideously computationally expensive) is to compute an embedding of each chunk of the training data and then query for the k nearest neighbors when Copilot generates some code, so you can see what the closest snippets in the training data are and evaluate for yourself if they're too similar.
- mcbuilder 4y agoYes, but Co-Pilot is more like a system that has spent a lot of time learning from open source code but it has the logical reasoning ability of that code of the 12-year old who just learned about the problem yesterday that the author mentioned. But will it be capable of generating a copyright infringement? Let's imagine I want to implement an algorithm to find the shortest path between nodes in a graph. Well I remember from my MSc that you probably need to implement it with Dijkstra's algorithm, and I've implemented this a few times over the years in kata and leetcode. Now let's say I'm programming Rust, which I don't really know the syntax for so I go read some GPL code while browsing the net for syntax. Do I need to state that my implementation of Dijkstra's algorithm is 10% GPL, because I spent a few minutes reading syntax of a GPL file? What if someone else's copyrighted code looks 90% similar to mine, because hell there's only so many ways to implement it, can I get sued because there is 90% overlap and I might have looked at this code? These systems aren't even capable yet of generating more than a a few functions, much less a coherent library. Even then, as they generalize more and gain more expressive power they will be less likely to copy in the training data to the output now, a possibility that I consider quite remote even given a relatively primitive Co-Pilot.
- gus_massa 4y agoThere is a problem because some licenses require attribution, but ignoring that... You can make a model that is trained only with BSD and MIT code, and IIRC/IIUC the result can be used with any license, including proprietary code. You can make a second model that is trained only with BSD, MIT and GPL2 code (and perhaps GPL2+), and IIRC/IIUC the result can be only with GPL2 code. You can make a third model that is trained only with BSD, MIT and GPL2+ and GPL3 code, and IIRC/IIUC the result can be only with GPL3 code. AGPL, Apache, WTFPL, ... Just add them to the correct model or create a new model for them.
- q-big 4y ago> You can make a model that is trained only with BSD and MIT code, and IIRC/IIUC the result can be used with any license, including proprietary code. BSD and MIT license still require attribution of the used source code (there exists a MIT No Attribution License, though: https://en.wikipedia.org/w/index.php?title=MIT_License&oldid=1091025780#MIT_No_Attribution_License https://en.wikipedia.org/w/index.php?title=MIT_License&oldid...).
- wolpoli 4y agoGreat. Now developers will need to check license compatibility before picking the copilot model. In reality through, we'll all just end up using the MIT No Attribution License copilot.
- q-big 4y agoDevelopers already have to check license compatibility. If copilot implemented such a feature, this would just represent the status quo.
- xigoi 4y ago> In reality through, we'll all just end up using the MIT No Attribution License copilot. Great, so GPL code would be left alone. I don't see how that's a bad thing.
- mbreese 4y ago
- Brian_K_White 4y agoAs much as I hate the entire concept, I would have a hard time articulating a substantive difference between this description of the ai mixing together bits of stuff it saw and what I do myself when I'm writing something that I fully describe as mine.
- amelius 4y agoThe difference is that the AI is doing it on a much larger scale. Analogy: the law has no problem with you memorizing the license plates of cars you see as they pass your street; however, building an automated system for recording and storing license plates in a database is a different story.
- alpaca128 4y agoYou probably don't sell code snippets that you collected from various sources you can't remember without knowing whether you're even allowed to sell some of them. That's what Copilot does with extra steps. I don't see why it should be different just because Microsoft mixed the code snippets in a blender to obscure what they're doing. It's code laundering.
- giovannibonetti 4y ago"code laundering" is genious!
- disconcision 4y agojust trying to make the comparison more precise: using copilot is like subcontracting to a company who you know has no institutional qualms about their employees liberally copy/pasting from a known corpus of github repos. you know these employees don't always literally copy/paste, but all they do every day is read the corpus, so their code is always going to be basically derivative. you can optionally ask that the company avoids matches to existing code, which you know means that before sending you the code, the boss will fuzzy search their code against the corpus, asking their employees to write it again if a match is found. so given that this is the way the company works, and that you know this is the way they work, your task is to decide what your ethical and legal liabilities are.
- omegalulw 4y agoThere are much simpler solutions than that. Simply invest or reuse plagiarism models and tag output snippets to source repositories. You can either exclude this code or let the users know it's licenced code.
- amelius 4y agoThat's no excuse. Attach all licenses if necessary.
- charcircuit 4y agoIt's not impossible. Amazon's CodeWhisperer is already able to tell you the license of code if what it generates is close to existing code.
- jamal-kumar 4y agoI wasn't aware that Amazon was developing something similar, how does it compare in terms of usefulness? I was finding the copilot trial was pretty good at reading something like an adjacent CSV file and building out a struct with proper data types from the data in it but beyond writing CSV->DB migration stuff I didn't really use it for a lot
- groovybits 4y agoPreviously: https://news.ycombinator.com/item?id=31851373 https://news.ycombinator.com/item?id=31851373
- treesprite82 4y agoWhat's described as impossible is "[keeping the license] attached to each chunk of code as it wends its way through the model". Checking generated code for similarity against training set is possible, and is now done by both Copilot and CodeWhisperer. But it'll include code that just happens to be similar, even if that code had no influence on what the model generated.
- ninjin 4y agoIt does not matter whether Matthew understands how transformer models work. He does understand that legally this is looking messy – or “foggy”, in his words. As a person that teaches and conducts research with these models, there is no reason why a tool like Copilot has to be a purely parametric, generative transformer like GPT-3. Where attribution, as you describe it, is nearly impossible. For example, if the model was to use a retriever component to obtain specific pieces of concrete code from a database (where the licensing for each piece is known) conditioned on the original source code context; then generate its output based on this, the context in your source code, and pre-trained parameters, then it could theoretically at least satisfy Butterick’s request and probably be more akin to how a human programmer operates. This still does not rule out possible legal issues entirely as the pre-trained parameters are opaque, but it certainly makes it less problematic. Alternatively, there is active research on attributing a given output to a transformer’s training data. But it is still very early days and frankly it is very “foggy” as to what degree this can be done. Lastly, I have seen a bunch of comments elsewhere alluding to it being somehow sufficient to reduce the level of verbatim copying. Sadly, it will not, as even if you for example replace the variable names it is still a copyright violation. Just like if you manipulate the RGB space slightly for an image. Determining fair use, etc without a license is something that today can only be done in court, regardless of how we of a heavier technical disposition may feel about it. After all, that is exactly why we have explicit licenses in the first place!
- flohofwoe 4y agoWhy not simply split the input training data into a separate "license sets", and then train on them separately? Instead of one Copilot you'd have a "GPL Copilot", an "MIT Copilot", and "BSD Copilot" and so on and so forth, at least this would simplify the burden that's currently put on the user. Attribution would still be a problem though. How to attribute generated code that's a random mishmash from thousands of inputs? The only solution can be to attribute all of them by doing some sort of reverse match. This "IP scanning" must clearly also be provided by Copilot and not be offloaded to the user. TBH I would have expected that the "evil geniuses" who came up with Copilot would have thought about such obvious pitfalls beforehand. It's not much of an achievement to innovate by ignoring the law.