5 ms·
GitHub Copilot may steer Microsoft into a copyright lawsuit
- PeterStuer 4y agoIt seems reasonable that if a tool is emitting non-trivial code fragments originating from licensed code that emission should comply with the origin code license. Whether or not there is an 'AI' in between seems irrelevant.
- williamcotton 4y agoIs the boilerplate for epoll across ten threads listening for incoming HTTP requests non-trivial? It sure is! Should the first person who wrote that code in C get to claim ownership over anything that is substantially similar? According to the courts in the US, the answer is “No”, and thank goodness for that!
- zugi 4y agoFrom most articles I've read on this topic, I haven't been sympathetic to the claims that GitHub Copilot violates copyright. After all, human intelligence learns by reading books and reviewing code and then producing novel code. The results are influenced by what the human inteligence has seen, and there's nothing wrong with that - all of the world's knowledge builds from prior knowledge. Trying to lock all of that away via copyright claims is not the intention of copyright law. Copyright law applies to explicit copying. But a recent article showed some specific examples where a 15-line code chunk was clearly copied and pasted from someone's github code, with a couple variables renamed and identical comments but with '//' instead of '/* ... */'. If this were homework, it would be reported as copied. So I suspect there's not as much "intelligence" in this system as advertised.
- fhd2 4y agoWell, LLMs, like other ML models, are just algorithms that were automatically derived from test data. Calling that "intelligence" seems fine for marketing, but I doubt it flies in a court of law. IANAL though.
- williamcotton 4y agoRegardless of Copilot, everyone should inform themselves on what is and isn’t covered by copyright. You can copy and paste verbatim anything that is utilitarian in nature. This includes basically all optimized algorithms, implementations of mathematical algorithms, boilerplate. It does not include class structures, code organization, comments, or any other “arbitrary” aspects of code. Rule of thumb: Useful? Not copyright. Useless? Copyright.
- prewett 4y agoThis is completely incorrect, at least for the US. As I understand it (not a lawyer, but did take an IP law class a number of years ago), all creative works get copyright by default. A list requiring no creativity, such as a complete alphabetical list of names and phone numbers, has no copyright, or at least "thin" copyright. However, lists that show creativity ("most important phone numbers in the US") have copyright even though the facts they contain are not copyrighted. "Optimized algorithms" are most definitely creative. Even a straightforward implementation of an algorithm has creativity, if only in naming, the use of `for` vs `foreach` vs `while`, and any machine/OS/language specific details. Class structures and code organization is definitely copyrighted: Google's (successful) defense in Google v. Oracle was not that APIs (class structures and organization) are not copyrightable, but that their use of the APIs was fair use. Comments are even more copyrightable, as they are creative writing describing the code. Unless your comments are mechanically generated, I guess, then the phone book ruling could hold, but in that case you should improve your comments. A more accurate rule of thumb is that everything is copyrighted.
- williamcotton 4y ago"In computer programs, concerns for efficiency may limit the possible ways to achieve a particular function, making a particular expression necessary to achieving the idea. In this case, the expression is not protected by copyright." https://en.wikipedia.org/wiki/Abstraction-Filtration-Comparison_test https://en.wikipedia.org/wiki/Abstraction-Filtration-Compari... This allows for verbatim copies if they are utilitarian in nature! As for why we should allow verbatim copies of utilitarian features... First, let's preface this with the substantial similarity of the structure, sequence and organization as established in Whelan v. Jaslow which amongst other things says that you cannot merely change the variable names if the expressive structure of the code remains the same. Now let's imagine 10,000 software developers who all implement Dijkstra's algorithm in C and then run it through clang-format. Aside from variable names, isn't it safe to assume that many of the implementations are going to be exactly the same? Now, this doesn’t mean that GitHub is not in violation of other copyright claims, such as clearly expressive parts like comments and more!
- jlawer 4y agoHow long until half the subjects in an IT degree are provided by the faculty of law?
- dang 4y agoRelated: GitHub Copilot investigation - https://news.ycombinator.com/item?id=33240341 https://news.ycombinator.com/item?id=33240341 - Oct 2022 (1212 comments) Copilot under fire as dev claims it emits 'large chunks of my copyrighted code' - https://news.ycombinator.com/item?id=33272786 https://news.ycombinator.com/item?id=33272786 - Oct 2022 (243 comments) GitHub Copilot, with “public code” blocked, emits my copyrighted code - https://news.ycombinator.com/item?id=33226515 https://news.ycombinator.com/item?id=33226515 - Oct 2022 (773 comments)
- williamcotton 4y agoI welcome every new discussion because it gives me a chance to correct some very wrong assumptions about how copyright works, even though it is incredibly frustrating to be silently ignored or outright told I’m “confidently incorrect”…