52 ms·
GitHub Copilot investigation
- Kiro 4y agoAbolish all copyright. We're all happily pirating movies and music but code is for some reason sacred.
- macintux 4y ago> We're all happily pirating movies and music but code is for some reason sacred. Speak for yourself. I pay multiple streaming services, music and video, because I prefer creators be able to eat.
- bobdvb 4y agoAs someone who works at a streaming service: thank you. We are people, we have ambitions and families. We're not just a faceless corporation.
- bigiain 4y agoFWIW, not all of us "happily pirate movies and music". I want there to be more good music and movies. I want to support artists who create entertainment I enjoy. I go out of my way to buy physical copies of music from artists, wherever possible from the merch table at their shows or from their own websites. I pay to go see movies on the big screen (partly because I like the big screen cinema experience, but also because I understand "opening week revenue" is a key performance indicator for the success of a movie). I thing copyright is old, outdated, and probably not really fit for purpose for forms of creative work invented in the last 50 years. But I also thing creative workers need to get paid for their effort (juist the same as software developers), and absent a FAANG-style set for companies employing teams of songwriters, musicians, authors, and the like - on FAANG-style salaries, copyright seems to be the option that is working (however badly). I'll join your "abolish all copyright" crusade as soon as there's an alternative that at least likely to possibly work as well (or better) than the system copyright allows. Just abolishing copyright and erasing the publishing/music/movie/art industries without a transition plan isn't a thing I can support. (At least a transition the artists/editors/producers/writers/etc. I'll admit there's a large chunk of management and legal in the fairly abusive parts of the music industry I wouldn't shed a tear if they all became homeless and destitute overnight...)
- suyash 4y agoThis investigation should not stop at GitHub Co-Pilot, large language models currently that are trained on huge amount of data should also be investigated as I'm sure there are lot's of problems to be found there.
- birthday 4y agoI wonder if it emits stable diffusion samples? ;-)
- ipaddr 4y agoDoesn't really explain how co-pilot is stealing your community. I've used co-pilot and it works great until you are past boilerplate than it falls apart.
- ja3k 4y ago> Over time, this process will starve these communities. User attention and engagement will be shifted into the walled garden of Copilot and away from the open-source projects themselves The author seems to be implying that since Copilot can reproduce the code of open source repository X in certain scenarios there'd be no reason for programmers to learn/use/engage with repository X. But this is silly. Maybe some open source repositories could be tab completed with a little prompting but people will presumably choose to add a dependency instead of tab completing the code of express or something.
- cercatrova 4y agoIt also doesn't make any sense. Copilot suggesting to me the signature of a function from some library is not actually the same as executing that library. That library still needs to be downloaded onto my computer to be executed. And who will write new features to a library if not for the people who are interested in that?
- remram 4y agoOpen-source licenses have terms that apply to the source code, not "execution", so I really don't understand your point?
- cowtools 4y agoIt strips GPL, or any license.
- bigiain 4y agoIt strips the (mandatory, in a lot of cases) licence text. But the licence still applies. (Or I guess more technically, the original authors copyright still applies, and the rights granted to use the work under the license as an exception to the strict limitation under copyright - do not apply...)
- comfypotato 4y ago[deleted]
- Thorentis 4y agoAll class actions are a mix of both
- authpor 4y agoI'm more worried about the status of freedom in software, open source feels like a mirage to divert the attention away from the original issues from the FSF.
- commitpizza 4y agoGreat, I hope it is tried in court. It should be. But unfortunately I have not a big hope that the courts will come to understand the issue well enough.
- otterley 4y agoMany courts, especially those in the Northern District of California (where a case would likely end up litigated), are very proficient and literate about software and copyright law. See Judge William Alsup’s cases if you want to see some examples that illustrate the court’s competence. And these judges frequently have technical consultants on staff to assist with technological issues.
- commitpizza 4y agoOk, well that sounds nice. I have little to no insight to how courts works in the states so I was talking about my experience of the court system at home :)
- echelon 4y agoBSD 5-Clause 1. Redistributions of source code must retain the above copyright notice, this list of conditions and the following disclaimer. 2. Redistributions in binary form must reproduce the above copyright notice, this list of conditions and the following disclaimer in the documentation and/or other materials provided with the distribution. 3. All advertising materials mentioning features or use of this software must display the following acknowledgement: This product includes software developed by the organization. 4. Neither the name of the copyright holder nor the names of its contributors may be used to endorse or promote products derived from this software without specific prior written permission. 5. Use of this source code for the research or training of machine learning models is permitted.
- jen20 4y agoThe issue with copilot is that it is not respecting clause 1.
- honkler 4y agoso what? You can put on a cowboy hat and larp as one, but that does not mean everyone else around you have to take it seriously. Same way with these made up licenses. If it's on the internet, it belongs to all. Or else keep it with yourself.
- jen20 4y agoSo, if you are not compliant with the posted license, you are in violation of copyright. This is no different whether you are honkler or Microsoft (other than in how vigorously or not someone may enforce it). If you really believe that anything posted to the internet ‘belongs to all’ then I don’t know what to tell you other than you live in a fantasy land where Oracle Corporation does not exist. We might all prefer it if things were that way, but they simply aren’t, and that’s just tough.
- wilg 4y agoOnly if you agree it is "redistributing" anything! Small enough pieces of code can't be copyrighted. No one would support an argument that I violated copyright by using the code "else if {" from some GPL library. So the question becomes what is the minimal unit of copyrightable code? What if you wrote a nice big function exactly (or almost exactly) the same way as someone else did? Whose copyright are you violating?
- cercatrova 4y agoDoes GitHub not have the right to view and train from your content when you agree to their Terms of Service and upload your code? People are conflating their open source license with the one they give GitHub when making a GitHub account, but they are two entirely separate and parallel licenses. The former is for other people to use your code, the latter is for GitHub to host your code. If you don't like it, you are free to host your code on your own servers. And anyway, as noted the other day about AI, it is often funny to see people not care about (or even enjoy) AI in other fields that they don't work in, but when it comes for their own field, they are suddenly very worried. See programmers on HN who argue for Stable Diffusion but against Copilot, and vice versa with artists on Twitter. As I commented then, it's an act of cowardice to think our own profession should be immune from AI while we enjoy the fruits of AI in other fields [0]: > Yes, many of us will turn into cowards when automation starts to touch our work, but that would not prove this sentiment incorrect - only that we're cowards. >> Dude. What the hell kind of anti-life philosophy are you subscribing to that calls "being unhappy about people trying to automate an entire field of human behavior" being a "coward". Geez. >>> Because automation is generally good, but making an exemption for specific cases of automation that personally inconvenience you is rooted is cowardice/selfishness. Similar to NIMBYism. We should want AI. That we then try to use outdated models like copyright to enforce holding back human progress is a true shame. In my view, so what if GitHub uses people's code for training data, we are all getting a better product because of that. [0] https://news.ycombinator.com/item?id=33226515#33228948 https://news.ycombinator.com/item?id=33226515#33228948
- sneak 4y agoNot all code on GitHub was uploaded by the copyright holder. The entire linux kernel is on GitHub and at least some of those copyright holders have never explicitly granted a license to GitHub beyond the GPL.
- wongarsu 4y agoThere are quite a few projects that didn't originate on Github. Some are mirrors of projects hosted elsewhere, some accept patches through other means, some include code that predates github. If get your linux kernel patch accepted by emailing it to the responsible maintainer, it will end up on https://github.com/torvalds/linux https://github.com/torvalds/linux. But you never agreed to the Github ToS, all you did was agree to publish it under the GPLv2. Linus agreed to the Github ToS, but he can't give away rights he doesn't have, so he can't be giving Github any rights to your patches that go beyond the GPL.
- williamcotton 4y agoCopyright only covers the expressive parts and not the utilitarian parts: https://en.wikipedia.org/wiki/Abstraction-Filtration-Comparison_test https://en.wikipedia.org/wiki/Abstraction-Filtration-Compari... https://en.wikipedia.org/wiki/Idea–expression_distinction https://en.wikipedia.org/wiki/Idea–expression_distinction https://h2o.law.harvard.edu/cases/5004 https://h2o.law.harvard.edu/cases/5004 Most of your code is probably not subject to copyright in the first place, regardless of license.
- jashmatthews 4y agoDoesn't Copilot reproduce the exact expression given the right prompt, though?
- gjsman-1000 4y agoEven if it does, it may not matter. For example, APIs are not copyrightable (see Google v Oracle), and if there is only one obvious efficient way to make something work, it does not follow that the user must be prohibited from using that way even if someone else did it first.
- williamcotton 4y ago“Expression” in the creative sense, as opposed to utilitarian in a functional sense. Copyright is meant to protect “useless” things like poetry and music.
- ok123456 4y agoSo does a random number generator.
- blibble 4y agoin countries where there is no fair use (most of the world outside the US) it seems quite likely copilot is willful, commercial scale copyright infringement
- wongarsu 4y agoFair use is unusually permissive in the US, but most countries have very complex copyright rules to allow e.g. a televised interview in a room with contemporary paintings, without getting permission from the copyright holders of those paintings. It'd certainly make for interesting cases.
- RunSet 4y ago> Microsoft characterizes the output of Copilot as a series of code "suggestions". Microsoft "does not claim any rights" in these suggestions. But neither does Microsoft make any guarantees about the correctness, security, or extenuating intellectual-property entanglements of the code so produced. Once you accept a Copilot suggestion, all that becomes your problem: > "You are responsible for ensuring the security and quality of your code. We recommend you take the same precautions when using code generated by GitHub Copilot that you would when using any code you didn’t write yourself. These precautions include rigorous testing, intellectual property scanning, and tracking for security vulnerabilities." I can't help but recall: "Linux is a cancer that attaches itself in an intellectual property sense to everything it touches." - Steve Ballmer, while CEO of Microsoft
- walrus01 4y ago> Steve Ballmer They have some really good blow in Redmond. If anybody could win an award for being coked up and sweaty on stage... https://www.youtube.com/watch?v=Vhh_GeBPOhs https://www.youtube.com/watch?v=Vhh_GeBPOhs
- Darkphibre 4y agoFun story: That was my first employee town hall, in 2000. I was concerned for the fellow (and so very glad when he left, Satya has been so so so much better for the company and morale). It was definitely an... interesting introduction to the company. See also this Domo video that turned it into a song. :) https://www.youtube.com/watch?v=f7ZDH45OAt8 https://www.youtube.com/watch?v=f7ZDH45OAt8
- walrus01 4y agoAt the time, I was doing Linux, OpenBSD and FreeBSD stuff in Bellingham. The reaction from the local and regional non-Microsoft community was really like "Holy shit what is going on down there?!"
- nomel 4y ago
- pr337h4m 4y agoIt's tragically beautiful how the copyleft crowd is putting so much effort into drastically expanding the scope of copyright. "I used the copyright to destroy the copyright." That sort of plot never works in practice.
- vlunkr 4y ago> drastically expanding the scope of copyright. I think you need to explain that more. The problem (or at least one problem) being explored here is that by using any code from co-pilot, you are responsible for making sure the licensing is correct. You could unknowingly be using and modifying GPL-licensed code in your non-GPL project, which is a violation if you don't publish your modifications. We're not talking about expanding copyright, just protecting the existing copyright systems from being trampled by microsoft.
- deleted 4y ago[deleted]
- boomskats 4y agoEverything else aside, the design on this site is among the best I've ever seen. Amazing typography, great to read on a phone.
- chatterhead 4y agoCan you elaborate on what makes it so? Changing font sizes, boldness, lines etc...
- ceroxylon 4y agoFor me it was the 'magazine' style with proper breaks and editing, formatted for a vertical screen but still reads great on my laptop. The effort in the content, links and emphasis make it feel like journalism I would normally get paywalled on.
- elfatizer 4y agoHe wrote the book on it. https://practicaltypography.com/ https://practicaltypography.com/
- Kiro 4y agoI think it's very hard to skim for some reason.
- civilized 4y agoWith this site you see about 50-100 words on a large mobile screen. On HN you see 2-4x that. Also, the section breaks and headers and boxes lack obvious rhyme or reason. It scans a tiny bit like a classy version of Time Cube. You keep getting hit with different font sizes and font styles and lines and ribbons and colors and you're not quite sure why.
- holler 4y agoThat was my first thought as well. Perfect font sizing, clean & elegant design.
- qu4z-2 4y ago
- mmastrac 4y agoI think the test for whether an AI is infringing or not should be: Can this AI regurgitate the vast majority of the creative aspects of an original/novel piece of software with minimal prompting, to the point where the output code looks mostly and directly cloned to a reasonable person trained in the art?
- bigiain 4y ago_Maybe_ software is fundamentally different to other "creative works" which rely on copyright protection, but it's not immediately clear it is, and as far as I know it's certainly not a "special edge case" as defined in copyright law in general. So "Can this AI regurgitate the vast majority of the creative aspects of an original/novel piece of software" is not the test that, for example, the music industry uses when determining if a sample is infringing. The test there is "is a sample, however small, identifiable as part of a copyright work by a reasonable person trained in the art?" You can't own copyright in a composition of a single middle c note. But lawsuits have been won for copyright infringement of melodies of 2 bars (fewer than about 16 consecutinve notes). Men At Work lost a copyright case for the flute melody in Land Downunder which is the same as a 90 year old tune Kookaburra Sits In The Old Gunmtree https://www.claytonutz.com/knowledge/2010/february/men-at-work-go-down-under-in-kookaburra-copyright-claim https://www.claytonutz.com/knowledge/2010/february/men-at-wo... Whether that's done by a flute player or an AI, really doesn't make any difference as far as copyright law sees things. (Whether copyright law is a "good fit" for source code, and whether it makes sense to apply laws meant for books/literature/music/film to software is a different but very good question. I don't have much in the way of other ideas which take original author's efforts and potential rights to benefit from then though...)
- mmastrac 4y agoFWIW, in this case I was trying to feel out what I think is a good fit for copyright in general. I think the same test could be applied for books, music, art, etc.
- footlose_3815 4y ago[dead]
- jasone 4y agoI really don't care if my code gets ingested and regurgitated by Copilot, but it seems rather a stretch to imagine that this is fair use, in part because it separates me from the legal protections afforded by the licenses I released my software under. In my ideal world, Copilot would be legally viable, and releasing my software without restriction wouldn't be risky. As a long-time open source software developer, I have favored the 2-clause BSD and MIT licenses because they are the simplest licenses that provide me some liability protection. I would release code into the public domain if that didn't increase the likelihood of being sued, whether for liability, or for someone else claiming intellectual rights to code I actually wrote.
- xxs 4y agoI still release under cc0, being copied verbatim is of no concern. Yet, I don't think reproducing somebody else code is 'fair use'
- armchairhacker 4y agoOne issue I see with Copilot is that they get free access to all open-source data on GitHub, but using GitHub APIs to download the data yourself isn't possible (rate limiting). This is an unfair advantage. Copilot is not only making money off of open-source, they are making money off of open-source in a way others can't. I would love to see a lawsuit which requires GitHub to provide their full Copilot dataset.
- keithnz 4y agowhy use the API? why not just use git to get the code? All you need the API for is repository discovery
- rcoder 4y agoI'm gonna guess that Microsoft GitHub (tm) would shut you down pretty quickly if you tried to clone tens or hundreds of thousands of repos in a short window of time, b/c of course that's sketchy/abusive use of their infrastructure, right? But of course if the data is already sitting in object storage inside your cloud environment and all you have to do is run some MapReduce jobs to get at it... Hence: unfair, anticompetitive, intellectual-property-right-abusing behavior. Microsoft GitHub (tm) can prevent anyone else from running the kinds of analysis they do by simple "operational security", while running literally any kind of analysis, model training, etc. they want. Don't like it? But their commercial services and products so you can run Microsoft GitHub (tm) on your very own Microsoft Azure (tm) infrastructure, using Microsoft Visual Studio Code (tm) and Microsoft GitHub Codespaces (tm) so work on _your_ code privately. Best of all, you can still still take advantage of the huge library of "free" code offered by Microsoft GitHub Copilot (tm) to ensure your private, proprietary codebase still has all of the advantages of Open Source Software, brought to you exclusively by the Microsoft GitHub Platform (tm).
- ketralnis 4y agoI don't understand. Your favourite boba joint can email every one of their customers a coupon. That's "unfair" to the other boba joints without access to their mailing list too, right? You're just describing a regular old competitive advantage
- bigiain 4y agoThis is a bit off-topic, but I wonder if there are people/teams right now creating git repos, doing the source code equivalent of "SEO" on it, and embedding backdoors in stupidly overoptimized for the training process code? I wonder when we'll hear about the first big hack that gets traced back to production code pushed live after CoPilot "suggested" eval(base64decode({webshell}))
- melony 4y agoIf that code managed to hit production, then the problem is with the management and engineering leadership, not Copilot.
- xdfgh1112 4y agoIt's best to treat copilot like an eager intern who can churn out boring code for you. It still needs to be reviewed.
- bigiain 4y agoAnd none of us here have ever worked under incompetent management or engineering leadership... (Hell, I've _been_ that incompetent management and engineering leadership at various times over the last few decades...)
- wongarsu 4y agoThe less overt version of that is to figure out what mistakes copilot already makes (either things that are common in tutorials but not good in production, or things that are outdated, like hashing passwords with md5), and then systematically looking for software that includes such copilot suggestions.
- perrygeo 4y agoIs there a technique to scan for software that includes copilot suggestions? Or is this just theoretical? Sounds impossible given MS/GH's monopoly on access to the model input data.
- shpx 4y agoIt seems like there's no good license that places absolutely no restrictions or requirements on people using your code (such as attribution and respecting patent rights) worldwide. I want my code to be used the way people treated text in the old days. There's texts that have been re-written, added to and edited by thousands of people over the centuries and yet they don't come with thousands of pages of attribution notices because why would they?
- dwgebler 4y agoYou don't need to license it, you can just publish it with a declaration that as the author, you are releasing your work in to the public domain. However, in terms of licensing I believe MIT is the most permissive.
- deleted 4y ago[deleted]
- slondr 4y agoMIT still requires the license text to be included with the source. Copilot, if it is not fair use, violates the license of MIT code it re-emits.
- dwgebler 4y agoOkay, so just take "The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software" out of MIT and call it the Do Whatever You Want Licence. You're not obliged as an author / copyright holder to impose restrictions on people using your work if you don't want to. Whether Copilot is breaching MIT depends on what constitutes a substantial portion, which I am not qualified to rule on.
- yjftsjthsd-h 4y agoCC0 or unlicense, probably. Although I concede that it's hard to do in all jurisdictions globally.
- 4y ago
- esskay 4y agoI do have to wonder if Copilot will last. It's going to become a legal minefield and I can't imagine for a second that Micrsoft will want to be in the crosshair for another antitrust case.
- scarface74 4y agoYes because it turned out so badly the last time. Microsoft went from being one of the three most valuable companies in the US in 2000 to being one of the three most valuable companies in 2022. Also back then, Microsoft had 90%+ share of the PC operating system market and was bundling IE in its operating system. I’m glad the DOJ forced MS to change its ways.
- thrillgore 4y agoSo they turned around and started buying everyone else. GitHub, Nokia, Activision. They're back to their old shit.
- scarface74 4y agoGitHub doesn’t have a monopoly on “hosted git repositories”. By the time MS bought Nokia, it was already a has been in mobile and the acquisition was a total failure and the game market is competitive.
- solomatov 4y agoIMO, It will last. It's impossible to roll it back. IMO, if the lawsuit goes to a point where it's likely to be won by copyright owners, Open AI could do the following: - Use less sensitive code from big corps they have partnership with for training. I bet MS and other have plenty of such code. - Buy training rights from copyright owners of OSS projects. Many of them have SLAs which allow the owner do much more than the license allows. - Buy rights to train code, and collect generated code with Copilot from a large number of smaller software companies, likely with exclusions for some sensitive parts. MS has a lot of leverage here (discounts, partnerships, etc).
- Uhhrrr 4y agoThe example given is "sparse matrix transpose in the style of Tim Davis", but someone who wanted something with such specificity would be able to just take it from Github anyway, perhaps with a little more searching.
- lilyball 4y agoAnd would therefore have to follow the license of the code they took it from. That's exactly the point. Copilot is reproducing the same code but without the license.
- edsouza 4y agoPlease look up the license in the source code that Tim Davis points too, it just mentions it's LGPL, but doesn't include the full license text. And none of the C code mentions a license or Tim Davis. And if you dig further, the whole repo is mixed with BSD and LGPL "licensed" packages together. It's probably best that CoPilot does not suggest from code that does not have an explicit license stated. I think originally Tim Davis was complaining about the non public sources for suggestions which Github CoPilot ignored.
- deworms 4y agoA simple search on github reveals that those functions were reposted verbatim thousands of times, most people just copy and paste snippets of code they find useful, ignoring licenses. This highlights how all the power a license promises to hold is completely fictional. Any "in the style of Tim Davis" modifier only shows some kind of unwarranted self-importance complex on the part of the guy, thinking his style is widely known and distinctive (it's not). It's not the job of Copilot, the team that builds it, or the programmers that use it, to determine where the functions that were reposted thousands of times under all kinds of licenses originated. This is the same case as with copyrighted photos in newspapers, a paper prints a photo somebody allowed them to use, but then it turns out that person did not have the right to use it in the first place. Did not stop newspapers from printing photos at all. Here are the search terms: https://github.com/search?q=cs_transpose&type=Code https://github.com/search?q=cs_transpose&type=Code
- Barrin92 4y ago100% correct takes in the piece, this is just ridiculous >"Tim Davis gave numerous examples of large chunks of his code being copied verbatim by Copilot, including when he prompted Copilot with the comment / sparse matrix transpose in the style of Tim Davis /." Copilot regurgitates code and blatantly violates licenses, not even sure what there is to argue about. Not only does it seem straight up illegal and sideline open source communities, I think the next logical step of this is that people who want to avoid having their work vacuumed up and their rights violated simply to move to proprietary software, which would be a huge disaster for open source.
- elfatizer 4y agoThere are lots of comments arguing for or against Copilot on a value judgment, and having an opinion on it being ethical or legal, etc isn't going to be the same for everyone. But I think regardless of where you stand, there should be some sort of legal ruling to clarify the gray areas that Butterick breaks down.
- epolanski 4y agoBingo, I feel so uneasy at the thought we could risk a lawsuit because a colleague put unlicensed code in our repos.
- hn_throwaway_99 4y agoAgreed, but I also hate how so much of our substantive law basically has to be created by the courts because (a) many of our legislatures, especially at the federal level, have become more and more non-functional, and (b) IMO legislatures are especially bad at implementing technical legislation. I think there is a good, fundamental legal/societal question of how copyright should apply to AI output. I just don't think our existing copyright structures handle this question well. Note there is currently a very important case before the SCOTUS that is related to this issue, [1] where the original photographer of a Prince photo is suing Andy Warhol's estate for copyright infringement. The fundamental question is whether the Warhol series of painting are "transformative" enough of the original photo. While there are always gray lines on what "transformative" means, if there is any chance that Warhol's painting are legal and not infringing, I don't see how Copilot could be in the wrong. Copilot's output, even if it contains a substantial amount of the original source, appears to me much more "transformative" than the Warhol paintings are compared to the original photo. 1. https://www.npr.org/2022/10/12/1127508725/prince-andy-warhol-supreme-court-copyright https://www.npr.org/2022/10/12/1127508725/prince-andy-warhol...
- CoastalCoder 4y agoI agree. Any law that's only clear after a court ruling is, de facto, an ex post facto law. Disgusting.
- beefman 4y agoOpen source? They used everything on github with no regard for license, which would have included plenty of code under conventional copyright. Microsoft is now profiting from that code.
- version_five 4y agoNew tech creates winners and losers, and losers inevitably complain. See looms, VHS, Napster, etc. The more of this complaining I see, the more it falls flat. The only interesting thing is which side different communities end up being on. To be fair, record companies were not in the least bit sympathetic. Open source contributors are easier to identify with, though imo it doesn't actually make their concerns more valid
- jamiek88 4y agoWhat does this comment even mean? I cannot parse what you are suggesting.
- typon 4y agoHe's saying you're a luddite for not wanting Copilot to steal all your code and use it for furthering the grand vision of AI generated code (which supposedly represents the forward march of progress)
- bakugo 4y agoAs far as I can tell it's just a more convoluted way of saying "new good, old bad, only old people disagree"
- TAForObvReasons 4y agoIt's more nuanced. Copilot exists publicly, which also means some copilot-lite thing trained on a smaller subset of repos probably exists privately in many different places. It may not be as good today, but these private instances will improve over time. Since the demand for a copilot-like service exists, eventually a large VC-funded public instance will show up. In that lens, it is more sensible on the individual level to prepare for a world where copilot thrives than to put all of your eggs in the "ban copilot" basket.
- fswd 4y ago“If I had asked people what they wanted, they would have said faster horses.” Henry Ford
- 19h 4y agoI'd be rather saddened if Copilot was shut down or neutered because of a few vocal few protesting against it. It's been a massive productivity improvement to our senior devs, and I got so used to it that it's an annoyance when Copilot doesn't respond.
- randomsearch 4y agoI tried copilot and found it an excellent way to inject subtle bugs into my code. It always had a seemingly plausible guess, that was never correct, and coding turned into a guessing game trying to spot the bugs it had injected and hoping I’d found them all.
- jfghi 4y agoI’d claim that it’s more than a “few vocal” protestors. If the system is illegal, it needs to become legal or disappear. If I’m writing code for a query optimizer, the SQL Server solution isn’t going to magically show up.
- tpmoney 4y agoIs there any evidence that the PostgreSQL or MariaDB solution will though?
- Spivak 4y agoIt’s not illegal, it’s at worst a fancy code search tool that Github has the right to show you the results via the license you grant them when you upload and make public code on Github which is way stronger than other search engines like Sourcegraph have to show public code. It doesn’t mean you have the right to use any of the code it generates but Copilot itself isn’t illegal in any meaningful sense.
- noitpmeder 4y agoThis is definitely not true. When your license requires you bundle said license with any reproductions of the code, and Copilot spits out said code sans license, they are breaking the law.
- mdswanson 4y agoI've been trained on open source code, and there are likely many algorithms that I've internalized that are very similar to the "standard" way of performing an operation. Is there a reason why an AI being trained on the same open source code isn't a similar situation? I agree that wholesale pasting of code chunks is an issue, but that hasn't been my experience with Copilot. I'm not arguing for Copilot here...I'm genuinely curious why this would be considered any different.
- layer8 4y agoHumans can reasonably distinguish between when they are plagiarizing and when they are just applying their experience and knowledge. Copilot presumably isn’t able to make that distinction, and, arguably, so aren’t the consumers of Copilot’s output.
- bakugo 4y ago>Is there a reason why an AI being trained on the same open source code isn't a similar situation? You are a human. You know what's right or wrong. You know you can't just copy code 1:1 from public repositories without respecting their license. The AI doesn't know and doesn't care. It's a common problem with creative AIs that they will occasionally regurgitate near 1:1 copies of their training data, and I don't think it's an easy problem to solve. >I agree that wholesale pasting of code chunks is an issue, but that hasn't been my experience with Copilot. The article provides several examples of it happening. Just because it hasn't regularly happened to you doesn't mean it doesn't happen.
- mdswanson 4y agoI don't deny that there are examples that appear to be wholesale copying, and that is definitely an issue to be addressed. No doubt. What I don't understand is why the rest of the service (where it doesn't appear to be pasting existing code) is being maligned when it behaves like a more powerful version of autocomplete.
- renewiltord 4y agoHonestly, Github Copilot seems fine. It's just a tool that you're responsible for using responsibly. If I Google something, and copy and paste that, then Google is not responsible for my infringing. It's just "intelligent autocomplete".
- gjsman-1000 4y agoI wonder what Dictionary companies thought about Autocomplete...
- janef0421 4y agoGoogle search doesn't return random snippets of text without indicating their source.
- Scoundreller 4y ago> No match for domain "GITHUBCOPILOTCLASSACTIONLAWSUITSETTLEMENT.COM". > Last update of whois database: 2022-10-17T23:07:12Z <<< Just sayin'...
- orsenthil 4y agoOh man. I want to continue using co-pilot. It has improved my productivity and made me excited to do things that I previously felt like a chore. Also, programmers please do not hinder on other programmers work. If you do, someone higher up in the ladder with eat your cake at every opportunity.
- ummonk 4y agoYes, we’d all find our work easier if we could just steal other people’s work.
- dwgebler 4y agoFor me, as a (granted very minor) contributor to some open source, I couldn't care less about attribution. The ethos of open source is specifically about sharing stuff (probably for free) for the benefit of everyone, take a penny, leave a penny. It's more of an interesting question if Copilot is suggesting code verbatim from source-available rather than open source repos though.
- typon 4y agoThis Copilot saga is another good reminder of why nothing is free. Developers have been using Github for free for years - now the chickens have come home to roost. The copyright licenses are just a formality - a form of kayfabe. If you aren't hosting your own code (GNU style), you should assume Microsoft owns it, for all intents and purposes.
- ClassAndBurn 4y agoMy view is the copilot is not stealing open source code. It is learning from it just as a human reader would. People's disguste is based on the assimilation of what they thought was a human trait being machine derived from their work. The copilot service backed by an army of actual humans wouldn’t be a story at all. Nor would anyone be angry, if an individual offered coding skills as a service, and had gone through the exercise of learning great amount to open source software to do so. No open source license was written with this in mind. Because previously learning was something only humans could do and no one had issue with sharing that knowledge. Until licenses take machine learning use into account I see no problems with Copilot. Source cannot be open if you restrict any viewing of it.
- slondr 4y ago> People's disguste is based on the assimilation of what they thought was a human trait being machine derived from their work. No, people's disgust is with Microsoft violating their legal privileges. > The copilot service backed by an army of actual humans wouldn’t be a story at all. Correct, it would be an open-and-shut lawsuit.
- Sirened 4y agoYou aren't allowed to just read code and regurgitate it in order to claim it as your own. That is, just because you memorized this great new novel you read, it doesn't mean you can go and sit down and hammer it out and sell new copies. People go to great lengths to do this sort of things (see: clean room reverse engineering [1]) in order to try and wash themselves of liability. [1] https://en.wikipedia.org/wiki/Clean_room_design https://en.wikipedia.org/wiki/Clean_room_design
- deworms 4y agoIf you think most people pay any attention to licenses or respect them you better think again. Snippets get copied verbatim with no regard to their source all the time. Licenses have no power and are routinely ignored.
- 4y ago
- gjsman-1000 4y ago"[W]e inquired privately with Friedman and other Microsoft and GitHub representatives in June 2021, asking for solid legal references for GitHub’s public legal positions … They provided none." Well... DUH. Why would they? You want to possibly sue them. Why in the hell would they, or anyone, provide crucial evidence for your lawsuit before you've sued them, regardless of the case and circumstances? Of course they aren't going to provide evidence, because you are obviously going to then try to prove hypocrisy, whereas you might not have enough to go on if they don't talk. No corporate lawyer in their right mind would ever grant such a request. (Edit: You are quite literally asking what their legal strategy is going to be, before the lawsuit has occurred, and then trying to spin the refusal as a proof of guilt.) That's like claiming that an alleged drug dealer who didn't talk without a lawyer present is obviously a criminal, because if he wasn't he would have talked. What a nothing of a point.
- olliej 4y agoI feel "stealing your community" is lawyer hyperbole, but people also seem ok with what MS is doing with copilot, and I am not. If you think what copilot is doing is ok, and there is nothing wrong with it, I'd love it if you could go through this small thought exercise, and see if it impacts your view at all: Say you write a bunch of code, and release it under GPL. For the sake of argument imagine it is something complicated that you care about. Now say another person is trying to do what your code does, and they find your code, a say "excellent". They then copy and paste it into their project, and release their code under a BSD license instead. Would you consider this theft of your IP? The law certainly would, and I think most devs would as well. What would you say if they instead release "their" code as public domain? Now we'll go a bit further. Another person is trying to solve this problem in some commercial software. They find your code, copy-paste it into their project, then sell their software and don't release the source, or even acknowledge you. Would you consider _this_ theft? again the law would. Now, what if instead they found your code through the invalid BSD relicense? or the invalid public domain one? To me every one of these would be theft, and every one would be required to required to release the source of projects that made use of my GPL'd code, under the GPL. That is literally the whole point of the GPL. But let's imagine a different route. A person is writing some code and can't work out how to solve a problem, so they ask on StackOverflow. Now another person comes along and answer the question by copy-pasting from your project into SO. The first person says "yay!" and then copies that code, and we repeat the above scenarios. In an even more extreme case, imagine both of the above people work at the same large company - so neither knows or is even aware of the other - how does this impact what is going on? It's two people, but fundamentally the company is copying the original GPL code into SO, then copying it from SO into its proprietary code. I get that MS and GitHub try to position it as if copilot is "creating code", but it is simply doing a statistical code completion that is demonstrably happy to copy and paste from the original source into the recipient code. To my mind all it is doing is providing a mechanism to launder GPL (or whatever) code into your own without the license, by slapping "ML" and "AI" on the process and requiring more than 3 keys to be involved.
- thorncorona 4y ago> Another person is trying to solve this problem in some commercial software. They find your code, copy-paste it into their project, then sell their software and don't release the source, or even acknowledge you. Let's be honest, copy-and-pasting happens all the time. In software, in engineering, in marketing, in everything. Whether people acknowledge it or not. Everyone looks at Stack Overflow all the time. You do it. I do it. Nobody reads the licensing terms. We all produce software with reskinned and taped together functions. A collage is still unique, creative, work despite being glued together with other people's art. Most songwriters will write a song with a part like someone else's song. The products you buy at a store are rip-offs of someone else's product. Everybody stands on the shoulders of giants before them. Such is learning, such is life. Get over it.
- WrtCdEvrydy 4y agoI wonder what will happen when a company pays some overseas developers $50 for some code, they copy it from Copilot and it copies a bug from a US developer and that company gets hacked for $10 million. Will the lawsuit fall on the overseas developer, US developer or Github?
- belltaco 4y agoNo one? They'd probably stop doing business with the overseas developer and that's it.
- layer8 4y agoIt seems to me that in principle it should be possible to maintain attributions through the training process, so that Copilot outputs could come with a list of weighted sources, possibly discarding those that fall below a certain weight threshold. Doing so would likely be much more expensive in terms of the computational power needed for training, and probably also in the size of the model. But it would be great to actually be able to see what went into a specific Copilot output.
- swhalen 4y agoWhy can't Tim Davis (or another software author whose code is emitted verbatim by Copilot) demand that Microsoft take down Copilot, or at least the part of Copilot that contains his code? Microsoft is distributing his software without a license, isn't it?
- gjsman-1000 4y agoIt could potentially under "fair use," which completely overrides any and all copyright claims if the conduct is found to be, indeed, "fair use." Fair use in code is broader than just copying. For example, in Google v Oracle, APIs were found to be not copyrightable. Even if you copied the names of, say, 86,000 different functions in a proprietary library, you did not violate copyright. Then comes the second problem. Let's say there is a function, say, `AddTwoNumbers(int a, int b)`. Just because John Fitzgerald in 1999 implemented that as `return a + b;" doesn't mean you can't too. There's a degree where you can copy the code that made a function work, even if that code existed earlier. It's fuzzy but it is legally real. Finally, there is your third problem, which is that you risk a "safe harbor"-esque judgement. Just because YouTube has occasional copyright-violating content doesn't make YouTube illegal. Similarly, the person suing here risks a finding that GitHub Copilot is legal as long as any occasional long proprietary code regurgitations are removed as needed. If your code falls under the first two conditions, copyright be damned, license be damned, it's all irrelevant. See also Linux copying Unix.
- swhalen 4y agoThere must be a line past which copying is no longer "fair use", otherwise no copyrights in code would be enforceable at all. I suppose it is up to a court to decide, but in the Tim Davis thread from yesterday it looked to me like Copilot was emitting entire, nontrivial functions verbatim.
- Havoc 4y agoI find it hard to see a scenario where MS doesn’t get absolutely wrecked in court.
- endisneigh 4y agoCopilot is trained on and returns AGPL code verbatim. It’s game over. If these licenses are not enforced it defeats the entire purpose.
- solomatov 4y agoIt might be the case that it is fair use to train the model on public data, but the code which it produces is covered by AGPL. Github limits liability in its TOS. (I am not a lawyer).
- toastal 4y agoFolks really should take their GPL code to a platform with similar ideals and stop propping up Microsoft GitHub.
- broodbucket 4y agoThe vast majority of GPL violations are not enforced, because those who would want to enforce them are small and their opponents are big. For Copilot to blow up, it'd need to be licensed code from a big company demonstrably turning up in a product of a competitor, or some similar event.
- abigail95 4y agoThat's a problem of the licensers, not for Microsoft or the CoPilot users. If you released AGPL code but never intended to ever sue anyone. Why did you release it like that? If you did and if someone is able to use your code without any damage to you, without reputation loss, and via a way they have access to the innocent infringer defense after you overcome fair use, after you sue them. How is that game over?
- f1refly 4y agoSo suppose you go out and about and a Microsoft representative punches you in the face. Now, the Microsoft representative has a billion dollar corporation backing him, willing to defend him at all cost through every institution, while you're just John Doe who went on a trip. If you ever went on a hike but never intended to sue anyone. Why did you go out in the first place? If you did and someone is able to punch you in the face without any lasting damage, without reputation loss, and via a way they have access to the myriad legal defenses you couldn't come up with if you tried, after you sued them. How is that game over? Just because someone corporation is, because of its sheer size, over the law (as far as a John Doe is concerned anyways), does that make it a right? We could probably do away with laws at that point and just accept getting punched in the face by Microsoft whenever they feel like it as the new reality.
- rafaelero 4y agoCopilot is great and this is a waste of time.
- yjftsjthsd-h 4y agoBeing a useful tool doesn't make it legal.
- rafaelero 4y agoTechnical progress takes precedence over pitiful intelectual property discussions. If you don't believe that, I am not sure what you are doing in a community like this.
- qu4z-2 4y agoBeing a hacker?
- yjftsjthsd-h 4y ago> Technical progress takes precedence over pitiful intelectual property discussions. Let's pretend for a moment that your value judgement is reasonable and the advancement of technology should reign supreme over minor things like rule of law. Do you really think that letting people ignore copyright is always good for technical progress? Say, letting people use GPL code in proprietary code that they then refuse to share with others? Because that sounds questionable even if we agree with your casual disregard for the law. > If you don't believe that, I am not sure what you are doing in a community like this. Being interested in tech without being a fan of breaking the law and running roughshod over other people's work.
- rafaelero 4y ago> Do you really think that letting people ignore copyright is always good for technical progress? Only in very, very, very, very specific circumstances would I say it is not good. And they involve thinking about the counterfactual: "would this thing be created if there wasn't intellectual rights in place"? Code doesn't pass this test because people enjoy writing and sharing code. Pharmaceuticals, maybe.
- rockemsockem 4y agoWhat do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield where you must ensure you properly attribute your model outputs, only train on opt-in data, etc, etc. Surely no one really thinks that a court case against Microsoft/OpenAI (even if they lose) would stop CoPilot? Most of these complaints seem to be extremely emotional and cherry-picked. "People's legal rights are being violated!" (you definitely don't know that, no one knows that, the article is 100% right about that), "look I prompted CoPilot for this piece of code that I already knew about and it spit it right out" (that's not how it's going to be used in practice). It seems to me that the longer-term implications of the outcome of a lawsuit like this are far more interesting, yet almost all the comments I see are nitpicking and whining about how the world isn't the way they want it to be. I wish the conversations around generative AI could be...just better.
- downrightmike 4y agoIt creates a body of knowledge, everyone can use and can't be sued for since it would be the industry standard way to do things.
- nordsieck 4y ago> It creates a body of knowledge, everyone can use and can't be sued for since it would be the industry standard way to do things. That's just not true. If the "industry standard way to do things" is to violate other peoples' copyright, then everyone doing that absolutely can be sued. And while it's not clear if using these AI tools constitutes copyright infringement, it looks to me like there's at least a very strong case that could be made. And at up to $10,000 per copy (register your code with the copyright office if you care about this issue!), that starts to add up very quickly. Even for a company like Microsoft.
- gfynt 4y agoIf they continue that path, the future will be that OpenAI, Microsoft, Google etc. will pay larger and larger fines at least in the EU, until they are blocked entirely.
- verisimilitudes 4y agoNever forget this is how people who dare to reverse engineer Windows are treated: https://www.theregister.com/2019/07/03/reactos_windows_research_kernel_claim/ https://www.theregister.com/2019/07/03/reactos_windows_resea... https://marc.info/?l=ros-dev&m=118775346131654&w=2 https://marc.info/?l=ros-dev&m=118775346131654&w=2 I don't use Github, but fuckers upload my code there anyway. Copyright is evil, but only large corporations having copyright, even more than they already do, is even worse.
- rockemsockem 4y agoThis I feel like is one of the better points in the thread. The asymmetry that exists in copyright law where large corporations can enforce their copyright to the point of breaking the law themselves (YouTube's content ID is another non-legal, but still very impactful example) is absolute bullshit. Unfortunately I think that if training ML models on Internet-data is found not to be fair use then things will get harder for individuals training models and corporations will be barely inconvenienced as they can afford to pay for sources, make deals with other large institutions for data, etc.
- belltaco 4y agoThey're treated with an email to the mailing list? Was there C&D or lawsuit? You make it sound like the ReactOS devs were thrown into prison.
- nmilo 4y agoGood bye and good riddance. Even just the idea that GitHub should be allowed to train their proprietary AI on other people's work is insane. Much less distribute that AI in a paid package which lets you spit out other people's code verbatim. Anyone who supports open-source and the (ab)use of copyright law to create free works should be vehemently opposed to Copilot.
- wilsonnb3 4y ago> Even just the idea that GitHub should be allowed to train their proprietary AI on other people's work is insane. You explicitly agree to this when you upload code to GitHub. FOSS folks shouldn’t have sold their soul to the proprietary devil but they did and now they have to deal with it.
- remram 4y agoThe fact that it's the first major development to be started at GitHub after their acquisition by Microsoft is such a hit too. Way to spend their social capital. I can't imagine the money they made from Copilot subscriptions was worth it since companies have certainly stayed away from this...
- DannyBee 4y agoTo set people's expectations, it is likely to take a bunch of lawsuits and a bunch of cases here to get to anywhere useful. The problem with lawsuits on copyright is that they are rarely precedential. I get that what people see is the large cases that try to tackle big topics. But for every single one of those, there are probably 10x or 100x equally large case that did precisely none of that. This is particularly true of fair use, it is very fact specific. A court is much more likely to answer a very fact specific question about copilot, tied to the very specific facts of the case (IE how is this exact thing used/etc) than more broad, abstract questions. In fact, standard Article III courts in the US are literally not allowed to issue advisory opinions.
- invig 4y agoWhat's with the default to "if it's not explicitly legal, it must be illegal"? Imagine if every new piece of software your wrote had to be tested for legality because you don't know that it's explicitly legal. Oh there aren't laws for this new thing, so I guess you should challenge yourself all the way to the supreme court? I get the author not liking Copilot, but I don't see that GitHub/Microsoft have any kind of obligation to figure this out just because they're GitHub/Microsoft. If I as an individual had this obligation placed upon me I'd just never write any more code. Ultimately I think, like open source, Copilot and the tools that will follow advance human progress in novel ways. Software getting easier to make is a good thing. If you don't like this particular implementation of something helpful, feel free to start an open source alternative without challenging yourself in the supreme court.
- j2kun 4y agoThe licenses in question in this issue make it explicitly illegal for Copilot to reproduce their code.
- abigail95 4y agoI don't care what your license is, I'm going to use it and I'm going to claim fair use. What's explicitly illegal about this?
- notacoward 4y ago> What's explicitly illegal about this? The fact that the work is copyrighted, rights withheld in the absence of a license and limited with one. Also, you should know that willful infringement can carry 5x the statutory damages compared to accidental, and spurious claims of fair use would be distinctly unhelpful to your case. Just by posting that comment, you have probably compromised your position in any future copyright case you might be involved in, or you might even have invited one. I really recommend being more careful when anything legal is involved.
- gspencley 4y ago> What's with the default to "if it's not explicitly legal, it must be illegal"? That's not how I interpret what's happening. People who produce things have rights over their products. Be it artists, craftsmen, inventors, entrepreneurs or coders. There is a legitimate question here as to whether CoPilot has infringed upon those rights. I don't see it being about "making something illegal." I see it about answering a valid question as to whether CoPilot is liable for measurable damages caused to creators under existing laws.
- ok123456 4y agoMaybe I'm in the minority, but I think the prospect of someone autocompleting and getting a snippet that came from me, they found it useful, and are going to incorporate it is great. It means my thoughts and logic are shaping culture in a mimetic feedback loop.
- janef0421 4y agoThat's not the problem. If you want to license your work in a way that allows that, you are free to do so. The issue is that Microsoft did that with code that was published under licenses that either did not grant that right, or which explicitly forbade them from doing so.
- skydhash 4y agoAlso charging for it.
- frankjr 4y agoI wouldn't be surprised if Microsoft lawyers didn't like the word "github" in the domain name...
- Andys 4y agoIf you're against Copilot as developer, you're shooting yourself in the foot. Locking up code under non-permissive licenses stymies the pace of code development and increases the costs of progress dramatically. We all stand on the shoulders of others before us. Including the organisations that stand to benefit the most from aggressive licensing.
- jacooper 4y agoNot my problem. I put my time and my effort to open source a program for free and I want to make sure that my code creates an incentive to create more free software, by using a copyleft license.
- Andys 4y agoOh, so you're saying Copilot doesn't actually go far enough. It should not only give you code snippets but enforce particular licensing of the code it is helping you create?
- qu4z-2 4y agoIf it feeds you code made available under the GPL, it should probably tell you that your code needs to comply with the GPL to use that snippet, yes?
- klllai 4y agoCopy & paste isn't "standing on the shoulders of others". It is more like being an intestinal parasite.
- rockemsockem 4y agoNot trying to overly advocate for copy-pasting here, but isn't copy-pasting just the ugly child of calling a library function? If it's a blind copy-paste it's pretty much the same effect. Surely you wouldn't call using a library being an intestinal parasite?
- kazinator 4y ago> Arguably, Microsoft is creating a new walled garden that will inhibit programmers from discovering traditional open-source communities. This is extremely far fetched. User bases (let's avoid one of the four dirty C words) are organized around something which builds, executes and is documented, not searches for snippets.
- chiefalchemist 4y ago> Arguably, Microsoft is creating a new walled garden that will inhibit programmers from discovering traditional open-source communities. Or at the very least, remove any incentive to do so. The walled garden bit I get. But I'm lost making the leap to "remove any incentive to do so." Is Butterick suggesting that someone is going to put aside their code and do a deep dive on GitHub looking for a snippet that might not exist? I'm not trolling. I'm sincerely trying to grasp the argument being made.
- not2b 4y agoIt seems that Copilot could address this issue by searching for matches in its source repositories for the strings it generates, with appropriate criteria, and give the user a link describing the origin of the code, who wrote it, and what the license is for cases where a match length exceeds a threshold. So, you wouldn't just get the Quake fast integer square root routine, you'd get a pointer to the Quake repository and license info from which it came. A separate model could be trained up that would find the closest match in source code repositories. A user could then use Copilot safely, attribute code correctly, and avoid code with incompatible licenses. This would be a better approach than "shut it down".
- aetherspawn 4y agoI wonder if people realize that letting GitHub train Copilot on their open source contributions is effectively de-valuing your own time, which (if repeated at a larger scale) devalues your experience, and that eventually has the effect of reducing the correlation between your experience and your salary. For example, if an overseas firm can just as easily use Copilot as I can write original code (or use Copilot myself), why would any company hire me locally?
- quickthrower2 4y agoMaybe I will start writing open source code intended to trick copilot. Stuff that just about works in the given context, but will fail badly if copypasta'd into another program. If we all did that.
- pmarreck 4y agoWould an opt-in system fix this? Where your code is only learned from if you opt into using Copilot to help you develop faster?
- BoppreH 4y agoMost of these points can also be raised against DALL-E 2, but software has one extra thorn: patents. It's a common advice to not read software patents[1] because the infringement penalties are lower if you did so unwittingly, that is, by reinventing the patented technique yourself. I wonder if using Copilot doesn't push the penalties back again to wilful infringement. Or worse, patent trolls poisoning the training data with patented algorithms. [1]: https://queue.acm.org/detail.cfm?id=3489047 https://queue.acm.org/detail.cfm?id=3489047
- larsiusprime 4y agoHas Dall-E 2 yet reproduced 1:1 anything from its training set?
- thedorkknight 4y agoIdk about dalle but with stable diffusion if you type in "Mona Lisa" or "Van Gogh" you have to fight pretty hard with your prompt to NOT get identical reproductions of those respective works
- jacooper 4y agoThis would for sure have affects on anything related to Ai content generation.
- bo1024 4y agoThere are two issues -- (1) feeding copyrighted material in to an AI model, and (2) getting copyrighted material out. The latter is obviously a violation of copyright, full stop. The former, to me, is obviously not a violation. If it were, that would massively tilt the playing field in favor of large corporations. It would become very hard to independently train your own models. Philosophically, I go by the principle that if it's (il)legal to do yourself, then it should be (il)legal to do the same thing with an AI's assistance. The massive complicating factor is that nobody knows how to do (1) without also doing (2) as a side effect, because we don't understand how deep learning works well enough to control it.
- deathanatos 4y agoThis was a comment made to me in a previous, similar discussion, discussing case law around Google's use of copyrighted books in building a search engine: https://news.ycombinator.com/item?id=32654478 https://news.ycombinator.com/item?id=32654478 I'm not sure I completely agree w/ the comment (nor do I think it vindicates CoPilot), but I think it does provide insight into why CoPilot is violating copyright.
- bo1024 4y agoBut fair use, as I understand it, is only about outputs, not inputs. (Licenses can apply to inputs.) A copyright violation occurs when one produces a work that infringes copyright. In this case, Google made digital copies of the books and showed snippets of the books to website visitors. Fair use refers to cases where producing that work is nevertheless legal. The only analogy I can see is that copying the code internally to use in CoPilot training could be a violation of copyright (like how backing up your own MP3s is a violation of copyright?), but the licenses on these public repositories probably already allow that...
- jarsj 4y agoBut is it illegal for AI to provide the said assistance ? That, I believe, is the bigger question.
- 4y ago
- rwalle 4y agoI don't understand why GitHub decided to run the project this way. This is a great idea but they messed the whole thing up. They could have make it opt-in from the very beginning and ask people to waive their rights, and I'm sure lots of people and lots of big projects would still be interested in joining the initiative. They could reward participants with, say, 3 year of Copilot access after it is officially launched, and people would love that. But instead they just take code without asking or attribution and keep pushing it, and now we are in this situation.
- z9znz 4y agoMS needs to give up and terminate Copilot. The potential legal issues are there, but that's not why Copilot should die. Copilot should die for any (or a combination of all) these reasons (and more which I don't mention): - the operator has to already understand the emitted code to be able to determine if it is what is needed, or to modify it if it is close but not quite right - the operator may have a false sense of capability, leading to bugs and other problems that would appear later (in production?) - wrong suggestions are a distraction from the careful mental structures which one maintains while writing software - any problem that Copilot can solve with guaranteed correctness is probably trivial or already met by a (battle tested) library Forgive the analogy, but effective automated code generation is like autonomous driving systems. Anything less than 100% accuracy is a risk, and in these examples risk of incorrect behavior is not acceptable. Copilot seems like a pointy-haired boss fantasy where they can hire only junior programmers and expect successful software products.
- siliconc0w 4y agoYou're missing the big picture, first you create a lot of licensing violations littered throughout internal code and next they can sell you an Azure hosted open-source licensing annotation AI to fix it.
- airstrike 4y agoITT: armchair lawyers go after GitHub
- NautilusWave 4y agoCan someone clarify if copyright violation is actually considered "illegal"? As far as I know it's a civil matter and not something a state or federal government would attempt to prosecute.
- ummonk 4y agoIt’s a crime in certain cases, as piracy site owners can attest.
- mark_l_watson 4y agoFirst, the author’s book Beautiful Racket is very cool, recommended. I largely disagree with this article, at least for MIT, BSD, etc. training code examples. The small autocompletions, even if they are several lines long, sort of seems like fair use to me. I do think that CoPilot should have an option to use a smaller model just trained in code that has very liberal use licenses, because I think the use of GPL, etc. licensed code is problematic - at least for me. For what it is worth, I have a lot of Apache 2 licensed repos on GitHub (largely examples from my books) and I am pleased if my code contributed a small bit to the CoPilot training data. I also publish my recent books under Creative Commons, allow reuse, even commercially licenses: basically anything I do that might help someone, I am all in for sharing.
- int_19h 4y agoSharing is distinct from attribution. Are you okay with your code being reused without attributing it to you? If yes, then why have you published it under licenses that explicitly require such attribution?
- BeefWellington 4y agoMIT requires attribution, which copilot does not seem to include in the cases where it fully reproduces existing code.
- 627467 4y agoHey, I despise bait and switch from large corps. But I also find it unsustainable this idea that societies and legal resources are wasted fighting for IPs. The code is out there. Millions of people are being trained and writing code based of the learnings of open data. Designers have "mood boards". Developers have open source. Right now I don't have sympathy for MS, but in a few years any you developer could just do what MS is doing with Copilot in their bedroom. Why would you care about the kid in their bed room training an AI with free (as in public) information?
- ROTMetro 4y agoThis is copyright. I put it out there with a government guaranty that I get to retain ownership out of it. Society at large benefits because more people put their stuff out their. You want to break that deal and you will end up with less sharing. Why are you for information silos? I am for open ideas and sharing. You are for stealing and breaking the moral agreement because distributing my code in the form of an AI analysis doesn't groke with you as the same thing as distributing it in another manner.
- Aeolun 4y agoThis right here is why we can’t have good things.
- TheMiddleMan 4y agoThere's a big difference between learning and memorizing. If the AI is "learning" how it works by studying public code then using its knowledge to create, that's okay. But if it's just memorizing code and reciting it back, not okay. Just like if a human were doing this. Of course we don't currently have ways to know the difference [that I know of] since AI is a black box. Interestingly, current AI is not capable of truly understanding how code works and how it will execute, so it has to learn in it's own way. I suspect it can learn what valid syntax is, but I doubt it is aware of how the code will execute. It's possible this is just a case of Overfitting. https://en.wikipedia.org/wiki/Overfitting https://en.wikipedia.org/wiki/Overfitting
- nonasktell 4y agoWhy would anyone want to stop Copilot is beyond me. Reinventing the wheel, millions of time a day, is an atrocity. Millions of (wo)man hours, wasted, every single day, on writing solutions to problems that have already been solved. There is a partial solution to this, and it's making people angry, it's crazy. If you put your code publicly on the internet, you should expect that people will reuse your code at some point, no one broke into your privates repositories. Why would anyone waste their time to make other people waste more of their time is really beyond me. Let go of your egos for once.
- Ecco 4y agoYou’re missing the point. It’s not an ego problem: if you put your code on the internet with a license you should expect people to respect the license’s rules…
- deleted 4y ago[deleted]
- macinjosh 4y agoNot everyone believes in intellectual property and good luck enforcing that license worldwide.
- gruez 4y agoSo? Most of the developed world have legal systems that does believe in intellectual property. The fact that a few people "don't believe in intellectual property" because they want to torrent movies/games is mostly irrelevant when it comes to the software engineering profession.
- mbf1 4y agoIt doesn't matter what you believe. It matters what the judge and jury say when this goes to trial, and it will go to trial because Microsoft has a lot of money.
- nonasktell 4y ago
- steve_taylor 4y agoOpen source has trained me and countless others. We have learned from it. Why shouldn’t machines learn from it too? Is co-pilot copy-pasting slabs of code verbatim? I see Copilot as a net positive. Open source is for sharing and learning. Copilot is sharing and learning on steroids.
- batmanturkey 4y agoYes, copilot IS copying code verbatim from GitHub hosted repos, right now. The license and attribution are stripped from regurgitated copied code snippets. Verbatim with no context, no attribution, no citation, no reference to the project it’s part of… If the people don’t known which project the code was taken from, how can they one day contribute to that codebase? Copilot is an interloper who doesn’t even tell you which project the code snippet was ripped off from!!
- redog 4y agoCopyright is the problem. The rest of this is just dancing around the legal framework built to support the bullshit.
- obiefernandez 4y agoHow can I personally and proactively fight against this effort?
- wkdneidbwf 4y agoit seems like copilot is simply a search engine in this context. when i search gh or google or <insert tool> i can get code snippets without seeing the license. how is copilot doing something fundamentally different?
- skytrue 4y agoIt's always interesting to see the buzz that occurs when Copilot is brought up as a topic. This place is called "HackerNews", yet routinely people forget that a "hacker" is somebody using technology to overcome novel problems. Doesn't GitHub Copilot fall into this category? Why is there such an outcry over a technology that has been in the public's hands for less than a year? I'm almost certain that the team responsible for Copilot is going to try to figure out how to avoid spitting out code verbatim, as that's obviously not a good look. It's most likely the case that in 1, 3, 5 years, Copilot won't be spitting out code blocks verbatim. It will generate rightsize code, trained on lots of publicly available code, and start reducing the surface area required to code/develop. Stable Diffusion doesn't get in trouble right now because the artwork looks like permutations of different works; text is easy to copyright, style is more challenging, but artists are facing up against the same reality. There's no rolling this back; ML models are going to remove a ton of cruft from creative/labor based endeavors, and people are going to need to evolve to stay relevant.
- stefan_ 4y agoNothing is in the publics hands; they have taken the public data and given back nothing but a blackbox that you pay money for. They don't even trust the thing to train it on their own code, yet their boss is over here telling us they are "learning". It's a damned insult.
- anigbrowl 4y agoI won't shed any tears for Microsoft if people liberate or reverse engineer the model weights.
- iudqnolq 4y agoPlenty of people approve of individuals doing things they disapprove of large corporations doing. For example, I wouldn't care if a small YouTuber used a copyright song in the background. I would care if Disney stole a small YouTuber's original song and used it in a movie. This is entirely consistent within my ethical framework: scale and power matters.
- 4y ago
- deleted 4y ago[deleted]
- hsuduebc2 4y agoIn this next episode of corporation name seems like a cool corp but reveals as selfish and malicious inc. we could saw corporation name act selfishly and maliciously as in every other episode. See you next time kids.
- naikrovek 4y ago> Why couldn’t Microsoft produce any legal authority for its position? Absence of proof is not proof of absence. They don't owe anyone anything beyond what they agree to provide to users of Copilot via its license agreement or to GitHub users whose code it has used in accordance with that license agreement. Those agreements define what they owe. That's it. The only way those license agreements don't hold up in court is if they are somehow deemed invalid. I do not see Microsoft making that kind of mistake. This website is designed to get people angry, and that's all it is going to accomplish.
- naikrovek 4y agoHere's the thing about GitHub that most people do not realize. I find this funny because of all the talk of following license agreements, very few have taken the time to read the terms of service for GitHub. from their terms of service: "Short version: You own content you create, but you allow us certain rights to it, so that we can display and share the content you post." emphasis mine. that's what they call the "Short version" of the following paragraphs, which are found here: https://docs.github.com/en/site-policy/github-terms/github-terms-of-service#d-user-generated-content https://docs.github.com/en/site-policy/github-terms/github-t... they allow themselves the right to display content you upload to others. GitHub does not seem to really put a cap on that in terms of what intentions it needs to have or for what purposes it needs to share your content. this seems to me that, by putting your code on github.com, that you are granting GitHub license to show it to others. period. IANAL, but it seems like all code anyone puts on github.com is dual-licensed, at least. GitHub gets their own rights to your code. I read this before I signed up, and while I can't remember if this exact passage was present at the time, I was ok with everything GitHub wanted at the time, and I continue to be. githubcopilotinvestigation.com doesn't seem to have much hope of doing anything except getting people mad. but you all were already mad anyway, weren't ya?
- gjkqfc 4y ago> but you all were already mad anyway ... This seems to be the line of argumentation agreed upon by several waffling pro-GitHub posters. Many comments have some variation on that diversion from the issue. > GitHub gets their own rights to your code. This is preposterous and false. GitHub has the right to display the entire work, properly attributed and licensed, to others. No new licenses are given, no dual-licensing takes place, no code-laundering is permitted.
- naikrovek 4y ago> No new licenses are given, no dual-licensing takes place, no code-laundering is permitted. I suggest you read the terms of service again. here, I'll link directly to the license grant: https://docs.github.com/en/site-policy/github-terms/github-terms-of-service#4-license-grant-to-us https://docs.github.com/en/site-policy/github-terms/github-t...
- scombridae 4y agoBut this has long been the deal. In order to offer their services gratis, Big Tech makes money on your data, which you've freely provided. Welcome to the last twenty years of the software economy?
- cryptonector 4y agoMaybe MSFT should have one instance of Copilot for each common license, and then the user gets to pick which licenses they want to deal with when using Copilot. If you're writing code for a BSD-licensed codebase, you might accept Copilot trained on BSD- and MIT-licensed code, as well as any other license that's compatible with BSD. If you're writing code for a proprietary codebase you might want to exclude Copilot trained on any copyleft licenses. And so on.
- black_puppydog 4y agoDon't give them a more complicated problem. It seems like they're already struggling with any distinction between code they can and can't use for this. :| Which is ironic given that this is Microsoft. Whatever happened to "don't use programmers' code without paying them" and the whole "proprietary software is better because it sustains the programmer"?
- pmayrgundter 4y agoCrazy idea.. have the automatic code generator check if the code is too similar to a source it was trained on, and if so, automatically include attribution as well. Ta da!
- nojvek 4y agoIs GH Copilot only on public repos? My assumption is that code from private repos was also showing up. I feel like I read an HN article about this previously. Don't have evidence but that seems like a much bigger issue and trust violation if that is true. Kinda like how gmail was reading everyone's emails and showing ads based on them.
- bloppe 4y agoI think it's important to realize the exact implications here: - MS absolutely has the authority to copy, use, and even train their models on your GPL-license code, because you agreed to let them do that when you signed their EULA when you decided to host your code on GitHub. - This authority does not extend to CoPilot users, who cannot republish your GPL-licensed code without respecting the license. But remember that people have always had the ability (not authority) to copy and use open source code in violation of the license. This simply makes it embarrassingly easy for a person to do so unknowingly (although, legally, this would probably be considered negligence, not ignorance). IANAL but I wonder if the extreme facilitation of copyright infringement here could be considered gross negligence on the part of MS, as they're almost entrapping their own customers in a minefield of copyright concerns. Can't wait to find out. The logical next step in this arms race is for the GPL camp to build tools to automatically search for copyright infringement in large codebases. Copyright holders could set up hotlines for insiders to blow the whistle on infringement in exchange for compensation, since AFAICT all litigation precedent in the US has so far resulted in settlement.
- moolcool 4y ago> MS absolutely has the authority to copy, use, and even train their models on your GPL-license code, because you agreed to let them do that when you signed their EULA when you decided to host your code on GitHub. What about GPL code which you don't own, but post to Github, Like the gcc mirror repo?
- naikrovek 4y agoread the terms of service. you must have the right to publish the code you put on github.com, and by publishing to github.com, you assert that you have the rights to do so. you also grant GitHub the right to show that code to others, no matter what license your code is licensed under. why does no one read the terms of service or license agreements? these questions are answered there and this "copilot is stealing" stuff won't even make it to court.
- broodbucket 4y ago
- cabaalis 4y agoIt seems a benchmark of a transformative technology is whether or not people attempt to use the legal system to stop it.
- egypturnash 4y agoIt's hilarious how when I express displeasure about AI image generators looking likely to take a huge bite out of my profession of "artist" and playing extremely fast and loose with fair use, I get told that it's completely inevitable now and I should either retrain as a prompt engineer or go join the buggy whip manufacturers, but now that this is clearly violating programmer copyrights, you folks are starting to get angry. I'll just leave y'all with my favorite of the things you keep telling me to STFU about art AI with: If you're the kind of programmer who feels threatened by this, then you're not a real programmer.
- bugfix-66 4y agoYou're absolutely right. Copilot and Dall-E (and so on) are all bad in the same way. Many of us agree with you.
- corndoge 4y agoUnless the image generators routinely generate specific works produced by you (or other artists) then it’s not a directly comparable situation to Copilot.
- kps 4y agoLike this? https://news.ycombinator.com/item?id=32573523 https://news.ycombinator.com/item?id=32573523 > I just got a Dall-E render with a very intact "gettyimages" watermark on it.
- int_19h 4y agoYes, this would be a good example of genuine copyright infringement that shouldn't be tolerated. Of course, it doesn't mean that all or even most DALL-E output infringes on someone's copyright. The same is true for Copilot. I think both have many legitimate uses if and when the "copyright laundering" issue is solved.
- klyrs 4y agoI'll admit it took me longer to connect the dots on this one but when I was tinkering with an image generator and it gave me a clear istockphoto watermark, I knew something was amiss.
- Imnimo 4y ago>Copilot introduces what we might call a more selfish interface to open-source software: just give me what I want! With Copilot, open-source users never have to know who made their software. They never have to interact with a community. They never have to contribute. >Meanwhile, we open-source authors have to watch as our work is stashed in a big code library in the sky called Copilot. The user feedback & contributions we were getting? Soon, all gone. I don't see how you square the above complaint with this: > First, the objection here is not to AI-assisted coding tools generally, but to Microsoft’s specific choices with Copilot. We can easily imagine a version of Copilot that’s friendlier to open-source developers—for instance, where participation is voluntary, or where coders are paid to contribute to the training corpus. Is an AI that was trained on opt-in or paid-for training data any less damaging? How would these choices have alleviated the problems described above?
- smegsicle 4y ago> GitHub Copi-lot inves-ti-ga-tion lol rarely see such aggressive use of soft hyphens in page titles
- semireg 4y agoIs there an AI system for those dot woodcut prints ie WSJ?
- mgraczyk 4y agoAll this discussion of legality is interesting to me, because I'm pretty sure that if Github ran a search in the background, found the corresponding license for the code snippet, then showed it to the user in some cookie-banner like annoyance, it would be completely legal. This is what Github already does on their website with a search bar. Yet somehow I think most people upset about Copilot would not like that outcome.
- culi 4y agoHow do you run a search on the product of a bunch of neural network weights? Do you just mean like double-checking that the code it produced isn't a direct copy of something copyrighted? If so I think that introduces a lot of other issues. There's an interesting phenomenon of different independent comedians suing talk shows for stealing their jokes. It almost always turned out that those jokes weren't actually stolen. Instead, there's only so many jokes you can make about a given news story and there's bound to be overlap between a whole room of comedians trying to milk every event of any comedic value and random independent comedians doing the same
- mgraczyk 4y agoYou run search on the text output. Yes I mean that they would search the code. It's certainly legal to show LGPL code snippets along with the license (It may also be legal without this, IANAL).
- stingraycharles 4y agoIt’s interesting and pretty much uncharted territory from a legal perspective, that’s for sure. It very much relates back to the discussion about accountability of machine learning models, in that it’s desirable to be able to explain how/why some output was generated. I don’t think a banner would be sufficient here, though; perhaps some references to the inputs that were used to generate the output, but that’s often very difficult to pinpoint. Whatever happens, if this ends up setting some sort of legal precedent it will have a big impact on the industry, and I personally hope it leads to more accountability and transparency of the models, rather than the black boxes they are now.
- schoen 4y agoHere are a few thoughts I haven't formulated before: It seems clear enough to me that training AIs on copyrighted works is typically or commonly a fair use under existing law, because the AIs can and commonly do learn non-copyrightable elements and aspects of those works. It's very obvious from enormous numbers of examples that current AI systems are capable of learning much more abstract features of human culture (grammar, concepts, facts, cultural tropes, and many others). A human being doesn't violate copyright in learning from a copyrighted work, including when that human being is later more able to produce other works based on that learning (e.g. reading fantasy novels and learning concepts, tropes, or vocabulary that one uses to produce other fantasy novels; reading a newspaper and learning facts that one incorporates into an essay; learning artistic techniques or stylistic conventions from studying existing artworks and using them when producing new artworks). Current AI systems are (amazingly) becoming capable of all of these things and may do them in ways that are somewhat akin to how human beings do them. (although I guess Jaron Lanier would object "that's what they want you to think") But there are also examples in existing copyright doctrine where people accidentally repeat enough of a prior work to get in trouble for infringement -- most often with song composition (like George Harrison's "My Sweet Lord") because relatively small pieces of melody (which a person might easily memorize) may be considered copyrightable. If human beings had much more accurate memories, copyright would be quite a bit more intrusive (and/or quite a bit less effective) because, following any exposure to some kinds of works, we could use our own memories to reproduce those entire works from scratch for our own use or pleasure without obtaining authorized copies from elsewhere. Computers do have such accurate memories, and machine learning systems, which are optimized for things like maximum likelihood estimation, can and do reproduce both copyrightable and non-copyrightable elements of works that they've been trained on. After all, the maximum likelihood continuation of a fragment of a text or a song is ... the complete original work. And the ability to reproduce the complete original work would, other things being equal, reduce loss in training. After all, that's something someone might specifically ask for, and if the system could oblige, it would be doing a better job of providing what the user wanted. It's relatively foreseeable that machine learning systems would potentially be able to reproduce both copyrightable and non-copyrightable elements of various works, because the distinction between the two isn't especially clear from an algorithmic or mechanical point of view. (For instance, facts aren't copyrightable, but the notion of what constitutes a "fact" for this purpose is a culturally-bound legal notion and not at all straightforward to make precise.) But if you had a human author or artist or scholar or programmer who was "trained on" exposure to an enormous body of works, and that person had an exceptional eidetic memory, you could imagine that he or she would be perfectly capable of recreating many of those works from memory (and that other people might request such recreations). (Again, in music in particular, it's already routine that someone could have unambiguously copyrightable material memorized and be subject to copyright restrictions on performing songs. Like if a singer or band performs a cover from memory.) If you wanted to avoid this ability then you might need to build in an explicit notion of copyright that limits the accuracy or level of detail inside of the model in some way. This is tricky because (1) I don't think people have really tried to do this much so far, (2) copyright applies very differently to different categories of work, (3) it obviously wouldn't satisfy critics even if it mitigated the most extreme examples of "regurgitation", and (4) it would be kind of weird because you would be intentionally limiting the quality and extent of learning that the system was allowed to do. (I imagine Jaron Lanier getting mad again about my repeated comparison between human learning and machine learning, and between human memory and machine memory) Some of the weirdness in point (4) is that accurate prediction is usually cool / great / impressive / accepted as an appropriate goal or capability, but if it's too accurate in certain contexts, it may be deemed a copyright infringement. Like if you said "what word comes next? FOUR SCORE AND SEVEN YEARS AGO OUR FATHERS", there's a clear correct answer and knowing it requires having a certain text memorized. OK, if you said "what word comes next? MR. AND MRS. DURSLEY OF NUMBER FOUR PRIVET DRIVE WERE PROUD TO SAY" ... same thing, but Bloomsbury Publishing may be unhappy if you have a system that can get all such questions right.
- woah 4y agoIt would be sad if someone succeeded in shutting down CoPilot for this kind of copyright stuff. It is genuinely useful. I don't care that it reproduces copyrighted content. The only way you can get it to do that is to bait it with the function names of functions that have already been copy and pasted thousands of times onto GitHub without proper licenses. Luckily, someone will probably come out with a "renegade" version trained on whatever makes it a useful assistant to my coding. I won't be afraid of accidently violating copyright myself, because I won't be trying to bait it into reproducing heavily copy&pasted cherrypicked examples, and I won't use 20 lines of its output with zero modification.
- MarcelOlsz 4y agoIf it was also free and open source, sure. But it's not, it's a paid product that one party reaps the profits from.
- culi 4y agoYeah I think it just means a non-commercial alternative would be made to replace it
- nightski 4y agoPersonally I'm not worried about the end user using copyrighted code. That is their responsibility. If you have verbatim GPL code in your commercial closed source code base that is a liability and it might be dangerous to use copilot. What I have more of a problem with is Microsoft charging for copilot which was trained on copyrighted code without any permission whatsoever which they really have no right to utilize/charge for.
- bruhhh 4y agoThis does more harm than good. If you set a precedent, then things like stable diffusion will also be illegal since it's trained on public data. OP just wants to make money from microsoft using fearmongering and false sense of righteousness
- slugiscool99 4y agoMaybe the software engineers are worried they're being made redundant, but it is super fair for them to not allow their own work to make them redundant without permission
- angusturner 4y agoI get the impression that many peoples' grievance with generative AI (text, code, images etc.) isn't _really_ about the data provenance. Or at least, it feels secondary, compared to the general disruptive nature of the tech. If tomorrow someone released a StableDiffusion, CoPilot etc with the same functionality, but respecting the provenance of the data (i.e. licensing etc), what concrete difference would this make? Programmers and other creative professionals would still (reasonably) be nervous about the implications for their livelihoods and communities. At some point it will be possible to prompt a model for music in the style of <random artist>, and having never heard <random artist>, the model will generate a convincing emulation, based purely on statistical knowledge gleaned from millions of unrelated songs and text pairs. (I give it 5 years). Now what? <random artist> should still be concerned (or not), but at least we're talking about the correct issue: How do we co-exist with generative models that massively disrupt/alter the process of doing creative or intellectual work?
- jackdaniel 4y agoSometimes when working on software the goal is not to get a competitive advantage, but to promote some ideas. Copyleft license is a tool that aims to help with granting, that the work derivatives are available to the public to read and modify (i.e to prevent closing the source code of the program that is commercially sold). The concrete difference you ask about is that the work derived from copyleft code retains the license and the source code can't be closed. If you scrap the license, then the code created by someone who had clear goal in mind when writing it for not making improvements over it closed source, ends up with possibility of being closed source.
- glouwbug 4y agoBig money here. Good luck
- blackoil 4y agoFrom all the discussions, it seems people are rooting for MPAA alike organization and ContentId like system for code.
- mkr-hn 4y agoA sizable, possibly plurality cohort of fully adult tech people is young enough to not know about United States v. Microsoft Corp. This would explain a lot of comments I see on this topic. If you don't know Microsoft's history, a lot of what more informed people are worried about seems overblown. Copilot was Microsoft's first test of people's trust after the GitHub acquisition. It's going very, very, very poorly. There were ways to do this with consent and collaboration with the people and projects it takes code from, but they're acting like classic Microsoft here. Too many people are focused on what's legal. It's fine to think of, but law is the last stop before the breakdown of society. Microsoft skipped society and went straight to sparking an inevitable test of and possible reshaping of copyright law.
- enlyth 4y agoOr some of us do, and just have our own opinions. I literally don't care if someone steals my code, and I think the current state of digital copyright is nonsense that does not benefit society in any way.
- gizzlon 4y agoThat is certainly your right, but very interesting in this discussion. Copyright laws should change if needed, but this is not the process.
- mkr-hn 4y agoYou're free to put your code under something like CC0 if you feel this way. Everyone else who puts their code under a license that requires at least attribution and expects Microsoft to follow it can continue taking Microsoft to task for ignoring that responsibility at scale.
- nonbirithm 4y ago> Too many people are focused on what's legal. It's fine to think of, but law is the last stop before the breakdown of society. Microsoft skipped society and went straight to sparking an inevitable test of and possible reshaping of copyright law. Maybe it's illuminating of a trait of human nature. On the stable diffusion webui repo many people have stated that they would continue to use the code even if it were stolen or unlicensed. These people aren't a part of a corporation; they are average netizens handed a technology essentially indistinguishable from magic with nothing in place to prevent its use. If the tech is simply so impeccable as to be irresistible then a higher order framework needs to be in place to teach people not to bite because they will be bitten back.
- low_tech_punk 4y agoSadly, I think this marks the beginning of a winner-takes-all economy fueled by AI. Just imagine how in a lawsuit like this, OpenAI can use GPT-3 to generate eloquent court speech with statistical confidence that it can defeat human lawyers? It just comes down to TPU power.
- braingenious 4y agoI’m really interested in seeing how this gets litigated. I imagine it will involve a lot of philosophical arguments about attribution and what the software is actually doing. I’m also curious to see if/how Amazon CodeWhisperer takes advantage of this whole debacle.
- Springtime 4y agoA bit meta but anyone know why the submission title contains unicode between various characters? It's hidden on both Chromium/Firefox when viewing the page but when saving the page it reveals them in the text field, eg: `GitHub Copi_lot inves_ti_ga_tion` Plugging the title into a unicode converter shows they're 'soft hyphen' characters GitHub Copi [0x00AD] lot inves [0x00AD] ti [0x00AD] ga [0x00AD] tion Edit: apparently they're for indicating to formatters where character breaks should be, though I can't understand the consistency here.
- kmeisthax 4y agoIf Copilot itself is infringing then so is GPT-3, DALL-E 2, NovelAI, and Stable Diffusion. There's no legal argument that would solely target one application of this technology, and you can't build generative AI using current ML tools without relying on a very large corpus of public data. All AI is built on free-riding[0]. While there is no US case law that explicitly says "training AI is fair use", the Second Circuit says that scanning books to make a search engine for them is. And the absolute worst interpretation of AI is that it's just a very well-compressed search engine index for its training set data[1]. So I'm not entirely sure if we can even thread the needle to only ban Copilot or AI training as a whole without also creating harmful precedent for search engines. Actual judges may try, I'm not sure if they'll succeed. Internationally, the EU already legalized training AI on copyrighted works[2]. So if we do win against Copilot in court, all we've really done is shift AI research over to the EU where laws are already more favorable. I fully agree that Microsoft is shoving too much liability onto their users, though. And this, again, also applies to all generative AI. My personal opinion with generative AI is that it's a nice curio, but not anywhere close to "production-ready", and Microsoft and OpenAI are trying to sell us on a lie that it's better than it really is. [0] This also implies that all y'all playing around with image generators are just as much of a freeloader as Microsoft is. [1] This viewpoint is also called "compressionism". [2] This was part of the most recent EU Copyright Directive update - the one that added a de facto upload filtering requirement. It also added a copyright exception for museums and historical preservation.
- AlexandrB 4y ago> And the absolute worst interpretation of AI is that it's just a very well-compressed search engine index for its training set data[1]. I don't think this parallel makes sense because a search engine links to copyrighted works, each of which is still governed by its original copyright, while these AI create derivative works or reproduce the original works without even attribution. Indeed, if an AI was just and index for the training set there would be less of a problem because the origin of a work could be found and its license honored.
- victor9000 4y agoPersonally, I draw the line at corporations profiting from the derived works. But if companies want charge for tools that make these models easier to interact with, then that seems pretty reasonable.
- deleted 4y ago[deleted]
- gamekathu 4y agoA bit of a controversial opinion: to those who are defending CoPilot saying it "boosted my productivity" and would miss it if it is discontinued, maybe you are not a productive developer to begin with. I fail to see how searching the same snippets on Google or saving commonly used macros in your favorite editor would not yield the same amount of productivity. I have used CoPilot for several months and I actively stopped using it, because I was afraid I will be dependent on it, and it would actually reduce my ability to do critical code-building. I'm happy without it - sure it takes some micro seconds more to type out my code instead of autogenerating it, but I feel much more self confident in my own coding skills. CoPilot is a great research work - it is indeed spectacular to see how pre-training can achieve such impressive code completion results. However, in my honest opinion, it should not be a tool for a serious developer.
- e-clinton 4y agoPart of me feels like this will help big tech and hurt potential startups that’d compete in this space. Microsoft has the resources to make this issue “go away” while smaller incumbents will not.
- minhazm 4y agoThere's been a lot of discussion around licenses but I'm not even sure if they matter for Copilot. I was reading their terms and conditions and there's a paragraph that basically says they have the right to display and share your code with other users. So even in the case where people are directly prompting Copilot with specific function names, I think the terms and conditions still cover them. > We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video. https://docs.github.com/en/site-policy/github-terms/github-terms-of-service#4-license-grant-to-us https://docs.github.com/en/site-policy/github-terms/github-t...
- brigandish 4y agoPerhaps the only way out of this is to start suing the users of Copilot, much as some jurisdictions target the users of a product (e.g. drugs, prostitution) as a means to shut it down when the providers are too difficult or numerous to challenge effectively.
- truth_seeker 4y agoI am feeling very greedy but .... With all that intelligence if GitHub Copilot can't produce easy to use and manage full stack framework yet with distributed database inbuilt in either any existing programming language or perhaps a new one created by itself then its not useful for me.
- a254613e 4y ago>But how will you feel if Copilot erases your open-source community? Jesus Christ, dramatic much? Are people that stumble upon a piece of code while googling how to do something, and end up copying and pasting the code from the repo, really building the open source community? Because that's essentially what it is. Whether I use copilot to generate a tedious function, or I copy it from your open source repo I'm on the same level of being a member of your open source community. This whole thing feels like artists screaming how AI generated art is horrible, trying to figure out how to sabotage it, or how to start lawsuits - just because their value went down just a bit. Same thing with developers.
- youssefabdelm 4y agoCouldn't agree more... It's very depressing that this post is popular, wouldn't want Copilot shut down over some drama queen lawyers that have no connection to the reality of software development and ALL creative fields. Creation requires influence: https://www.youtube.com/watch?v=nJPERZDfyWc&feature=emb_title&ab_channel=KirbyFerguson https://www.youtube.com/watch?v=nJPERZDfyWc&feature=emb_titl... The entire fucking concept of intellectual property and copyright is flawed from the get go. The issue people are wrestling with beneath the surface is not copyright but the monetary system itself which incentivizes this "chisel off one another" behavior and "MINE!" behavior because otherwise how will you survive if you can't monetize your actions?, but intelligent socioeconomic alternatives exist: https://www.youtube.com/watch?v=lBIdk-fgCeQ https://www.youtube.com/watch?v=lBIdk-fgCeQ People are trying to solve this problem in an ass-backwards way. Either move to universal basic income or a resource-based economy and make all ideas 'free', 'copyable' and 'remixable' since it doesn't matter either way you have access to some resources (in UBI) or all resources for free (in resource based economy) and don't need to monetize anything since you have access to everything...instead people are content with making life shittier. "We stand on the shoulders of giants" said Newton, but oh no.. this piece of paper called 'the law' knows better!
- mahogany 4y ago> The entire fucking concept of intellectual property and copyright is flawed from the get go. Many people are upset because Microsoft is hiding behind copyright and lawyers to enforce it, while at the same time ignoring the concept of intellectual property when it comes to smaller players. I'd imagine that if Microsoft removed copyright on all their code and released it and Copilot as open source, there would be much less outrage. The issue here, for me at least, isn't centered around copyright as a concept; it's about the asymmetry of the situation. Microsoft is exploiting those without any recourse in order to sell a product. Your ideas about a new economy and no copyright are interesting, but they will not happen any time soon. In the meantime, in reality, Microsoft is making millions based on an enormous pile of community code while not offering their code back to that community.
- perryizgr8 4y agoI am glad all the legal bs didn't stop MS from making the product. Copilot is surprisingly effective, it truly makes life easier for me, as a developer. The fact is that if you give your code away publicly, you cannot finely control what the world does with it. If this is not acceptable to you, keep your IP private. If these guys manage to shut down or cripple Copilot using legal mechanisms, you can bet there will be a Chinese/Russian alternative that will be even more indifferent to your LICENSE.md, and you won't be able to get it shut down using the courts.
- welder 4y agoThis reminds me of pirating music. Lawyers tried futilely to stop it, but if something is technically possible people will find a way to keep doing it. Maybe you set some legal precedent on fair use with AI, but it won't prevent the real world usages if there's a benefit to the technology.
- hbarka 4y agoDo the same copyright issues arise with AI-generated videos learned from Shutterstock? https://news.ycombinator.com/item?id=33239706 https://news.ycombinator.com/item?id=33239706 https://waxy.org/2022/09/ai-data-laundering-how-academic-and-nonprofit-researchers-shield-tech-companies-from-accountability/ https://waxy.org/2022/09/ai-data-laundering-how-academic-and...
- welder 4y agoTime for a new open source license specifically allowing fair use for machine learning?
- javajosh 4y agoIt's too bad we can't experiment with interesting things like Copilot without worrying about remuneration and the respecting of rights. But that's the way of the world - we must think of these things. MS/Github should give code copyright holders a simple and easy way to opt-out of contributing their code to the Copilot corpus. Currently the only way to opt-out is to make your repo private. That's not good enough. It would be better, of course, if Copilot was opt-in, but they'd never go for that.
- t43562 4y agoThey have done so already with their license and there's no legal reason for them to have to opt-out.
- rickydroll 4y agoRSI took away my ability to write any significant amount of code 30 yrs ago. co-pilot plus speech recognition restored that ability. what impressed me most was that from a textual description,co-pilot gave me code that could have been written by my mind and pre-injury hands. from the comments here, if I push copilot into giving me code that I would have written for a given problem and that code violates licenses, then who is responsible for the copyright violation? co-pilot for giving me code that looks like copyrighted code or me for tweaking co-pilot commands to give me the code I envisioned which looks like copyrighted code? also consider that the very tools used for solving problems in code lead coders to a small number of solutions for a given problem. is it plagiarism or parallel original thought? also consider that when I wrote code, if I was solving a similar problem to what I solved before, I recreated that previously used code fragment (or larger) and use it to solve the problem at hand. I had zero issues leaving a trail of duplicate code behind me especially if the code was a major part of a software patent. I didn't care, my code was lauded for it's readability and reliability. reuse the same concepts in multiple variations, you get real good and writing code correctly. maybe co-pilot like programs could scan existing code bases and find examples of code fragment plagiarism with the goal of showing that software copyrights are useless.
- carom 4y agoThe issue isn't that something like copilot shouldn't exist. The issue is the disregard for open source. You want a proprietary code completion tool? License your training set properly. If you want to build on open source then open source it and what it produces.
- scotty79 4y ago> how will you feel if Copilot erases your open-source community How will you feel if greed of a lawer erases progress of your tools? Lawyers are a detriment to anything they touch. Letting them into software was the biggest mistake we ever made. We should kept them away same way they are kept away from math.
- sergiotapia 4y agoIs there a license that explicitly forbids corporations from ingesting my code and making a billion dollars off of my work for free? The AGPL? I've been using the MIT license for more than a decade, but it's time to change that.
- josephcsible 4y agoAlmost every FOSS license requires attribution, and Microsoft already seems perfectly happy to violate that, so I don't see why they'd be any less happy to violate whatever other license you'd come up with.
- modernerd 4y agoI love that Matthew is investigating this and agree that Copilot warrants more scrutiny. His suggestions that Microsoft let developers opt-in to having source used for training purposes, to pay for source it uses, and to attribute or credit it appropriately all seem reasonable. Can someone help me to imagine a reality in which these points are viable concerns? > …how will you feel if Copilot erases your open-source community? > …Copilot will become not just a substitute for open-source code on GitHub, but open-source code everywhere. > …Copilot is merely a convenient alternative interface to a large corpus of open-source code. > With Copilot, open-source users never have to know who made their software. They never have to interact with a community. They never have to contribute. Is the author suggesting that Copilot will be used in place of `npm install next react react-dom` or `cargo add tokio --features full` or `raco pkg install pollen` — that developers will be content to use augmented autosuggest in place of large, well-tested, well-documented open source libraries? Does he see Copilot's final form as some kind of AI package manager that drops a library of untested unattributed undocumented files into our projects? Or is it more that he thinks those libraries won't exist because open source contributors will grow to feel more abused than they already do, perhaps quitting the scene or developing in private, like certain artists have already done in response to the AI art movement? There is already such a huge disparity between paid package consumers and unpaid package contributors. I haven't seen that change since Copilot launched in beta or under general availability. I see the same ratio of help/feature requests compared to code and documentation contributions that I always have. And package usage has not declined so far for the open source things I work with. It would be nice to learn more about the “Copilot will lead to the death of open source communities” line of reasoning — what is the author's perceived timeline to open source's decline and fall as a result of Copilot's current path?
- belter 4y agoTL;DR: GitHub(Microsoft) declared that: “training [machine-learning] systems on public data is fair use”. When asked for the relevant jurisprudence to support it's position, could not provide any.
- kioleanu 4y agoUhm, strangely I get a Connection Reset error in the browser when I try to access the URL from the corporate network, but it works without problems from my phone
- gareth_untether 4y agoAn intrim update to Copilot could link to where the code as pulled from. Or maybe there's a way for open source devs to add a comment to the code that links to their community/repo. If it was standardised then any data gathering would need to follow the collection rule. I agree with the article's long term outlook about community and code quality, it is a very long term outlook though. It makes me wonder if humans will actually be writing code.
- w10-1 4y agoDon't confuse what you want with what the law says "Your work is under copyright protection the moment it is created and fixed in a tangible form that it is perceptible either directly or with the aid of a machine or device" [https://www.copyright.gov/help/faq/faq-general.html https://www.copyright.gov/help/faq/faq-general.html] A copy is made whenever that text is displayed, e.g., in GitHub's UI. Even that copy is subject to copyright. Is there an excuse/exception? In this case, there is no "fair use" exception, because exceptions have to be litigated case-by-case to be recognized, and there are no remotely similar situations. Don't forget: Lexis is a multi-billion-dollar business built on protecting the copyright to the page numbers in the otherwise public court opinions. Does the law actually protect people if it's too costly to enforce? Not really; hence the blase attitude. Congress is considering a "small claims" system for copyright, to remedy the big-firm bias. [https://www.copyright.gov/title17/92appm.html https://www.copyright.gov/title17/92appm.html] In the ML era, data is the new gold. Many, many firms nowadays get a good chunk of their revenues from selling their private view of "public" data: Facebook, LinkedIn, credit reporting companies, ADP, etc. Microsoft has gone all-in on stealing that gold from open-source developers. It's not just that the code replication reduces any need to get the code from the source. But removing any link to the source destroys the value most-commonly sought in open-source software: recognition. Salaries are the biggest expense of tech companies. They do everything they can to increase labor competition and reduce reputational rents: outsource, cross-train, promote open-source (for competition) and destroy any reputation networks or systems that justify higher rates. And, of course, standardize on containerized copy-paste or AI-generated software if they can. So, no: copilot is not legal, it's socially and economically destabilizing, and it presents structural challenges to developers. It's not good, but most will keep using it because although the vast, vast majority of developers are wage laborers, they aspire to be founders. They see it can make code fast, and they'll think it make them better.
- abigail95 4y ago
- melonmouse 4y agoAs a joke, I made a webpage where you can do attribution to ALL GitHub repositories: http://thanksforthecode.com http://thanksforthecode.com It scrolls past all the repos movie-credits-style. Doing it that way takes several days! It shows how abstract and absurd giving contribution to such a large body of works is.
- toastal 4y agoYou're missing your <noscript> tag
- chronolitus 4y agoTo stay sane: for myself as a developer, I consider github copilot as a (much) faster google/code search work-flow. I can copy / or re-mix code I find in a google search, but it's my responsibility to figure out the copyright situation of that code. Imagine if something like google didn't exist, and then it suddenly did. People would be saying: "This newfangled computer algorithm is giving everyone copies of my code with a misattributed licence, just by typing the function name and site:github.com !"
- hsbauauvhabzb 4y agoHow can I, as the lead of a small team, make sure none of my code ends up on copilot (or any other submission of our IP to third parties)? We use Devops internally, and IDE decision is up to the developer. Im unsure if vscode etc submit samples or just interact with GitHub. Edit: and furthermore, make sure it doesn’t import code from third parties. I don’t want my code being infringed upon, but also don’t want to accidentally infringe on others’ work. Legal or not.
- captainmuon 4y agoOh god please no. GitHub Copilot is a wonderful technology. I am not taking anything away from you if Copilot suggests code that is similar or identical to your copyrighted code. You were not going to sell it to me anyway. The following is supposed to be OK: somebody reads your GPLed code, learns abstract concepts from it, teaches it to me, I write code that uses the same algorithm. But it's not OK to abbreviate the process and reach the same result directly with Copilot. That is some Talmudic level reasoning. In a sane legal system, one would note that it is legal to do when jumping through pointless hoops, so it should be legal per se, and the system should be adjusted. Copyright is increasingly at odds with technological development. Not just since AI applications, at least since Napster or since floppy disks. Of course Matthew Butterick as a lawer would disagree - "It is difficult to get a man to understand something, when his salary depends on his not understanding it."
- HarHarVeryFunny 4y ago> The following is supposed to be OK: somebody reads your GPLed code, learns abstract concepts from it, teaches it to me, I write code that uses the same algorithm. But it's not OK to abbreviate the process and reach the same result directly with Copilot. The trouble is that this apparently is not what Copilot is always doing. If it had only "learned abstract concepts" from GPL'd (or any other form of copyright) code, then that would not be a problem, and of course that is kind-of what Copilot purports to be doing, supposedly learning the association between concepts described in comments and corresponding forms of implementation. However, apparently Copilot is sometimes NOT generating it's own code based on the concepts it has learned, but is instead just regurgitating chunks of potentially copyright-protected code verbatim. It'd be interesting to know if it is doing this deliberately (to maintain the coherence of what it is generating) or not - I guess the more of something it has already copied exactly the more it is likely to continue copying since that is the best "predict next word" continuation. Of course while it would be interesting to learn more about the mechanics of Copilot, that doesn't change the legality, or not, of what it is doing, another aspect of which (although IANAL) is how much of the original work is being copied. At the end of the day it shouldn't matter whether it's you or Copilot either learning from or copying someone else's code - exact same copyright protections apply.
- trasz 4y agoOne way to fix the problem would be to somehow feed Copilot a corpora of closed source code. This would either force Microsoft to add necessary copyright protections, or - which is imho more likely - would prove that those protections are already in place, but disabled for open source code. A good start would be to take a leaked code of Windows, and then mechanically adjust all the names, constant values, and code formatting, and then publish it and observe.
- tapia 4y agoI think that Microsoft should train copilot with their own code (they own certainly enough lines of code after all). If they think that that would not be a fair use, then why should be a fair use to use somebody else's code?
- amai 4y agoNobody would like to use Copilot if the quality of the code it produces would be like code from Microsoft. Garbage in = garbage out.
- random_kris 4y agoLet us just work on cool technical things without having to worry about this kind of bullshit. Knowledge data should be free to copy and do whatever we want with it
- preisschild 4y ago> Knowledge data should be free to copy and do whatever we want with it I'm more of a copyleft fan. Feel free to copy my stuff, but you have to make it open source as well.
- rapht 4y agoAll this just shows one thing : copyrighting / licensing "code" is meaningless... but of course that was already known by all those people who think that the US laws about copyright should not have been propagated to the rest of the world. "Code" is merely an algorithm put to work. There should be nothing inherently copyrightable about this, no more so than the recipe take a chocolate is just a way to put chocolate and a few other ingredients to work.
- spookyuser 4y agoThis is so stupid I can’t believe how this community has become toward some of the most inspiring new technology I’ve seen in a decade.
- amelius 4y agoHas anyone spotted licenses in the wild that specifically prohibit AI tools like Copilot?
- cool-RR 4y agoWhile the moral and legal discussions here are interesting and worth exploring, I find this text hyperbolic. Its premise is that the main way that people currently interact with open-source projects is by digging into their source code, copy-pasting away a snippet of code that solves a particular problem, and then of course giving the authors the required attribution. This is far from the truth. The main usage of most open-source projects isn't as code, but as a product. The median user of an open-source project wants to think about the project as little as possible. They want to be as unaware as possible of the code that makes up the project. They're happy to add the project to their `requirements.txt`, add a few lines to import and use it and then never think about it again.
- Sydneyco 4y agoI agree with that. Also, if we agree that GitHub copilot enables you to be more productive as a developer. Can we argue that it could help open-source communities by helping them finish projects faster?
- alexchantavy 4y agoI haven’t used Copilot but do its samples give links on where it was from? If so, that seems to be a sufficient funnel back to the OSS repo itself without the community harming aspects mentioned in the article.
- cool-RR 4y agoThey don't. It would be a profoundly difficult problem to find the right links for each suggestion.
- freeqaz 4y agoNot at all. There isn't even a way to get the "source" if you wanted.
- thamer 4y agoIt doesn't, because that's not how it works. Copilot doesn't recognize what you're trying to do and then paste a code sample from a repo it has in its index. Just like DALL·E 2 doesn't produce images that say "I picked these pixels from this image and this part from this other one and these colors from this third one". When a model is trained, it's effectively a set of hundreds of millions of numbers that when combined in just the right way can produce a specific output. In my experience the vast majority of the time Copilot doesn't write code that already exists. It actually uses the variables you declared, the functions that already exist in your code base, etc. It's not an index of best matches from GitHub for what you're trying to do.
- nonrandomstring 4y agoFunnily, for an article all about copying (lots) everywhere the author writes Copilot it appears as "Copi lot" in text browsers. Also the HN title appears the same (check it in hex dump) For example, from TFA file: 0005e10 o f C o p i 302 255 l o t
- qwerty456127 4y agoIt makes sense to copyright a book, but it doesn't makes sense to copyright a phrase (unless you are using it as a trademark motto or something like that), normally phrases are free for anybody to re-use. It makes sense to copyright a program, but it doesn't make sense to copyright a piece of code.
- PAMANOCH 4y agoThe problems have almost nothing to do with deep learning stuff. They are on the companies who develop such products. If a company use someone's code for a commercial product (a normal app), they do need to follow the license accordingly. If a company use someone's code for a commercial product (model training), they don't need to follow anything. If a company use someone's art piece for a commercial product (a normal game), they do need to get consent, and pay for the right to use to the hosting platform or artists themselves if it is not royalty free. If a company use someone's art piece for a commercial product (model training), they don't need to get consent or pay for anything. All the problems actually happen before the technical details, making the entire pipeline questionable.
- shp0ngle 4y agoTraining AI on copyrighted works is literally what Google always did. Look how Google News enraged news orgs. Now they come for the programmers. So now it’s a problem.
- albertizzley 4y agosuing Github for this seems like a neat idea to make money on our open source projects
- Dave3of5 4y agoQuestion for all those who are pro Copilot in this argument and are claiming fair use. Do these same rules apply if I manually copy someone elses copyrighted code into my codebase ?
- fartsucker69 4y agocould this be solved by MS brute-force shipping all the licenses (w/ references to their original projects) of all the repos they used to train to copilot along with copilot itself? it wouldn't cover cases where people illegally copy pasted some code into their projects with dubious / not explicit licenses, but this is the same as using any open source project in general.
- benced 4y ago1. New player shows up, changes value chain and creates abundance 2. People who benefitted from old value chain whine 3. New player throws them a bone with a small fund or maybe a setting box, doesn’t change 4. (A few years later) no one cares about the kooks who whined I’m not even 30 yet and I’ve seen this happen again and again - it’s frankly boring at this point. We’ve seen this with Spotify and music, newspapers and the internet etc. The practical truth is that Copilot is a useful tool for humanity to have. It is exceedingly unlikely it will be stopped because a small percentage of programmers - themselves a small percentage of people who benefit from code - feel their interests have been hurt. Change or get left behind (but make sure to enrich some lawyers on a pointless suit in the meantime).
- fregante 4y agoIf this becomes illegal, it will pretty much mark the death of free/open ML and its sets. If you can't train on data before asking for permission, the data set becomes sparse. Thje only people who will be able to afford this will be, you guessed it, established giants who can build their own sets.
- wkz 4y agoIf they really don't think that they need to comply with any license, then why not include all private repos in the training set? Could it be that they're worried about legal repercussions, whereas OSS is easier to (ab)use for this purpose because there's much less legal muscle behind it? It is also very telling that they have not included any of their own proprietary code in the training set. If it's merely suggestions that are generated, why not also train on the NT kernel? Office?
- 0xferruccio 4y agoSad to see people trying to make copilot illegal Using it is exactly like using Google. Google scrapes the internet and trains a model that gives you results for search queries on their website. The results may be copyright protected Copilot scraped the internet to train a model that gives you results for code snippets in your code editor. The results may be copyright protected
- HarHarVeryFunny 4y agoSurely the difference is that if you find something via Google/search you can then read any copyright notice to determine whether the code is OK to use for your use case (no joke if you're a corporate developer). If you're using Copilot to "generate" (but sometimes regurgitate?) code then AFAIK Copilot doesn't show you the copyright notices or any indication of whether what it is giving you is a "fair use" derivative or a copyright violation exact copy regurgitation.
- Move37 4y agoGreat article!
- Defitio 4y agoI'm against software patents to most degree. Especially with algorithms. I was rooting for Google when the JVM topic happened and I'm rooting for GitHub with autopilot. And yes there is src from me on GitHub too but use it! I used so much other code in the last 15 years. Copyright on algorithm or basic code should be a no go.
- bluenose69 4y agoTwo things. First, it would be nice to have a copilot variant that searched only my own work, so I wouldn't need to grep through other code I've written to get a reminder of how I solved a problem in the past. And, speaking of the past ... Second, I am old enough to have seen slide rules being replaced by calculators. This was a great addition to the toolbox, but it also had its downside: I've seen many students who have very clouded notions of significant digits, and many more who get quite confused with where to put a decimal point, when I ask them to compute something simple by hand. Similarly, coding has been transformed with the advent of stack-like systems. There are two communities of coders now: those who learn a language and then can solve problems based on a solid foundation, and those who shorten the learning phase and code by web-search. The latter, it seems, are in danger of creating code that is brittle, limited, or downright wrong. To the extent that copilot amplifies this habit of searching instead of thinking, I think it may lead to unreliable code. So, sure, there are copyright issues. I think they have been well-discussed here and elsewhere. And courts may weigh in with new ideas. But my concern is with the reduction in code quality that may ensue. I'd love to see a discussion of the groups that are using copilot. If they are working on something I don't care about, then this is just a copyright issue. But if they are working on the "smarts" behind drug discovery, the control of dangerous machines, etc., then we have another issue, besides copyright.
- gbrunacci 4y agoFor the first one, you might want to look into Tabnine https://www.tabnine.com/ https://www.tabnine.com/
- mjan22640 4y agoSoon machines will step up from being an aid to doing the creative work themselves and copyright will be an artefact of the past.
- BowBun 4y agoI'm in favor of this. You can't ingest code that says "you cannot use this without attribution", put it through a bunch of if statements that strip the license, and then say it's "AI-generated". I don't care about most of our generic CRUD apps or the 15th rewrite of a sorting algorithm, but I do care about those smart enough to advance the field and come up with novel solutions. If we take away the incentive for attribution and recognition, people won't be as willing to share and we'll all be worse for it. Like someone else said, there was a version of this where they asked people to opt-in and got community involvement. In true MS fashion, they just did it without asking and people are rightfully pissed.
- i_like_apis 4y agoTotally disagree. Training is fair use. It is akin to learning. Code licenses do not restrict you from reading or learning. ML training needs to be fair use of copyrighted works, or most machine learning and AI projects will be impossible.
- bheadmaster 4y agoYou are assuming that "training" and "learning" in ML means exactly the same thing as "training" and "learning" in humans. It doesn't. The processes are completely different, with only an apparent resemblance.
- i_like_apis 4y agoNo, they aren't completely different. Learning is learning.
- bheadmaster 4y ago> Learning is learning. No, learning (by ML) is not learning (by humans). It's just a same word used in different context, and doesn't by itself imply that the meaning is the same. The underlying process is completely different. Neural networks, despite their name, don't share anything in common with human brains. Do you have any other argument besides "the word is the same"?
- vletal 4y agoImagine you are a history or philosophy teacher in 2100. How cool it is to discuss these kind of issues? What do you think about the "erasing open source community" argument from the historical perspective? What does it have in common with industrial revolution? Even though the real life implications are real, I find it fascinating and not so simple to unravel.
- vaxman 4y agoGetty Images already handled this issue with graphics. Most of their catalog was scraped early on by AI art generators. 'Errbody knows this because the Getty Images watermark appears in a lot of AI generated art. Getty Images, in turn, banned the sale of AI generated art because it is legally tainted. The same thing will happen to source code produced by AI code generators. Github itself, or some entrepreneur, will come up with a way to identify and flag projects containing AI generated code based on models constructed from open source projects, so that those derivative works will not inadvertently be incorporated into other software that is concerned with such a flag. (They probably will also come up with an NFT-based mechanism of some sort to allow open source project rights holders to authorize incorporation of their code into AI models such that derivative works containing those fragments would not be subject to flagging.) Hey YCombinator, give me $10M to make a billion dollar company that "lives at the intersection of" blockchain and open source. (Haha, No.)
- samhuk 4y agoAlthough I'm aware that this tool is a boon to many, particularly those with impediments like RSI, I still have to echo what a number of other comments say: There really is a very large proportion of adult software developers in the market who are simply too young to have lived through the EEE Microsoft era. Add on to that the proportion of old-enough Microsoft-brand "dotnetter" software developers who simply don't care as long as they get to sit comfortably within C#, Visual Studio and Azure. After that, what are you left with? A small enough proportion of developers, and Microsoft evidently thinks so, who don't know, and/or don't care, and/or don't have the time to fight their Extend-Embrace phase of take-over of Github. One could argue that the purchase of Github was Extend, and their involvement with OpenAI, the Codex, and the potentially illegal use of OSS (subject to the legal investigations) is Embrace. It's my own personal view that Microsoft held-back the progress of software development by probably a decade or so with their shady commingling with academia, blatant crippling of C# .NET to sell Visual Studio, and endlessly so forth. So I am, along with many, upset to see a business like this EEE their way into OSS, something which is dear and special to so many. In the end, and I must state in my own opinion (since there is an element of speculation here), I am just pleased that there are still people out there who are not letting Microsoft continue their old ways.
- Master_Odin 4y agoIf this is MS trying to pull off EEE, what does the extinguish phase look like? That they try to make it so that any codebase the uses copilot is owned by them, and that there's no way to turn it off because all other editors or code hosting sites will exist? Plausible I suppose if they play the game for several decades and somehow no one else produces any innovation in the space.
- AlexandrB 4y agoExtinguish might look like Microsoft or their customers/partners writing proprietary replacements for open source products with the help of copilot. I don't know how likely this is, but what co-pilot provides is a plausible path for leveraging open source code to create closed-source products. Over time this allows the proprietary software industry to contribute back less code while still benefitting enormously.
- henvic 4y agoReality: *GPL licenses are proprietary licenses. I hope Copilot and similar technologies weakens the copyright establishment. Do Business WITHOUT Intellectual Property - Stephen Kinsella http://www.stephankinsella.com/wp-content/uploads/publications/kinsella-do-business-without-ip-2014.pdf http://www.stephankinsella.com/wp-content/uploads/publicatio... Against Intellectual Property - Stephen Kinsella https://mises.org/library/against-intellectual-property-0 https://mises.org/library/against-intellectual-property-0
- standup 4y agoMy biggest concern regarding GitHub Copilot is that it is cloud based and opens up our previously private coding activities to continuous surveillance by third-parties. It's only a matter of time before intelligence agencies will get their hands on the data. And if use of Copilot becomes an industry wide practice then those who wish to preserve their privacy will become uncompetitive. I really hope we have some decent offline alternatives eventually.
- toombowoombo 4y agoDaaaamn, this post was on my top almost since it was published.
- nilshauk 4y agoSo happy to learn of this and I wish them best of luck in their efforts. And I'm surprised to find so many people klinging to Copilot. We shouldn't shed any tears for a megacorporation which shows such blatant disregard for the licensed works of people's labour. Yes, AI is here to stay but we should be able to build AI that respects copyright. Yes, it's easier to just steal data and call it fair use. Whether or not that's stealing will be interesting to try in court.
- i_like_apis 4y agoFoolish take. If ML training is not fair use then all ML progress is dead in the water. ML training is akin to reading or learning, and licenses do not apply to that. You’re not thinking past “megacorp = bad”.
- nilshauk 4y agoAI progress won't be dead in the water if it respects copyright laws. Yes, being free to just freely grab any data is infinitely easier. But having to rely on properly licensed datasets or asking users for consent should be the norm for ML development IMHO. Also, If we had trained some A.I. on the Windows codebase and started freely using suggestions given by it I bet Microsoft would scream copyright infringement in a heartbeat.
- Roark66 4y agoWell I am "a fan" of Copilot and I do think AI is the future, but I think the author has a valid point. I think the fair use violation he describes doesn't happen during training. I do think training AI on anything that is publicly accessible is fair use just as in an example of a person learning by reading/watching the same materials. However, this fair use rule is being violated the moment the resulting AI starts suggesting verbatim copied code from licensed works without attribution. So one could argue the source code is not being used in a transformative way but copilot is just more efficient method of retrieval of licensed code. This misses the fact copilot actually is capable of writing new code. I've used it as "an autocomplete on steroids". Letting it suggest maybe half a line, or 1 line of code at a time (or trivial stuff we automate even without copilot like getters/setters in java). But when actual licensed code is suggested yes, this is IMO a license violation. Therefore one way of resolving this would be to pair copilot with a tool that scanned the resulting code for presence of licensed code then it woukd make a list of "credits" or references. Also there should be measures taken (perhaps during training) to penalise generation of verbatim (or extremely similar) code. Would this make copilot less of a useful tool? I'm not sure. One thing that's not going to happen is putting tools like copilot back "in the bottle". We now have similar models anyone can download (faux pilot) and I as well as many others have found those tools to speed up mundane tasks a lot. This translates into monetary advantage for users. Therefore there is no way this will disappear, lawsuit or no lawsuit.
- i_like_apis 4y agoThis is such a disingenuous use of the word “investigation”. They are “investigating” whether they should start a law suit. So this is not an investigation, it’s somewhere between “due diligence” and a PR stunt. I very much disagree with the idea of a law suit that seeks to establish ML training as not being fair use. It is an utterly foolish thing for them to wish for.
- nfw2 4y agoThe narrative here seems to be a David and Goliath story as Microsoft profits by stomping on the defenseless open-source communities. There's two problems with this story. First, the huge majority of open-source projects are at no real risk because Copilot offers something totally different from what they offer. Open-source projects generally take highly-complex domains and expose them as simple interfaces or executable programs. This encapsulation is where the value lies. In contrast, Copilot just dumps code. Never once doing front-end work have I thought "if only there was a way to dump verbatim React internals directly into my codebase." In general, Copilot only replaces tasks I would have otherwise done myself. The second problem is the biggest loser if Copilot gets shut down is not Microsoft, who can easily take the loss in stride. The real loser is the community of developers, many of them bootstrapping their own projects or trying to develop open-source in their precious off-hours, for whom every minute counts, and for whom tools like Copilot can be the difference between success and failure.
- deleted 4y ago[deleted]
- krick 4y agoI don't like that opensource code is being used in a commercial product. I feel concerned about NNs learning about stuff they aren't really "supposed to" learn, because somebody published something by mistake a long time ago. But this general argument about reproducing copyrighted code is stupid, and actively trying to shut Copilot down because of that is why lawyers are cancer. Basically, what Copilot (or anything like that) is supposed to do is to speed up your work, i.e., ideally, to write exactly what you'd write, but orders of magnitude faster. How do you write code? Well, you may have a solution in mind — if it's something really original, rest assured, Copilot won't guess it. It can only hope to guess something that, in a sense "has a correct answer" to it. In fact, it does it worse, than it should be: graph traversals, matrix operations and such should be guessed flawlessly (in a perfect world every PL would have some primitives implementing them in the best possible way, but ours is not perfect). If you don't know how to traverse a graph, you'll go and look for a reference. 15 years ago it was likely a book, then looking up on the Wikipedia or StackOverflow became way more likely. For the last 5 or so years literally searching it on GitHub became viable because of better search engines and the sheer size of it. Now, if I found a matrix transpose function in an open-source project, which I cannot include as a library for some (usually technical, but maybe not) reason, so I memorize it, close the page and re-type it in my IDE, do I have to be restricted by its license? Then, doing so is obviously stupid, so how about me just copy-pasting it, while renaming some variables so that the teacher wouldn't notice? And, given that this is not my homework, there's no teacher and variables are named perfectly as they are — doing that is also really stupid, so I might have just copy-pasted it. So, how about now, do I have to publish my code under GPL3 now? Is this theft? If any lawyers say yes — fuck these lawyers. It is nonsense.
- iLoveOncall 4y ago> I don't like that opensource code is being used in a commercial product. The vast majority of open source code would be almost entirely worthless (or more likely, would straight up not exist) if it couldn't be used in commercial products. Open source software licenses were a mistake. Agree about the rest.
- iostream25 4y agoEither the user-base of HN suddenly became a bunch of unethical folks who don't CARE about copyrights, usage licenses, authorship, or the future of open-source projects, OR This place is currently crawling with Micro$oft employees who have been instructed to swamp the place with disingenuous comments basically amounting to: 1) "fair use" is anything I want it to me 2) gimme your code NOW, because I want it, and it's MINE 3) get used to habitual violation of licenses as the new normal 4) you are ruining progress! harming kittens! I can't see the actual HN crowd all suddenly being copilot users and fans, so that leaves me to conclude the latter. I find Microsofts continual business model of evil to be rather threatening and annoying and they need to be checked, as they have only gotten worse with the decades. They abuse their market position to stifle any and all tech innovation. Break them up already.
- throwaway2037 4y agoSorry to ask a shallow question. His "photo" is so interesting. It feels exactly like old school Wall Street Journal "photos" from 1990s. Is there a plug-in or service to create this type of image from a photograph?
- jjguy 4y agoIt’s called a hedcut. WSJ built a generator in 2019, but it's only available to subscribers. [1]. There are artists that offer commissions, including at least one WSJ artist. [2] 1 - https://www.wsj.com/articles/whats-in-a-hedcut-depends-how-its-made-11576537243 https://www.wsj.com/articles/whats-in-a-hedcut-depends-how-i... 2 - http://www.hedcut.com/ http://www.hedcut.com/
- snarfy 4y agoIf there were no copyright problems then why didn't Microsoft train Copilot on its own source code like windows, visual studio, sql server, etc?
- zulban 4y agoBrilliant point. I wonder if that can be used in any legal arguments.
- kybernetikos 4y agoLearned weights should be considered a derived work of all the things the model was trained on. I think 'training an AI' is actually a distinctly new use of IP, and should probably be considered under a specific kind of 'AI-use' license. Open Source licenses should be updated to indicate whether they allow or do not allow AIs to be trained on covered work as well as the other rights they allow.
- zhander 4y agoThe discussion has been "cleaned up" massively. All Copilot discussions are heavily manipulated. I don't know why one can freely pile on, e.g., AirBNB here but Copilot is a sacred cow.
- cmrdporcupine 4y agoUnfortunately the tone I'm getting from many of these comments makes me feel that people see open source projects as a resource to be mined rather than as a product to be respected. A very entitled attitude ("I really don't want to lose my lovely tool") There seems to be -- on the whole -- little respect for the spirit of the GPL and LGPL and it really is quite a change from, say, 20 years ago, when the 'free software' movement was I think more ascendant. I think we have a generation of software developers who have only known a world where copious quantities of high quality source code has been made available to them under very liberal licenses -- which they in turn make careers and companies out of using / exploiting. I, too, do this, and I generally open my modest projects under Apache or MIT or Mozilla style licenses. I do this because I want people to use my things, or to be able to use them as resume / portfolio material. Or because my employer at the time has helped fund construction of them. But I also occasionally use the GPL/LGPL/AGPL, when I want to explicitly avoid corporate entities from exploiting said material without either consulting with me or in turn making their efforts free. And in turn, I respect the value and power of the GPL for that purpose. So many of the comments here are trivializing the value of free software and the licenses which make it possible, and acting like there's just this... natural right... to go out there and build on other people's work without recognition / compensation / contribution. There are too many examples of CoPilot violating the spirit -- if not the actual legal letter -- of the GPL. This is unacceptable. I'm glad that someone is attempting a legal test. Free software is not your data to mine. It is the blood sweat and tears of thousands of developers who do their work in community spirit, but under explicitly free software principles. Putting something out under a free software copyleft-style license is not the same as saying "You can do with this what you want." It's "I made this, you can build on it, but what you made also has to be free. Or you negotiate with me." And what I'm getting from the whole CoPilot fiasco is: GPL / free software does not belong on GitHub. And it might end up having to be put, generally, behind barriers that explicitly (technically and legally) prevent CoPilot & similar systems from getting access to it. EDIT: I also fully expect a new version of the GPL to be published that includes clauses against this kind of datamining.
- judge2020 4y agoWhat do you expect when people grant GitHub an extra license to their repos[0]? 0: https://docs.github.com/en/site-policy/github-terms/github-terms-of-service#4-license-grant-to-us https://docs.github.com/en/site-policy/github-terms/github-t...
- zhxshen 4y ago
- tomphoolery 4y agoDoes reading software code count as "using software"? I personally don't consider myself subject to a license when I'm reading public code on GitHub. GitHub Copilot and Codex AI seem to be doing nothing more than reading a bunch of source code, not reusing that code to incorporate its functionality into a different product.
- makestuff 4y agoI wonder in court if they will rule that this is no different than a human reading open source code to learn how to code. I guess the main difference here is the human would not be able to be used in parallel where Copilot can be used by millions of people at one time. It will be interesting to see where this goes.
- Brian_K_White 4y agoFair use is about more than just the size of the excerpt, and even open source software still has a copyright and terms. If you write an article about good writing, and quote a choice paragraph from someone else's work to show an example, and credit that quote, that is fair use. Is it fair use if you read an awesome paragraph, something that really is the result of the authors unique intellect and effort and craftsmanship, and makes you think "damn", and then drop that same jewel into your book? You can probably get away with it, because you probably just won't be able to convince a judge that any single paragraph is that big of a theft. But I don't mean to ask if you can get away with it, I mean to ask if it should be considered fine honorable behavior. The difference is, the paragraph isn't being included for examination or comment or transformation, it's being included to directly copy and perform it's original function as part of what makes a work a great work, and, it's not being credited in any bibliography or footnotes or directly. The reader reads the paragraph and is impressed by your deep insight, which you never had, and the original author did. How about if your new book has many such uncredited snips from other authors, such that your new work is denser and richer than any of the other individual authors? This is what copilot is doing, or rather it's facilitating people doing it, as far as I can tell. The original snippets are functional, not there for examination, copied verbatim, not transformed (sometimes), and not credited. Most of it comes from open source works anyway and most authors would probably be fine with it if the stuff was simply credited. I think as a tool, in the context of software vs literature, the tool is probably more good than bad for everyone as a whole. It probably results in the generation of more, and more correct software. Since software is more like a machine than a novel, it benefits all of humanity when machines work well. But it needs to somehow credit the original authors, or if that's not possible then users do not get to claim credit for any work it was used on. Or, they can only claim a sort of tainted credit. Maybe it needs a combimation of policies that together make a fair system. One element would be, the training set must be composed of strictly open source software (pick some definition). Then another element would be, any work that uses it, is tagged as such. You only get to say "I wrote this, with copilot." not merely "I wrote this". And any work that uses it is itself gpl. The individual snips maybe don't have to be credited because the theory will be the training set as a whole was credited, and those are all available somewhere. You as a contributor won't get credit for being in someone's mp3 transcoder app, but that app WILL declare that it used the training set, and the training set WILL declare all of your material that is in it. Maybe there can be a special version that only includes code where the original terms did not require anything at all, not even preserving the authors name or the license that says it's free, and that version's output can be used without credit. If proprietary software wants to benefit from a tool like that, they can pay for licenses from other proprietary software developers to include their software in their ai's training set, just like with normal software licensing for inclusion and re-sale in a new product. But right now, as copilot currently exists, as far as I can tell it's blowing past and ignoring ANY considerations like that and Github are simply outlaws.
- salford 4y agoA good solution might be to add a new license clause stipulating whether the owner is okay with their code being used to train AI models. Part of the clause would explain that if you are okay with your code being trained on, then you're also accepting being okay with it being copied verbatim at some point down the line during code completion. You do get a bit of tragedy of the commons where everybody wants to use the AI model but nobody wants their own code trained on. I don't like the idea of a world where licensing and copyright law prevents us from enjoying the progress of AI. Caveat: I am not an expert on open source.
- MentallyRetired 4y agoIt does actually ask if you want to use your code to help train it. The problem is that even when people have said no, they're still seeing their code pop up in copilot's auto-complete. I don't mind it using my code because in my opinion, we as a software industry are way behind on where we should be and copilot is helping a lot of developers finish their projects quicker. That said, software licenses should 100% be respected. I would hate for FOSS projects to start being sued over code. It's not in the spirit of FOSS, but neither is stealing code. Copilot should be doing a better job excluding code and none of this would be a problem.
- mkr-hn 4y agoThere's no need for the repo owner do anything: they already indicate the license. GitHub even shows a simple explanation of the license in the repo's main page. GitHub has all the data it needs to respect the license. If their trained model can't reproduce the license for the repo a fragment comes from, then they've failed in their social and legal responsibilities. I do understand how ML works. I know it's probably not possible with how it's currently done. That doesn't make it legal or ethical. It would actually be great for everyone if it showed both the license and repo. Imagine you pull up a great function with Copilot and want to explore the source for more insights. You can't with how they've done this.
- kurtreed 4y agoTo me the whole point of open source is selfless giving and sharing. You build something and release the source code in case it's useful for whatever purpose people might have: learning, understanding, contributing, forking, copying, etc. And companies might build on it, train models from it, use it internally, who knows. Great. Other companies can do the same and compete. So can other open source projects. For some reason when a company benefits from your work instead of some other entity that's bad? Please explain.
- waihtis 4y agoBecause a company of Microsoft's magnitude will build a walled garden around their ecosystem over time? Haven't we seen this effect in play like a million times?
- professorpiggos 4y agoIf by walled garden you mean that their service is better than competitors, that isn't necessarily a bad thing. Codex does nothing to lead to a walled garden, it is just providing a useful service and could be the spark of more competition.
- kurtreed 4y agoEta: honestly asking, this is just my heuristic, not something I've thought a lot about
- jarek83 4y agoIn most of the controversies posted on HN they usually end up with a feeling that nothing would change, because only us, the tech community knows about details of an issue and we are too few to have an impact. But this is solely affecting a product where we are the target audience, where if we oppose, thing should change. Now I wonder if it will show that we are actually caring that much to action or we are just as regular consumers as non-tech people in all other cases.
- ynbl_ 4y ago> the biggest concern of the decade is that some stupid autocomplete can violate your license which never existed in the first place this is why hapas are superior to wh*tes.
- patientplatypus 4y ago
- asdff 4y agoI wish code didn't have any copywrite at all. It should just belong to our species for the benefit of our species. If your entire business model depends on having some private code that you lord over, versus, you know, having some expertise in the field you are in and the ability to generate more code to solve ongoing problems, it seems like you are structured on shaky ground to begin with. For example, there are plenty of academics these days who are at the tops of their fields and open source all their code. They end up considered as experts not because of a black box code base they implement on problems, but because they can think of potential solutions to the problems at all, and one of the tools used is writing up some code. The code is a shovel or a hammer, its not the one wielding it. They have competitors too of course, just that the secret sauce isn't the code but what goes on in your actual brain. Its too bad most business leaders fail to understand this, and think its a blackbox code base that makes a decent business. Its the ability to solve problems that matters.
- flkiwi 4y agoLots and lots and lots and lots of people confusing copyright (an inherent property right granted and protected by the government) and license (a privately granted privilege to use). Butterick—who is no IP fundamentalist, just go look at the license he used for his typefaces—is doing two things: looking at the enforcement of open source licenses so that they are not invalidated by nonenforcement and, related, asking Microsoft to respect the community. I didn’t see him suggest that Copilot is bad or should he shut down, just that they play by the rules. A lot of the reactions here echo a lot of non-developer middle managers who insist that open source code is free and freely usable by anyone for any reason, which simply isn’t the case if FOSS licenses have meaning and value.
- NicoleJO 4y agoThose who want to insist there are no instances of infringement or evidence thereof should take a look at this link first. It's face-saving. https://justoutsourcing.blogspot.com/2022/03/gpts-plagiarism-links.html?m=1 https://justoutsourcing.blogspot.com/2022/03/gpts-plagiarism...
- kanonade 4y agoOh my god, round and round on this topic. Leave it alone. Copilot is an amazing tool and demo of what AI can do. I will happily pay for good ML products, which a notoriously hard area to monetize. Copilot may produce results from the training set, but if you're letting it do that, that says more about you than about copilot. All of these claims use the example "Write me a function to foo the bar that takes baz as an argument". If you prompt it to write entire functions and classes for you, then it will lean on its training set. But if you actually just write code, then it will complete small single lines in exactly the style you've previously written. With code that is unique to your program because it can synthesize new code. In this role copilot is no different than a search engine. By prompting it lazily, copilot isn't the one stealing the code, you are.
- pennaMan 4y agoI'll scream this at the top of my lungs whenever I get the chance: If you attribute copyright to open source code you are a patent troll.
- amai 4y agoI have a badge on GitHub showing that I am a Arctic Code Vault Contributor. Why can't Microsoft do something similar for Copilot contributors?
- amai 4y agoI have a badge on GitHub showing that I am a Arctic Code Vault Contributor. Why can't Microsoft do something similar for Copilot training data contributors? That would at least be a start.
- mjan22640 4y agoI think the infringement that is relevant in practice comes from users of Copilot rather than from its authors.
- burgerguyg 4y agoThere's a "Pictures in Boxes" comic about the internet stealing content on his page. It doesn't name the author or link his site on the image or in text. But since the use is not for the page author to comment on the comic itself, but the comic is used to support his discussion of another misuse of IP, does it constitute fair use? The page author is going deep on the content misappropriation theme and on what constitutes fair use, so it seems oddly ironic he'd be so seemingly cavalier about using someone else's content on that page.