8 ms·
The larger issue is that anyone using GitHub is donating their work for re-use without attribution through Copilot. In return, you receive hosting from GitHub.
by 37ef_ced3 5y ago
The larger issue is that anyone using GitHub is donating their work for re-use without attribution through Copilot.
In return, you receive hosting from GitHub.
The writing is on the wall. You MUST host your own code on a stand-alone website.
Here is an example. This Go program (a compiler) generates and serves its own website: https://NN-512.com https://NN-512.com
It runs on a Linode shared CPU cloud instance that costs $5 per month: https://www.linode.com/pricing https://www.linode.com/pricing
Another example, look at what Fabrice Bellard does: https://bellard.org https://bellard.org
- wellthisisgreat 5y agoIs the copilot thing true for paid accounts as well?
- judge2020 5y agoIf you treat copilot like a human, and you ask it a question ( like "#find median of list") that human either (A) will put together all the things they've learned over thousands of hours of simply looking at code in the programming language and how the syntax works, then provide a new code block, or (B) remember "oh, I remember this extract string of text and what comes after it. Here's the next 10 lines from that snippet". In that B scenario, would the human be infringing on the original's copyright? Arguably yes; but is situation (A) also copyright infringement, or is it like getting code help from a friend? In these situations, whether it be originating from a human or robot, the knowledge comes from looking at public code on GitHub and it's always been a risk that your public code might not be used with proper attribution at some point. Think about how many OSS projects have core code and algorithms copied daily by companies with no public name and keep all of their source code private - it's surely caused more damage than CoPilot ever will.
- pbhjpbhj 5y agothe following is my personal opinion, unrelated to my work >In that B scenario, would the human be infringing on the original's copyright? // Depends on jurisdiction but in general I'd say no as quoting, particularly an insubstantial portion, is allowed (certainly under USA Fair Use, possibly under UK Fair Dealing). However, and again depends on jurisdiction, copying a whole work to use for AI training or searching can be restricted (though it's on the GitHub license in this case). The snippet might be allowed but copying into a training corpus may be excluded (I think UK have an exception for AI training, but I may have misremembered, it might be a suggestion??).
- fault1 5y agoI think the answer is: we don't know the copyright consequences of copilot yet. It's certainly a legal gray area, and it has not been challenged in court yet. There are quite a few companies where copilot and copoilot-like technologies seem radioactive at least for the moment.
- chongli 5y agoAs a human, if you memorize a book word-for-word and later reproduce whole passages of the book in your own writing then that is considered plagiarism and copyright infringement. It does not matter that you stored it in your human memory before reproducing it. The test to use is to imagine you’re writing an essay in college. If your essay contains unattributed passages from another work then the professor will not care that you memorized them rather than just copying and pasting the text. In order to be in the clear you need to properly quote and cite the original author. GitHub copilot does not provide attribution. All it does is obfuscate the original source to make attribution impossible. Copyright lawyers ought to have a field day with this one.
- 999900000999 5y agoTo be honest, can't humans go through various open source repositories and copy the bits they like? If you don't like that,don't release your code as open source. I'm not going to host my open source projects on my own website. That's just hard, it's much more difficult to get people to check it out if it's on Broblog.net I personally don't believe extremely short code snippets, like the ones copilot tends to copy are problematic.
- nerdponx 5y agoThis is the same fallacy behind the argument that we shouldn't care about privacy, because it's always been possible to track and surveil individuals.
- 999900000999 5y agoDo you want to write open source software or not? Here. def add( x, y): return x + y I don't want to live in some dystopia where we have dozens of lawyers deciding who owns that above code snippet. If Amazon wants to use that snippet, fine, you can use it, etc. All Co Pilot does is optimize taking code snippets from different sources.
- nerdponx 5y agoI have no idea what point you are trying to make.
- rustc 5y ago> The writing is on the wall. You MUST host your own code on a stand-alone website. How would self hosting your code prevent Microsoft/GitHub from using it in the Copilot training dataset? If using content from GitHub irrespective of their license to train Copilot is legal, so is training from code available on your website.
- nerdponx 5y agoIt's a question of handing it over and giving them a ToS shield to hide behind, versus making them work for it and risking license violations "in the wild".
- rustc 5y agoDoes their ToS give them additional rights to the code uploaded to GitHub? There are several unofficial copies of projects like glibc [1] uploaded by people who definitely do not have the authority to grant any additional rights to the code. [1]: https://github.com/bminor/glibc https://github.com/bminor/glibc
- 37ef_ced3 5y agoMicrosoft owns GitHub and you accept their terms of service. Copilot is trained on public GitHub repositories of any license: https://en.m.wikipedia.org/wiki/GitHub_Copilot#Technology https://en.m.wikipedia.org/wiki/GitHub_Copilot#Technology If they scrape your website, that's different.
- judge2020 5y agoIf it's legally sound to do that, then it's legally sound to scrape code from Stack overflow as well. No special license is granted to GitHub to train CoPilot; the hosting license in the ToS[0] specifically doesn't allow its use outside of GitHub itself[1], so i'd argue that applies to running copilot in VSCode for code not destined for GitHub - and i'm sure MS's lawyers reviewed such a product launch. 0: https://docs.github.com/en/github/site-policy/github-terms-of-service#4-license-grant-to-us https://docs.github.com/en/github/site-policy/github-terms-o... 1: > This license does not grant GitHub the right to .. otherwise distribute or use Your Content outside of our provision of the Service
- marcodiego 5y ago> The larger issue is that anyone using Github is donating their work for re-use without attribution through Copilot. I think this can only be valid of Github's terms of use clearly specify that or the chosen license allows it. I actually see copilot benefiting GPL projects: suppose a programmer uses copilot to develop a proprietary software and copilot regurgitates GPL'ed code: now the proprietary software is a derivative work and must be GPL'ed too.
- anonymousab 5y agoUnless copilot serves as a legally effective "laundering" of gpl code. Which sounds silly, but we're now in a situation where that outcome is super desirable to GitHub/Microsoft. Another potential outcome is that so many projects could now end up unknowingly using gpl code that enforcement becomes an impractical whack-a-mole, far more so than today. Being told to rework or relicense your project because of copiloted gpl code could easily end up with hobbyists wholesale begrudging the gpl license rather than copilot itself.
- fault1 5y ago> Unless copilot serves as a legally effective "laundering" of gpl code. It's quite interesting how it seems to have come full circle, at least according to this stackoverflow comment about Stallman releasing a paper in 1992 on how to effectively launder AT&T code via textual changes (I have no idea if he did, I just remembered the comment): https://unix.stackexchange.com/a/591221 https://unix.stackexchange.com/a/591221
- nerdponx 5y ago"Not Github" is also a good place to start. I am a happy Sourcehut customer. Roughly the same cost as a VPS, but comes with tech support.
- smorgusofborg 5y agoHow does your license prevent someone redistributing source code over GitHub? What does copilot claim as far as licensing rights beyond fair use? You may be practically limiting copilots use of your code but I don't see any licensing difference if Microsoft hosts Copyright code or scrapes copyright code.
- deleted 5y ago[deleted]
- peter_retief 5y agoThis seems to be the new reality. Sad we get fooled over and over by the same companies.
- coliveira 5y agoI have good memory. I always stood one mile from anything that MS does. The new generation that is easily scammed by flashy stuff like VS code is in for a treat.
- emilengler 5y ago> The larger issue is that anyone using GitHub is donating their work for re-use without attribution through Copilot. Wrong. Lets say a GPL project is not hosted on GitHub officially. I can easily setup a mirror for it though on GitHub as the GPL doesn't prevent me from doing it... Point is that anyone can put my work on GitHub, even if I don't want to.Assuming the project is under a free license though.
- 37ef_ced3 5y agoYes, you can "donate" someone else's code, knowing Microsoft will violate a use-with-attribution license. You can do it, but it's WRONG. Copilot is trained on public GitHub repositories of any license: https://en.m.wikipedia.org/wiki/GitHub_Copilot#Technology https://en.m.wikipedia.org/wiki/GitHub_Copilot#Technology We must all stop using GitHub.
- edoceo 5y agoMy team has code, MIT and GPL, on GitHub. We know the risk of this kind of theft. We remain on GitHub for the discoverability. It's not so absolute.
- _0ffh 5y agoI wonder what kind of language you would need to add to your license in order to explicitly forbid the ingestion of your code by Copilot and/or like projects.
- ghoward 5y agoIANAL, but I have written licenses for that purpose. [1] (I'm trying to get them reviewed by a lawyer, but can't afford to; maybe I'll do a GoFundMe.) What I did is say that if you feed copyrighted software to an algorithm that itself outputs software, then the license applies to the output. This covers the output of compilers and such, but it would also cover Copilot in my opinion. We'll see what a lawyer says. However, even with a license, I wouldn't doubt that Microsoft would just put it through GitHub anyway because finding them out would be extraordinarily hard. [1]: https://yzena.com/licenses/ https://yzena.com/licenses/
- phendrenad2 5y agoThat "issue" is unrelated. You've just hijacked the thread to talk about it. Also the point you're making is controversial, and definitely isn't widely agreed on. As a human, I can read public code on Github, and use my internal neural network (brain) to regurgitate sections of code, and don't need to attribute anyone (who can say which codebase I'm recalling code from? I certainly can't). So a neural network doing the same thing, but external to a human, is certainly questionanble, but it isn't a cut-and-dry case of copying without attribution.
- 37ef_ced3 5y ago"Unrelated"? Microsoft's automated violation of "no use without attribution" licenses? You can't see the connection? Please. Surely it's obvious to everyone but you.
- deleted 5y ago[deleted]
- Zababa 5y agoYou're making a false equivalence between copilot and a human brain. Neural networks are programs, programs don't have the same rights as humans. Implying the opposite is even more controversial than saying that copilot should respect licenses. Also, humans sometimes have limitations on what code they should have seen. It's common in the emulator community to forbid anyone that has seen proprietary code from working on black-box implementations to avoid legal issues.
- coliveira 5y agoThis is a feature in fossil since the beginning. And fossil is good enough to be used by sqlite, it could be used by other projects.
- CyberShadow 5y ago> Here is an example. This Go program (a compiler) generates and serves its own website: https://NN-512.com https://NN-512.com Not a very good one - clicking the link produces a download dialog on Firefox (I'm guessing because the website neglects to indicate a Content-Type). Ironically this is an argument against trying to do everything yourself - you might waste time chasing the long tail of thousands of little details that had been solved many times over elsewhere.
- y4mi 5y agoIt does work with Firefox on my end... Also hosting it yourself doesn't mean that you do everything yourself. It would definitely spawn build plugins that do most of what you need if this practice actually establishes itself.
- 37ef_ced3 5y ago
- CyberShadow 5y ago> It works for everyone except you. The Content-Type is right there, in the header. I don't know what to tell you. It's not there. https://dump.cy.md/9b2a31fa5184397159fe42a85c197242/1640447625.283727360.png https://dump.cy.md/9b2a31fa5184397159fe42a85c197242/16404476... It is there if I check with cURL, though. Edit: looks like this bug is triggered by the absence of Accept-Encoding in the request. When the server compresses the response, it neglects to include the Content-Type of the compressed content. > You defend Microsoft, here and in other parts of this thread, as they sell the work of open source authors without attribution and in violation of explicit license statements in the code. I did no such thing. Please stop.
- 37ef_ced3 5y ago
- CyberShadow 5y ago
- echelon 5y agoI'm including licensing notices in several of my private repositories to prevent inclusion in ML training data sets. I might put this in my public ones as well.
- petilon 5y agoAnyone who puts anything on the web is at the same risk. For example, ask Google how old Queen Elizabeth is [1]. Google tells you the answer in a big font at the top of the result page. Google sourced the answer apparently from usnews.com but you didn't even have to click the usnews link. Google "took" the answer from them and deprived that website of a click. So yeah, you are donating your work to Google when you put it on a publicly accessible web site. [1] https://www.google.com/search?q=how+old+is+queen+elizabeth https://www.google.com/search?q=how+old+is+queen+elizabeth
- fault1 5y agoisn't that covered under a fair use 'right to quote'? https://en.wikipedia.org/wiki/Right_to_quote https://en.wikipedia.org/wiki/Right_to_quote
- petilon 5y agoIf all of the useful information in a web page is mined and given away by a third party in an automated fashion, such that the copyright holder is deprived of revenue, then that's not the intention of 'right to quote'.
- fault1 5y agoI'm pretty sure Google pays these types of content publishers (like Reuters) very handsomely. In this case, the content distributor (USNews) also went out of their way to SEO their site with microdata[1], so I'm going to guess a lot of their inbound traffic also comes from Google searches. [1] https://developer.mozilla.org/en-US/docs/Web/HTML/Microdata https://developer.mozilla.org/en-US/docs/Web/HTML/Microdata
- kstrauser 5y agoThe following is pure conjecture with zero evidence: I think that’s on purpose, not to “steal” code from GPL’ed softwares’ authors, but to get GPL’ed code into commercial projects. Then Microsoft can say “oops, we were right all along! The GPL is so dangerously viral that you can’t even host your software on the same server!” I know that sounds stupid, unless you were there for the Halloween Documents. After all these years, it seems to me that Microsoft hasn’t done the 180 they portray.
- dragonwriter 5y ago> The writing is on the wall. You MUST host your own code on a stand-alone website. Even if Microsoft currently only uses GitHub-hosted code as an input to Copilot, their theory of the legality does not depend on code being hosted there and applies to any code they can get their hands on. The idea that hosting code publicly at some non-GitHub location is going to keep it out of Copilot is not well justified.