6 ms·
GitHub confirmed using all public code for training copilot regardless license
- randomperson_24 5y agoHere before this blows up
- vladharbuz 5y agoDo you remember when someone made a vim extension that would autocomplete code with Stack Overflow answers, as a joke? Why is it that in 2021 we're taking this kind of tool seriously? Have we reduced our field and craft to something that can just be autocompleted?
- DivisionSol 5y agoBecause 99% of our job is putting together standard logic that should've been included in most language's standard library? Automate the rote software engineering so it's built upon a shared, healthy, secure, etc codebase... Leaving only the harder, more-fun 1% for us to noodle on instead.
- yarg 5y agoThe standard patterns that frequently emerge within languages are often low complexity special cases of potentially more powerful constructs and usage paradigms. I might start with a simple for loop: for(int i = 0; i < limit; i++){ ... } And then decide that I can make it more concise: for(int i: [0 .. limit]){ ... } But I still need the base form, for less likely scenarios: for(int i = 1; i < limit; i <<= 1){ ... } A language or its libraries can support new cleaner forms for fundamental structures and frequently repeated boilerplate constructs - but it increases library/compiler complexity, cannot always be predicted ahead of time and often sacrifices power for specificity and apparent language simplicity.
- shagie 5y agohttps://gkoberger.github.io/stacksort/ https://gkoberger.github.io/stacksort/ https://github.com/drathier/stack-overflow-import https://github.com/drathier/stack-overflow-import https://github.com/james9909/stackanswers.vim https://github.com/james9909/stackanswers.vim
- coding123 5y agoI think a major goal behind software engineering will one day to be models only, rules follow. Right now we write code that "creates" those models and rules and UIs. At some point in the future we won't need auto complete because the only thing we'll be coding is the model itself, and then we'll provide rules (in some format, not a language) and the UI will know what to do. We see a lot of that happening in the no-code movement today.
- LeonB 5y agoI agree. Things are becoming more declarative over time. And it progresses over decades, as it takes that long to work out the right things to be able to declare.
- yarg 5y agoIt's actually a statement about what complexity paths remain outside the realms of what can be ~comprehended by AI (or whatever it does instead), and what still requires the capabilities and creativity of the human mind. As long as you're smart enough that you are still required to at least delegate some tasks to a machine, you're still worth keeping around - even if you're not always busy.
- jxidjhdhdhdhfhf 5y agoAre you seriously asking why software developers would want tools to help them write boilerplate code faster? Is it an affront to our craft to use an IDE over Notepad? Maybe drafters should refuse to use CAD tools because the craft originated with pen and paper.
- ipaddr 5y agoWhy not select a language / framework that reduces biolerplate? I use laravel a lot and boilerplate code is taken care of in the framework itself.
- jxidjhdhdhdhfhf 5y agoWhy would you assume I'm selecting the language of the projects I'm working on?
- Cipater 5y agoIf such a tool were built well and worked well, wouldn't that be a good thing? I think "commoditizing" programming would be great for humanity.
- marcosdumay 5y agoIt's because this time it uses AI!
- chrismcb 5y agoWord processors have autocomplete. Gmail has it for whole sentences. Is the field of writing a novel it screenplay something that can be autocompleted? There are lots algorithms that can just be copy and pasted. Of course many of those get turned into a library. You still have to know how to string those together.
- Tenoke 5y agoDupe: https://news.ycombinator.com/item?id=27769440 https://news.ycombinator.com/item?id=27769440
- NmAmDa 5y agoSorry I missed this. I use HN on mobile client and it doesn't give filter by latest posts when searching so missed this among many other topics about copilot
- chadlavi 5y agoAnyone can search and browse public code and copy and paste it into their own project regardless of license, too.
- tut-urut-utut 5y agoYes, but that doesn't make it any more legal. Good luck copypasting a chunk of GPL code into your company closed source project and getting caught. You can claim that it was proposed by the tool and that you didn't know the license terms, but lack of knowledge of the actual license still doesn't mean that the license is not relevant. I would rather not use such a tool than worry about infringing someone's license.
- cjohansson 5y agoYes but it might be illegal
- kadoban 5y agoYes, and then they can get DMCAed or sued later if they get found out. Is the same true if they use Copilot? Seems like it would be to my non-lawyer understanding of copyright. If using Copilot is playing roulette with a lawsuit, can't imagine that'll help adoption.
- abarringer 5y agoIf anyone ever attempts to acquire your company they'll hire a third party to do a software composition analysis audit and look for improperly licensed code reuse. And you'll be screwed.
- deleted 5y ago[deleted]
- oauea 5y agoWhy wouldn't they use the code you gave them permission to use by agreeing to their TOS?
- dekken_ 5y agoPlease provide a source, that this is legal.
- themanmaran 5y agohttps://docs.github.com/en/github/site-policy/github-terms-of-service#:~:text=4.%20License%20Grant%20to%20Us https://docs.github.com/en/github/site-policy/github-terms-o...
- dekken_ 5y agoI meant legislation. One set of terms doesn't invalidate another, at least in EU.
- roywiggins 5y agoFirst off, if this were a good argument GitHub might have made it, but they haven't, have they? They're claiming fair use instead. If they want to make the argument it's covered by their TOS then we can consider it. Secondly, if I release my code under a license that doesn't permit this sort of thing, and someone else publishes the code on GitHub, I haven't agreed to GitHub's terms and the person who did can't give GitHub any more rights over it than I gave them. Not every piece of code on GitHub was written by people who agreed to the TOS, it's quite easy for open source code to move around. Suppose the maintainer of a project moves it from GitLab or whatever, with your contributions intact. The maintainer only has your code because you contributed it under the AGPL or whatever- they can't "license" it to GitHub under more lenient terms than that!
- sundarurfriend 5y ago[1]: > You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video. > This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service The core of it (IMO, IANAL) comes down to "as necessary to provide the Service". It claims a right to "otherwise analyze it on our servers", but the clause "as necessary to provide the Service, including improving the Service over time" likely still applies to that. So it comes down to whether Copilot still counts as part of their "Service" that the user agreed to when using the site, or if it's different enough that it's unreasonable to apply that proviso. [1] https://docs.github.com/en/github/site-policy/github-terms-of-service#:~:text=4.%20License%20Grant%20to%20Us https://docs.github.com/en/github/site-policy/github-terms-o... (thanks themanmaran)
- randomperson_24 5y agoSo, if any public code is allowed and we let it slip through. Google can say for example see all text on Google Docs and __sell__ a solution that generates text for you. Is this allowed? What about images?
- chadlavi 5y agoGoogle docs aren't public
- jdavis703 5y agoYes it is allowed. This is in fact how GMail’s autocomplete came to be. Of course this raised all sorts of interesting problems. One of the funnier ones is when users typed “happy” it would suggest “thanksgiving” since “happy thanksgiving” was the most common phrase starting with happy in their email corpus.
- mherrmann 5y agoCould a solution to the license problem be that the auto-completion also shows you the code's license?
- ipaddr 5y agoThat's a solution with more problems. Different liceases with the same code would be a conflict. Different liceases making a bigger function would have mix liceases. Probably would have been smarter to just use bsd. Perhaps they tried that and failed Is it too late to rerun the model? It feels like it.
- jjoergensen 5y agoGoogle has trained their web search on much of the internet. Is it problematic?
- ipaddr 5y agoIf google is generating content that you will sell as a book yes. Legally you could get sued. Remember google uses the content under fair use for there service. You can't claim fair use when you try to sell someone elses code.
- jjoergensen 5y agoI agree that what Google does is considered fair use. They show your content as snippets and they gain knowledge about topics and improve their search engine based on your content. I don’t see a big issue with what Github is doing either, as long as it’s not a near duplicate of the original content.
- JohnWhigham 5y agoI really don't know why people thought a README file is going to stop one of the largest companies on the planet from slurping up all its hosted code and doing what it wants with it.
- sundarurfriend 5y agoDoes the licence make a difference regarding whether or not Copilot code is legal to use? To my understanding, the crux of the argument comes down to whether data crunched down and regurgitated by machine learning algorithms retains its licence - regardless of what it is. I suppose if we assume it retains it, then TFA's claim creates a further question of how the hell anyone using copilot can conform to this unholy mixture of all the licences. But that's a big assumption at this point.
- ipaddr 5y agoThe output is legally toxic to any company who uses it and opens up legal avenues to delay your software release in court. Github may be legally okay because it shifts the legal issues to their paying users. A legal case could be made that they are assisting in ip thief by selling a product that makes copyright breaking easier (we have seen those cases) the majority of legal issues will be on the user. Imagine the Google / Oracle legal battle replayed with this tool involved..
- deleted 5y ago[deleted]
- mullikine 5y agoThe cat (this technology) is coming out of the bag one way or another. It's just too useful. Where is it written that inspiring a language model with data is not just as infringing as copy and paste? It's a very grey line. Public facing open-source code & media is going to be learned by language models because they're exposed to them. I'm fully expecting that if I begin a story and put it on my blog or on github, and if I go away for 5 years, I'll see it completed for me when I return.
- deleted 5y ago[deleted]
- shireboy 5y agoIs there nuance here in what is used to train vs what is used in completions? For example, if all public code was used to train a ML model, but then the autocomplete feature only pasted in uniquely generated or licensed code, that would be different than if it pasted in verbatim licensed code without attribution or whatever the license requires. It could be like a dev reading public code enough to understand it, but then coming up with her own implementation. Not saying copilot works that way - I haven’t tested it yet. But could be one nuance here.
- Tenoke 5y agoThere's many things you can potentially do at a loss of performance. E.g. train it on all the code initially to learn the general rules better but then finetune only on the subsection that is of the right license or write something into your inference that checks how close to existing code the output is etc.
- qayxc 5y agoI don't understand the argument here. Yes, sometimes code is returned that is a verbatim reproduction of the training data. This can be prevented if need be. What I really don't understand is how some people are complaining about GPL'ed code being used for training. What's the difference between a machine looking at the code and learning from it and a human being doing the same. As long as the code isn't patented, there's no reason why I shouldn't be able to look at GPL'ed code and implement the idea using my own code. In other words, is - according to those who think using GPL'ed code for ML training - every implementation a derived work if I looked at GPL'ed code that implemented the same algorithm? Where's the line that separates plagiarism from original work? Is there even such a line? Does it matter whether the GPL'ed code is encoded in human neurons or network weights after looking at it and if so, why?