12 ms·
GitHub scraped your code. And they plan to charge you
- gdsdfe 5y agoNothing is free people! ... People are outraged by GitHub but nobody is going after Facebook or Google for training their AIs on your personal data. Facebook used your face to train some algos, google your personal emails etc.
- nomercy400 5y agoWhat if you paid to keep it private?
- ericmay 5y agoI mean.. I am? I care much more about Google or Facebook profiling and profiting off of my data (especially when I don’t consent to giving it to them in any meaningful way) than I do letting GitHub do things with code I knew was freely available and that other entities could use in profitable ways.
- void_mint 5y ago> People are outraged by GitHub but nobody is going after Facebook or Google for training their AIs on your personal data. ...what?
- SamWhited 5y agoIt's not about it being free, it's about GitHub taking something you licensed with conditions (ie. attribution or keep this copyright notice and license file, etc.) and blatantly ignoring your license because they know you probably can't afford to sue them for copyright infringement. Open Source doesn't mean you can reproduce and copy the code freely, licenses exist for a reason. Also: of course people care about Facebook et al. (not enough, I'll grant you). Plenty of people complain about Facebook violating privacy every single day.
- gdsdfe 5y agoYour license means nothing if you can't defend it. If you paid for the service they wouldn't dare using your code regardless of the license.
- wizzwizz4 5y ago> Your license means nothing if you can't defend it. It's okay to commit crimes if the victims are poor enough? (And yet what you've said is still true.)
- code_duck 5y agoI wish that I could somehow turn my personal data into IP, but I’m not sure how to accomplish that.
- joe_the_user 5y agoWell, "nothing is free" doesn't mean someone can blatantly violate a license 'cause they have rent to pay. Maybe people should be mad about what Facebook or Google do but that stuff doesn't involve taking stuff outside their terms of use. Maybe Github could try attaching a "we can relicense all your code whenever we want" condition to their hosting but they'd lose all their business.
- goodpoint 5y ago> nobody is going after Facebook or Google A lot of people dislike them and minimize their use. More importantly, we are seeing a bait-and-switch. People agreed on GitHub storing, showing and indexing their code and issues, not using the code for Copilot, regardless of what the fine print in the usage agreement says.
- shakow 5y agoBecause people accept a EULA when they give their data to FB or Google. Github is exploiting a grey area to leverage a big fat chunk of GPL (or other) licensed code in ways that are perceived to be a technically probably legal, but morally very ambiguous understanding of these licenses.
- macintux 5y agoWell, some of us try to keep Facebook and Google out of our private lives, but our faces aren't copyrightable.
- iliekcomputers 5y agoIf you open-sourced code and allowed it to be used for commercial purposes, I don't see the point of being pissy about Github using it, I'm saying this as someone who's written quite a lot of MIT code. (And charging for a product which adds value to your developer experience and needs money to be run is not a bad thing)
- SamWhited 5y agoIf you write MIT code you expect them not to strip your license out in derivative works. This is exactly what license are for and GitHub is blatantly violating it while people applaud.
- dopaminefasting 5y agoMaybe read the MIT license before you grab the pitchforks: "The above copyright notice and this permission notice shall be included in all COPIES OR SUBSTANTIAL PORTIONS of the Software." Reusing a snippet doesn't require reproducing the MIT license. People who publish MIT software know they're basically giving their code out with basically no strings attached. However, GitHub should be careful with the GPL variety.
- deleted 5y ago[deleted]
- JBorrow 5y agoSUBSTANTIAL PORTIONS can mean the core five lines of some key algorithm buried deep in a 1000 line wrapper library with a bunch of language wrappers.
- arp242 5y agoWhat counts as a "substantial portion"? Personally I'd say that a function is substantial, whereas one or two lines would not be.
- ClumsyPilot 5y ago
- deleted 5y ago[deleted]
- seph-reed 5y agoWell, for the most part my code isn't going to do anyone a ton of good. I don't use much in the way of popular frameworks, but I also guess this means I'm gonna be out of a job for not writing "normal" enough code at some point. Time to move on to the carbon age I suppose.
- rikroots 5y agoI don't know - I think a variety of coding approaches may make for some interesting AI code suggestions in the next year or two. I do pity the poor algorithm that has to parse sense into my coding idiosyncrasies.
- seph-reed 5y agoPerhaps. My experience with ML is that it's more like: * a person that follows popular trends than * a person who finds/dissects clever/unique solutions to add to their tool belt Honestly, I'm pretty sure ML hates my guts. Anything I've ever used involving it ends up burying my voice and slowly trying to etch away that parts of me that aren't normal enough.
- deleted 5y ago[deleted]
- st_goliath 5y ago> Hi. I know you’re excited about copilot. > ... > It’s truly disappointing to watch people cheer at having their work and time exploited by a company worth billions. Huh? Over the last few days that I've watched this "copilot" story unfold on various news aggregator sites, I've first seen people point out copyright and other issues with it, then the fast inverse square root tweet happened, and then more articles and tweets like this one and the discussion that we are currently having. But I somehow don't really recall anyone besides the Microsoft marketing department being overly excited about it. Did I miss something?
- ghoward 5y agoThere have been a lot of tweets from developers with access raving about how cool it is.
- dinglejungle 5y agoThe comments on the original announcement here were pretty positive: https://news.ycombinator.com/item?id=27676266 https://news.ycombinator.com/item?id=27676266
- Nuzzerino 5y agoWell, for whatever reason, the HN thread for it is in the top 30 most upvoted threads of all time. That probably counts for something. https://hn.algolia.com/?dateRange=all&page=0&prefix=false&query=&sort=byPopularity&type=story https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
- nomercy400 5y agoHow amazing would it be if you could ask a search engine for a piece of working code based on a short description? That would be exciting tech for me.
- rvz 5y ago> But I somehow don't really recall anyone besides the Microsoft marketing department being overly excited about it What you just saw 3 days ago was a hype driven unveiling of a cherry picked contraption by GitHub, OpenAI and Microsoft. Open source became the loser once again and got taken advantage of this clever trick and will soon become a paid service. (With lots of code that is under copyright of various authors.) Anyone who critiqued the announcement three days ago was drowned out, downvoted and stamped on by the fanatics. I wanted to see those who had access to it (Not GitHub or Microsoft fans) to demystify and VERIFY the claims rather than blindly trust it. Those suspicions by the skeptics were right, and lots of questions still remain unanswered. Well done for re-centralising everything to GitHub. Again.
- andrewjl 5y agoHere's the brutal and ugly truth: why isn't our personal data treated as private property? It's because those who write the laws governing its status either lack the requisite understanding or else practice a form of, to put it mildly, motivated reasoning.
- scrollaway 5y agoYour public source code is not "personal data".
- ineedasername 5y agoCan't you host code on GitHub that is not "free" for commercial use? If GitHub scraped these projects then it's a problem. Otherwise, Into honestly trying to have a conversation on this to understand the objections because I haven't made up my mind but struggle to see the problem. So pease consider the following: if the code was not encumbered by restrictions I don't see an obvious problem with this. Using code or data or anything like that in the public commons for a meta analysis doesn't strike me as wrong, even if the people doing it make money off of that analysis. If I scraped GitHub code and then wrote a book about common coding patterns & practices I don't think that would be wrong. I used the Brown corpus and multiple other written word corpuses (corpi?) Along with WordNet and other sources to write my thesis in Computational Linguistics Word Sense Disambiguation, later applying it to my job, which earns me money. Is this wrong? Public datasets have been used extensively for ML already. I don't see this as much different.
- joe_the_user 5y agoThe difference appears when copyrighted material is repeated verbatim. And because it's obvious Github has no control over how much copyrighted material is being repeated verbatim. And that copyrighted material is intended to be used by commercial companies who copyright their own material and don't want to have their copyright challenged.
- ineedasername 5y agoUnder US rules at least, everything automatically falls under copyright. But if the license allows full use even for commercial purposes, verbatim repetition is no different in this context than if I included it in my own new piece of software. (Of course attributions and distribution of the code & modifications would frequently be requires... with Pilot it's... potentially a gray areas on whether code lumps together with countless other software is actually modified in the traditional sense of the term, but attribution should still be an obvious requirement. And where the other license requirements are strictly required or not, I still think it would be the right thing to do for GitHub to honor the spirit of those strictures, especially considering their entire business model is based on a majority of users trusting their code to them. If they're not going to act accordingly, there's no reason someone couldn't roll their own GitLab instance, or a competitor with more respect enter the marketplace.
- bphogan 5y agoHi HN. Didn't really expect to see this tweet make it here of all places. But that's cool.
- yongjik 5y agoThere may be discussions to be made about licenses, but "to watch people cheer at having their work and time exploited by a company worth billions" is a disappointingly myopic take, especially from a developer. Information that is aggregated and organized for easy retrieval is worth more than the sum of individual bits of information. I thought that was common sense. We might as well complain that billionaire supermarket chains are pocketing all the profit while not growing a single potato by themselves.
- pessimizer 5y agoWe would, if they weren't required to pay for potatoes. Are you making a claim that Netflix shouldn't be required to pay for individual movies because they sell a collection of movies?
- ChrisMarshallNY 5y agoThis is my take-home: > We are obsessed with shiny without considering that it might be sharp.
- deleted 5y ago[deleted]
- tedunangst 5y agoWhat if I don't pay?
- speedgoose 5y agoThen you can't use the code completion tool made by Github using open-source code.
- greatgib 5y agoI'm a big proponent of open source and I'm usually not nice with bad moves of GitHub. For example, i find stupid to use vscode and believe that it is open source when it is a lie. But, in that case, I think that the things that are put to charge GitHub are not right. I think that the idea is nice and it is fair from open source code. Anyone is free of downloading free software and doing something similar, and it is nice. I just find the product itself is stupid, and it is for users to be smart enough not to use that knowing that their is a risk of them being sued for involuntary violating copyright. And GitHub might be at risk if it is a paid service as the companies could sue them back by pretending that they expected the code generated by GitHub to be safe for commercial use. Also, I would think that GH would have abused if they used 'private repo' codes to train their model without permission.
- ghoward 5y agoUnfortunately, just because code is open source doesn't mean that there aren't terms of use attached with it. One of the simplest and most widely used terms is attribution. This means that if Copilot does not attribute code when it copies and modifies it, then it is violating most open source licenses. Full stop.
- greatgib 5y agoMy point is that somehow, it is not copilot that is violating the licenses, but the code that is generated is. So, if you just use copilot to generate random things, you are ok. But if you try to use the generated code for anything: distribution, selling, eventually usage. Then, you are violating the licenses in the same way as you took yourself the parts of code to reuse. It is possible users of copilot that should avoid that or be very careful to check any line produced (that is almost impossible). Also, by itself, copying one or two lines of code can hardly be limited by copyright. But, as we so, copilot can spit big full block of code from existing projects.
- ghoward 5y agoYou have good points, but I would argue that Copilot itself is the entity distributing copyrighted code since there are times when it copies code verbatim. That puts the legal onus on Copilot itself.
- hekec 5y agoOn their website they say that "GitHub Copilot is a code synthesizer, not a search engine: the vast majority of the code that it suggests is uniquely generated and has never been seen before. We found that about 0.1% of the time, the suggestion may contain some snippets that are verbatim from the training set." So it won't copypaste your code. It had just read code from open sources and learned from it - similar to what humans do. So I don't see any problem with this.
- ghoward 5y agoGitHub has 56 million users as of September 2020 (according to Wikipedia). Let's assume that only 1 million of them use Copilot at an average of once a week. That means that every week, there will be 1000 verbatim copypaste of code by Copilot. Then multiply that by a year or more as Copilot gets older. 0.1% may not seem like a lot, but at the scale of Internet companies, it always is.
- rhn_mk1 5y agoFirst, it does copypaste code: https://docs.github.com/en/github/copilot/research-recitation https://docs.github.com/en/github/copilot/research-recitatio... Second, we can't ignore that if someone deliberately tries to make it spit out copyrighted code, the chances are going to be much greater. Why would anyone? Plausible deniability: "I didn't copy this GPL procedure, the copilot gave it to me!"
- einpoklum 5y ago> the vast majority of the code that it suggests is uniquely generated and has never been seen before. Original code in somebody's GitHub repo: int x = y + z; Copilot code: int Eisaa7ha = Wu8iazo7 + Roh0Eesh; Not copy pasted! Uniquely generated! Never before seen!
- macintux 5y ago> So it won't copypaste your code. You might want to check out this video... https://twitter.com/mitsuhiko/status/1410886329924194309 https://twitter.com/mitsuhiko/status/1410886329924194309
- justbored123 5y agoI absolutely despise these lot of super entitle people that have absolutely amazing free services provided by companies like GitHub and Google/Youtube and they get all agro because they dared try to monetized or impose some basic rules that don't really affect them at all in any meaningful way.
- throwaway2048 5y agowhat about the "super entitle" companies that take code and violate its license.
- ghoward 5y agoFor the record, I moved all of my code away from GitHub because I didn't like them. Now I self-host, with all of the work that entails. So I think I have a right to be mad when they do something like this to code I previously stored on GitHub.
- eqtn 5y agoGithub should list all the projects they scrapped the code from to make copilot.
- tasubotadas 5y agoGod forbit somebody profits from the code that you've posted publicly. It's a NET POSITIVE FOR EVERYBODY.
- rhn_mk1 5y agoCopyleft exists for a reason.
- TeMPOraL 5y agoTo the extent Copilot is doing something illegal, or making its users inadvertently engage in illegal behavior, it is copyright infringement, as (most) license violations are copyright violations. Copyright cuts both ways. Free Software and Open Software exist in context, and because of, copyright laws. This means that a person or a company using output from Copilot may be engaging in copyright infringement. In other words, Copilot is enabling software piracy. I might be sympathetic to it, and even consider it mostly positive, but then if companies can use my code ignoring the license, I want to be able to Torrent their products in peace too.
- lmarcos 5y agoA colleague of mine: "I either remove all my (useless) repositories from GitHub or I ask GitHub to pay me if they want to use my code in Copilot". It's not that crazy.
- superkuh 5y agoIf you're hosting at the free github service, or even paid, github did not scrape your code. They just accessed the information on the hardware they owned. HTTP wouldn't have to be involved at all. They could just look at the disks. Additionally, "The third-party doctrine is a United States legal doctrine that holds that people who voluntarily give information to third parties—such as banks, phone companies, internet service providers (ISPs), and e-mail servers—have "no reasonable expectation of privacy."" The above isn't to say I agree with this but just to highlight the dangers of outsourcing and the cloud.
- blibble 5y agobelieve it or not there's more countries in the world than the United States > "The third-party doctrine is a United States legal doctrine that holds that people who voluntarily give information to third parties—such as banks, phone companies, internet service providers (ISPs), and e-mail servers—have "no reasonable expectation of privacy." this is definitely not the case for 100% of the rest of the world
- vinay427 5y agoIt's not even true in the US. It's not a specific law, but rather a doctrine that is more of a vague legal idea as I understand it. There are always laws that don't follow it, such as the CCPA in California or several (narrower) approaches in other states, or even cases such as Carpenter v. US that rejected a possible application of it. This is without even considering more obvious holes in this concept including IP law, healthcare data, etc.
- superkuh 5y agoGithub isn't incorporated and mostly centered in the rest of the world.
- smoldesu 5y agoThe good news is that Github then also has no reasonable expectation for me to use their service. Most developers can just as easily set up a Gitlab or self-hosted alternative with zero friction.
- haolez 5y agoSource Hut[0] is getting more attractive with each passing day, but I'm not sure I can adapt to it's weird e-mail centric pull requests (and I know that this is a standard Git feature, but the UX seems bad). [0] https://sourcehut.org/ https://sourcehut.org/
- ghoward 5y agoThe other thing is that Source Hut is immensely complicated to set up. I wish I could love it, but Gitea was just so much easier to get going.
- sfg 5y agoIs there no licence with any sort of model training clause: "If this licence or the source code it covers is used to train a statistical model, then the model and code used to create the model are covered by this licence (which has terms like the AGPL)"? If not, will anybody quietly slip something like this into Copilot's training data?
- einpoklum 5y agoHey, how come what is essentially the equivalent of an HN comment, only made on Twitter, gets to be an HN story? :-(
- mrkramer 5y agoIsn't Google's BigQuery also scraping GitHub and is making it accessible/available for commercial use.
- mensetmanusman 5y agoIs this analogous to gtp-3 reading every sentence ever written without attribution to all of mankind?
- wizzwizz4 5y agoA little. But all works output by GPT-3 are provided in “source form” to everyone who uses them – whereas lots of the output of Co-Pilot (trained on copyleft code, among other things) is going into proprietary software projects. (Also, GPT-3 wasn't trained on nearly as much writing as that. Even if you ignore lost writing, GPT-3 was trained on a small subset of the 'net.)
- nlh 5y agoI have a genuine question about this whole thing with Copilot: A similar product, TabNine, has been around for years. It does essentially the exact same thing as Copilot, it’s trained on essentially the same dataset, and it gets mentioned in just about every thread on here that talks about AI code generation. (It’s a really cool product btw and I’ve been using and loving it for years). According to their website they have over 1M active users. Why is this suddenly a huge big deal and why is everyone suddenly freaking out about Copilot? Is it because it’s GitHub and Microsoft and OpenAI behind Copilot vs some small startup you’ve never heard of? Is it just that the people freaking out weren’t paying attention and didn’t realize this service already existed?
- moocowtruck 5y agoit's just what the community does these days, bored and have to be upset about something, and being upset at big companies is trendy
- echelon 5y agoBecause the repository trusted by millions is starting to do things we never anticipated. It's growing in ways that are a touch uncomfortable for some. I think some are also beginning to feel an Amazonification happening. We built all the stuff and made it free, but now a company is going to own it and profit off of it. Edit: If we want to prevent this, we need a new license that states our code may not be included in deep learning training sets. Edit 2: if private repository code is in this training set, it may be possible to leak details of private company infrastructure. Models can leak training data.
- ghoward 5y agoI personally have never heard of TabNine until now. Now that I have, I don't want my code to be part of that.
- Lariscus 5y agoYes, and it is also not an OK thing to do for the start-up. They were just lucky that nobody noticed their licence violations.
- 5y ago
- skc 5y agoI think the product is pretty cool, but I wish it had been announced by GitLab instead so there would be less of this brouhaha.
- abetusk 5y agoHere is the relevant portion of GitHub's terms of service (section D.4) [0]: """ 4. License Grant to Us We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video. ... """ Note that the relevant detail is that this applies to public repositories not covered under some free/libre license. I also assume this excludes private repos which might have more restrictive terms of use. GitHub has a section on it I just haven't read it in detail and so maybe the above covers private repos as well. [0] https://docs.github.com/en/github/site-policy/github-terms-of-service#4-license-grant-to-us https://docs.github.com/en/github/site-policy/github-terms-o...
- deleted 5y ago[deleted]
- yangff 5y agoSo.. I can see that this ML model is generating some code exactly same as the original dataset, which definiately a problem. A defect model, sure. Beside that, I cannot understand why the overall idea, using open-source project to train a ML model that generates code would ever be a problem. We human beings are learning as the model, we read others code, books, articles, design patterns... and it becomes part of us. Even the private code, I mean like you join a company, you read their codebase, methodology and it becomes something yours. Copyrights generally not allow you to "copy" the original, but you can still synthesize your own code -- cutting, combination, creating based on whatever you have learnt. The method of how a ML model works is differ from human brain for sure, but I cannot see why this would be a problem, or why an organic would become something superior that what they do is a creation and a ML mode is scraping your code. What is the difference here???? And also recently we saw GPT that generates articles, waifulabs that generates ... waifus... to be honest I cannot perceive the difference since all of them are "learning" (in a mechanical way of human created knowledge.
- dathinab 5y ago> A defect model, sure More like a defect approach, behavior like that is well known(1) to be basically guaranteed to happen with GPT-3 and similar approaches. (1): By people involved in the respective science categories (Representation Learning/Deep Learning, NLP, etc.).
- deleted 5y ago[deleted]
- wildmanx 5y agoThe difference is that it's a judgement call when to include attribution, whom to attribute with how much, and overall whether something is too close to be counted as a copyright or other license violation or not. Intelligent humans sometimes, or even often times, have a hard time doing this judgement call. An artificial intelligence would too, and a somewhat simple ML model (no offense) certainly does. I'm really waiting for this to blow up from the open source license angle. Freely combining code with different license is a hellish undertaking on its own. But already just re-using some, say, GPL code, even staying under the same license, but without proper attribution, is Forbidden with capital F.
- Traubenfuchs 5y ago> I have a SoundCloud and books and whatnot I could promote here. You just did.
- yayr 5y agoTo me it seems, the whole subject requires additional consideration in licensing. It is a little like applying telephone based law to the internet. It will not 100% fit. If the creators interests are not clearly expressed anymore with a license, we need updates to the license texts. Let's look at MIT: ____________________ "Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. [...] ____________________ From the license text alone, it would not be clear to me, why anyone could claim that the OpenAI codex or the Github Copilot would require attribution to any of the used MIT source code to generate the AI model. The AI model is simply not a copy of the source or of a portion thereof. It is essentially a mathematical / statistical analysis of it. Now what about any generated new source? How similar does it need to be to any source to be a copy? At what size of the generated code it qualifies to be a copy instead of a snippet of industry best practice? Where does the responsibility for attribution lie? Should we treat the AI code generation models like a copy & paste program? Usually you cannot really say where the copy came from 100% - how do you know what factors influenced it?
- TeMPOraL 5y ago> Now what about any generated new source? How similar does it need to be to any source to be a copy? At what size of the generated code it qualifies to be a copy instead of a snippet of industry best practice? Let's handle the simplest case first: Copilot can and does regurgitate large pieces of its training dataset verbatim. This is a well-known and trivially demonstrable property of all ML models in this family. Would such exact copy fall under the license of the code being copied? This of course needs to be tested in courts, but my gut says "yes". The problem now is, if you're using Copilot, you may end up with such copied code in your codebase without ever knowing, and this might open you to liability.
- yayr 5y agosee some analysis of the scope of this issue here: https://docs.github.com/en/github/copilot/research-recitation https://docs.github.com/en/github/copilot/research-recitatio... especially: Conclusion and Next Steps. This investigation demonstrates that GitHub Copilot can quote a body of code verbatim, but that it rarely does so, and when it does, it mostly quotes code that everybody quotes, and mostly at the beginning of a file, as if to break the ice. But there’s still one big difference between GitHub Copilot reciting code and me reciting a poem: I know when I’m quoting. I would also like to know when Copilot is echoing existing code rather than coming up with its own ideas. That way, I’m able to look up background information about that code, and to include credit where credit is due. The answer is obvious: sharing the prefiltering solution we used in this analysis to detect overlap with the training set. When a suggestion contains snippets copied from the training set, the UI should simply tell you where it’s quoted from. You can then either include proper attribution or decide against using that code altogether. This duplication search is not yet integrated into the technical preview, but we plan to do so. And we will both continue to work on decreasing rates of recitation, and on making its detection more precise.
- temac 5y agoThe way "AI" works for now, Copilot never comes with its own ideas, as it is incapable of deductive reasoning. It basically just detects from the context then mixes variations of things it learned. If there is nothing to mix (that is if there is a single source), the risk of spitting verbatim is high. But if there are multiple sources and some mixing and some amount of tiny differences, said differences better not be not too trivial because I don't see why we would suddenly drop Abstraction-Filtration-Comparison approaches... So their defense of the like "oh it's fine it very rarely emits verbatim things" is bullshit anyway. That's an answer to a wrong question, at least given the answer is in this direction (would there be ton of verbatim recitation, they obviously would not try to wave away the problem like that -- however we can not conclude anything from verbatim output being rare, despite them stating that as if it a quite central and strong argument)
- ricardobeat 5y agoI don’t understand this mentality. The AI is trained (or at least supposed to be - that’s fixable) on code that was published under open licenses. The “exploited by the man” trope after publishing OSS feels entirely backwards.
- maxbendick 5y agoWhat's hilarious about auto-generating the GPL license is that it's provable Copilot is trained on GPL code, but it's essentially impossible to tell which code it came from. Any legal battle will be strange... Is it enough for Copilot to not regurgitate GPL licensed code exactly? Is it enough for Copilot to create a slightly modified version? Laughably, as soon as slight variation is added, there is so much code in the world that it'll be impossible to prove wrongdoing for HTML or JavaScript synthesis. A model trained on all permissively licensed code on GitHub looks a lot like your own GPL code? Are you sure your code is so unique? Microsoft of course will implement compliance standards as necessary (they genuinely do not want to break the law), but what does this mean for smaller companies and individuals training models?
- lokl 5y agoHow does Copilot avoid training on malicious code? There might be bad actors who would love to have their code scraped for this...
- coliveira 5y agoNewsflash: all open source means that you're already doing free work for the largest corporations in the world! It seems like developers, as a group, decided that it would be better to spend their nights writing free code for FAANG, so they would be able to keep their day jobs. Bezos and friends thank you all. #genius
- stakkur 5y agoGithub is (owned by) Microsoft. This is just an appetizer.
- SergeAx 5y ago> It’s truly disappointing to watch people cheer at having their work and time exploited Maybe it's my information bubble, but I don't see anyone cheering. Currently Copilot churning out rather bad code. I am definitely would not use it. And my prediction about it that it will go like Tesla's autopilot for years.