91 ms·
Copilot sells code other people wrote
- habibur 4y agoWe stand on the shoulders of giants. That had been the way for decades. A newer stack over the older one without much thought. And someone in the future will build even a newer stack over the current ones.
- deleted 4y ago[deleted]
- JacobiX 4y agoIt’s the same problem with those ML models, the other day someone generated a children’s book using GPT3, turned out that there is a real children's book with the same name and a very similar content: The Very Lonely Firefly by Eric Carle.
- bartq 4y agoOther thing I'm worried about: how to retract facts from ML model? I guess it's impossible, you need to retrain from scratch with part X removed from training set. Or... people could invent layered ML models similar to docker - each layer would be marked what data it was trained with. Then at least you'd have some cache of trained model to re-use in next training session. Nasty stuff.
- icoder 4y agoInteresting, it's a big question I've had for a while, how 'original' stuff coming from these AI systems is, and also the distribution of uniqueness over many answers. I haven't dived into it yet, but I find it surprising how little this comes up when these systems are discussed (ie here on HN). Does anyone even know? Can we even check? What if 1 in a thousand, or one in a million outputs is (very close to) something existing? I find this especially relevant when generating faces.
- k__ 4y agoIsn't that what Web2 is all about? Someone creates content for free, and companies monetize it.
- WesolyKubeczek 4y agoThe real Web3 is companies sue original creator for infringement.
- deleted 4y ago[deleted]
- maxbaines 4y agoInitially not thought about co-pilot and other ai generators this way, but now I have I’m finding it hard to ignore.
- seydor 4y agoProgrammers are fine when their creations, pretty much all of tech, resells content that other people wrote for free, but no, not code, that one must be expensive
- onpensionsterm 4y agoThe only one making money here is github. Very few programmers are selling open source code. And programmers are (in)famous for not buying software.
- zx8080 4y ago%s/programmers/tech capitalists/g
- anonymoushn 4y agoI also don't think it's acceptable for TurnItIn to monetize content without paying the authors. My opinion about whether students should have their work stolen and monetized by a company doesn't seem to have much impact though.
- capableweb 4y agoIf GitHub could guarantee that the code Copilot had ingested was only made with OSS licenses, then I don't see what the problem is. But as far as I understand, GitHub trained Copilot on any public repository on GitHub, meaning even if it doesn't have a license specified (so the user publishing it still has the copyright to it), then I don't see how it can be OK.
- hooby 4y agomany OSS licenses require attribution
- saghul 4y agoEven if it was trained with OSS licenses, some of them require proper attribution, which copilot doesn’t do. Now, where the threshold is for substantial derivative work in order to require attribution is an interesting question.
- thelastbender12 4y agoIt is hard to see how verifying licenses is a solvable problem, when licensing for code dependencies can be transitive. For ex - if I copy code from a GPL codebase like Linux and create a Github repository with an MIT license.
- danuker 4y agoYou should be able to choose flavors of the model trained only on public-domain code which does not require attribution, for example. But that would mean Microsoft acknowledging license violations.
- thelastbender12 4y agoSorry, to be clear, I meant even if a Github user asserts their code is public-domain/no-attribution/unlicensed, they could have lifted it off a codebase that doesn't allow it. It would be tricky for Github to establish the code was indeed original and hence their agreement with the user allows them to train their models on it.
- jarenmf 4y agoI guess the question is where you draw the line between a derivative work and "learnt by an AI algorithm"
- asimpletune 4y agoWho needs a line when there are plenty of obvious examples lifted verbatim?
- triknomeister 4y agoIf the media copyright industries and their ContentID is anything to go by, it doesn't matter. It's all derivative.
- dgb23 4y agoIs it smart enough to: - respect attribution - respect copyleft - respect proprietary licences - give the user appropriate hints about the above Or does it just copy code without doing any of this?
- spupe 4y agoNo, it doesn't do any of that. However, it does not "copy code" except in marginal use cases, the far more common scenario is that it will suggest you very basic code that is akin to a Stack Overflow reply.
- dgb23 4y agoI read a lot of open source code and might subconsciously absorb techniques and patterns that are common. When I write code I might be influenced by what I read, not line per line, but rather generally. Is it like that?
- spupe 4y agoKinda, but I think you are imagining something bigger than it is. At least in my experience, it works well for simple stuff like "iterate over x and extract y" or similar queries that I imagine are well represented in its training data. When you get to very specific functions, its answer will be less reliable and more likely to be a wonky rehash of the few examples it has for that case.
- pabs3 4y agoI wonder if FOSS folks could copyleft originally public/leaked but proprietary code using CoPilot.
- spupe 4y agoI disagree. Copilot is selling content-aware code suggestions, which is a result of code that other people wrote in their platform, and which in no way affects the work of these people.
- skc 4y agoI get the feeling this entire debate would have been non-existent had this been a Jetbrains product instead. The whole thing is just bizarre when the vast majority of developers constantly look at OSS code daily and lift ideas/patterns/snippets from there regularly without once looking at whatever license is attached.
- Luc 4y ago> the vast majority of developers constantly look at OSS code daily and lift ideas/patterns/snippets from there regularly Perhaps in your circles, but that's certainly not something I've encountered over a 25 year carreer.
- skc 4y agoSo when you google a problem and it leads you to a code snippet that solves it that just happens to be OSS, you immediately scrub your brain and pretend you never saw it and instead instead come up with your own completely independent solution after the fact?
- avereveard 4y agoGoogle usage is outright forbidden for work in institutions that care about intellectual property rights, so the brain scrub issue is just arguing at the wrong level. If you're googling solutions around you're already not taking intellectual property seriously enough to care about what happens after you lift ideas around.
- anonymoushn 4y agoCan you name these institutions? I am surprised to hear that some institutions would prevent devs from viewing e.g. documentation of the APIs they are using or academic papers about algorithms for computing the multiplicative inverses of 64-bit integers, if they accessed those things via google
- 4y ago
- Separo 4y agoGitHub provides the repo hosting and tools for free on public projects. I'm happy with this deal.
- jalfresi 4y agoThis does raise a point - do we now have to assume that all those services that provide free hosting/access/service to open source projects will be strip-mining the work of the open source community to sell them back to us all? I almost feel stupid believing it was an altruistic move to contribute back to the shoulders of giants they were already standing on...
- eloisius 4y agoI feel scammed too. At this point it should be obvious, but I’m finally savvy to the fact that every tech company that offers anything free, and you use it to create “your” content, is not your friend and you don’t even own the works you host with them. I feel scammed that GitHub was cool about 10 years ago. It was like the professional/cultural center of gravity in my career. GitHubbers we’re cool people. Everyone cool hosted their site on GitHub Pages. I didn’t want to see a resume; what’s your GitHub? Now I feel stupid for having contributed whatever tiny bit of brains I did to this AI by thinking that I was using the cool, developer-first code website.
- Separo 4y agoNo. You still have the option not to buy Copilot and still use GitHub's services for free on public projects. Or, if you're not comfortable with your open source code being perused by an AI, you can set up your own privately hosted public Git repo pretty easily. I honestly don't understand the general outrage at this fair seeming deal to me.
- spupe 4y agoIf you assigned a task to a junior dev, and he/she used some code from open source projects and Stack Overflow to develop a custom program for the task, would you say that this person is selling you other people's code? Is it common or expected for this type of use to be acknowledged?
- XCabbage 4y agoPeople I've worked with have different philosophies on this, but personally, if you check in code that is distinctive enough that I can identify the source you copied and pasted it from, and you provided no indication (whether in a comment or a PR description) that you copied it, I will really get quite grumpy at you about it. Way too often I burn half an hour needlessly during review in one of two ways: * trying to figure out how the heck someone figured out some "magic" code that achieves something by invoking a bunch of poorly documented library or framework internals, and trying to reverse engineer WTF all the magic does by diving into the framework's source... only to eventually think to google the whole snippet rather than each individual method call, and discover it's copied from a Stack Overflow answer * trying to figure out why something was written in an unidiomatic or overcomplicated way rather than a more obvious approach, and commenting at length on how I'd simplify it... only to eventually realise it was copied from a Stack Overflow answer Attribution isn't just about making sure the right person gets credit, or about license compliance; reviewers and maintainers frequently need to be able to see where stuff was copied and pasted from in order to do their jobs effectively, even for snippets of just a few lines.
- spupe 4y agoI understand where you are coming from. However, I think you are making the assumption that this person simply copy/pasted some code with no understanding of it, or that this code is then very different from your codebase and needs to be refactored. If using Stack Overflow did not add to your overall development time but subtracted from it, because it was used as an appropriate piece of a much bigger puzzle - a far more realistic scenario for both Copilot and our general use of SO -, then I see no issue with it whatsoever. Certainly no moral or copyright issues as this person on Twitter implies.
- whywhywhywhy 4y agoSame deal for Dall-e if they ever productize it.
- abdulhaq 4y agoThat's like saying a plumber just sells parts that other people made
- WesolyKubeczek 4y agoExcept that a plumber buys them first. For money.
- gtf21 4y agoWhich the plumber has bought and paid for and then installs for you, which makes this pretty fundamentally different.
- antihero 4y agoI mean, if it's autocompleting a fairly simple line, and can do that because it's analysed a lot of lines, I don't really see that as "stealing anything". If you are using it to write whole complex functions thatare the same as other people's, I guess that is copying. But if you do the second thing you are not a great dev, and would have probably ended up copy pasting it anyway. I think the first use case is far more common, and creating boilerplate that is so generic you could never really attribute it anyway.
- alpaca128 4y agoThe first can be automated without ML though. And once you use ML you cannot guarantee it won't copy-paste existing code. This whole thing would be fine if GitHub hadn't just used all public code on their platform, ignoring all involved licenses.
- rob74 4y agoThe problem is, if they had used only code with a license that allows copying without attribution, there wouldn't have been a lot of code left...
- alpaca128 4y agoDifficulty doing something legally doesn't justify breaking the law.
- xupybd 4y agoIt changes the code for use. I'm not sure it can be considered a copy. It much like reading someone else's code and drawing ideas and patterns from that code.
- mojuba 4y agoCan I suggest a hypothesis that if you find Copilot useful it means the problem you are solving is a boring one? I might be wrong of course.
- workingon 4y agoSeems like a narrow vision. Is every line of code you write to solve a problem “not boring”? I solve problems I find interesting, but writing matplotlib code to visualize data never is.
- trention 4y agoThis is true for the current iteration of the model. Probably won't be true at least to an extent in 5 years. Besides, there is nothing wrong with solving boring problems. Not everyone can be Bjarne Stroustrup.
- viraptor 4y agoThe most interesting problem will have extremely boring bits. If you write a cli tool to solve all of world problems by changeling magic, you'll still need to add the parameter handling and do some error management. Which is repetitive and likely well generalised and predictable based on other projects.
- para_parolu 4y agoThe problem may not be boring. Typing boilerplate code is. I work on games as hobby. Sometimes I implement mechanics requiring vector math. Working on mechanics is interesting. Writing down math is not. Copilot helps with later.
- mojuba 4y agoThen another hypothesis: you probably haven't found the right tools for it yet. I find myself writing biolerplate mostly around some obscure system framework calls (iOS/macOS), but that's rather rare. But even OS API's and frameworks do evolve over time into requiring less boilerplate. Just take the evolution of CoreAudio, the modern Swift interface is so much better. So at the end of the day it's about the tools and interfaces: boilerplate is rarely absolutely necessary with the right tools.
- lakomen 4y agoI don't understand what's going on there. I don't use github. Can someone explain what the author means? Edit: in detail
- niek_pas 4y agoGoogle “GitHub copilot”
- lakomen 4y agoNice, being downvoted for asking questions. Nice asshole culture on HN.
- tjpnz 4y agoYou can ask questions but they can't be low effort and need to add something to the discussion.
- martin_a 4y agoJust like with StackOverflow, people are expected to invest some time or amount of work in getting familiar with the topic. Your question seemed to lack this kind of work and was probably therefore downvoted. I don't think that's so much about "asshole culture" but more like time management, as not everything can be explained to everybody in every topic.
- npteljes 4y agoGitHub Copilot is a paid feature, but that's a red herring in this discussion - people are free to monetize free software, neither or the major licenses forbid this. GitGub Copilot is an advanced autocomplete / code generation system, based on a machine learning model. The code used for training the model is taken from projects hosted on GitHub. These projects were published under different licenses. The main questions are: Some of the licenses need something from you if you create a derivative work. Does the Copilot training itself count as creating a derivative work? Sometimes the autocomplete basically quotes the original code. Does the original license then apply to the autocompleted / generated code too? How much of verbatim code quoting does it need for the result to be considered a derivative work?
- deleted 4y ago[deleted]
- parhamn 4y agoPretty soon the world is going to come to realize art/creation is just blending, incrementing and repurposing prior art. No book, painting, codebase, sonnet, design is theft-less. The art is the space reduction, otherwise we’d just bruteforce away.
- pera 4y agoI'm not sure what do you mean by "theft-less" but I believe you might be conflating inspiration with derivative work: Copilot can produce verbatim copies of open-source code, this would make it more similar to how some musicians sample other people's music to create new music.
- wnkrshm 4y agoSo the only thing left is handiwork I guess. Engineering isn't different from art in any way, the constraints are just stricter.
- Agamus 4y agoThis idea has been around for a while - why... "pretty soon"? And I'm sure I couldn't disagree with you more. Or are 'influence' and 'theft' the same now?
- TremendousJudge 4y agoThe idea has been around a while, but the legal system doesn't reflect it. I don't think it will any time soon though.
- coldtea 4y ago>Or are 'influence' and 'theft' the same now? They have been the same for most of history. People could openly copy titles, plots, parts, phrases, etc from prior work. Same for mechanical designs. The only thing preventing them was obscurity (e.g. the inventor trying to make it hidden) not any law or ethical idea that it's bad (there wasn't any). That's how things from math to gears to tunes got better (or changed over time, in the case of art, as better/worse is subjective there). E.g. globally and historically folk music has been basically taking whatever you want from tunes and songs where everybody does the same with no "permission" asked or needed to be given. Like 4 verses but want to add a fifth or change some part? Go ahead. Want to play it exactly like you've heard it? Go ahead again. The idea of "theft" in that regard came in the last 2 or so centuries, and was enforced with artificial legal barriers and new "ethical" concepts that are neither "natural", not present for the vast majority of history (including golden ages of art production).
- tpoacher 4y agoDoes this mean I can steal stuff if I say I trained an AI to do it for me?
- yaseer 4y agoTechnically, programmers search, copy and modify code all the time. One might argue copilot puts into software an algorithm that humans are already doing. Software like that is usually inevitable. Still, it sucks there's no benefit for the contributors. The most ethical thing I can think of is some kinda 'Spotify-like' revenue sharing model, based on how often their code is used by others. Not that they'd ever implement that if they can get away with it!
- kaibee 4y ago> The most ethical thing I can think of is some kinda 'Spotify-like' revenue sharing model, based on how often their code is used by others. Not that they'd ever implement that if they can get away with it! Based on my understanding of how NNs work, I'm not sure its even possible to implement something like that.
- omnicognate 4y ago> One might argue copilot puts into software an algorithm that humans are already doing. That argument only works if you think what Copilot is doing is meaningfully similar to what humans are doing. The debate about how these models relate to human thought might have legal implications. As I understand it (IANAL) copyright doesn't protect ideas and concepts. It protects the content itself. In theory, if I read some copyrighted work, understand some idea in it and then create a new work using that idea, without copying that original work, then that is not a derivative work. (I think this is at least how it's supposed to work - would love to be corrected if that's wrong.) So if I took a copyright work and rot-13ed it before distributing copies, I think that would be clear copyright violation, but if I made my own works using concepts I gleaned from reading it, it wouldn't be. So should Copilot be treated like the rot13 algorithm or like me understanding concepts and generating new works using them? That sounds like a fascinating legal debate to be had.
- teakettle42 4y ago> Technically, programmers search, copy and modify code all the time. When following the license terms, preserving the original copyright, etc, sure. However, honest, ethical people (including programmers) do not plagiarize. Copying and pasting code without attribution is plagiarism. Doing it without following the licensing terms is a copyright violation.
- nathias 4y agoCopilot is a new way for corporations to break copyright while enforcing it for everyone else, this will be the first big use for AI when other corpos follow.
- bmacho 4y agoOn a side note, I do believe that short programs or functions should be copyright free by law. Or we as a community need to create a better bsd, a cc0 for everything. Almost everything is nontrivial, and almost everything is copyrighted, at least with the pressure to name the original author (BSD, GPL, other major permissive licenses). Say you want to use a library, then you check for examples in the documentation, now you have to denote somewhere that the example is from the documentation (best if you put it in the source code, so you don't lure other people to copy what you copied and refer you as the author). It is a major PITA at least for me.
- stagas 4y agoWhat about a law that makes all code available but then requires you to use a portion of your earnings to compensate the people their dependencies you used?
- iLoveOncall 4y agoGithub Copilot is selling code other people wrote as much as the author of this thread is profiting from words other people invented. Absolute nonsense.
- nextaccountic 4y agoThe difference is that words aren't copyrighted and doesn't come with an open source license.
- coldtea 4y ago>Hector Martin: If you use Copilot, you are basically playing Russian Roulette that the random mashup of existing, copyrighted, hegerogenously licensed code that you get out of it qualifies as an original work, mostly by chance. Or that nobody will ever sue you otherwise. Well, that's already the case with Stack Overflow copypasta enterprise code. If anything, use of Copilot would be an improvement...
- moffkalast 4y ago> If anything, use of Copilot would be an improvement What do you mean, Copilot regularly pastes stuff directly from SO. One of those automatic doc generators was able to point me to the exact answer where one of them was from.
- coldtea 4y agoThat it doesn't just "copy and paste" but does more involved "AI" mixing
- moffkalast 4y agoI don't think renaming variables and adjusting spaces holds up in court.
- t0suj4 4y agoThat quote applies to any creative work. Be it code, audio or video.
- coldtea 4y agoHe talks about code, and Copilot works with code, so I'm not sure how it "applies to any". If you mean that if you make a "random mashup of existing, copyrighted, hegerogenously licensed" works of art (audio/video), it also applies that you might be sued for it, then yes. But that's not much of an issue with Copilot if you're using it for enteprise code that's already a mashup of copypaste "existing, copyrighted, hegerogenously licensed" and that you wont release and nobody will see anyway. Whereas audio/video you generally want to release. If you make them for your own consumption, then it's my response that rather applies: since nobody will see it, and you don't release/sell/circulate it, you can go ahead and mix Michael Jackson, Disney and Star Wars material - nothing will happen to you.
- fimdomeio 4y agowhat AI is showing is the fuzzy line between creating and copying. The truth is they are both always present in everything we do, we've just been trying to hide it. So it should be as simple as if you're using other people's content for your own profit you should properly compensate them. Or we could just abolish copyright law and assume that everything humans create emanates from culture so its always collectively built and everything should be open source. Or we just do the same we've been doing. Create even more complex laws trying to define this fuzzy line in a way that companies can keep profiting from it a lot more than individuals.
- FeepingCreature 4y agoAll I can think of is Steve Yegge [1]: "They have no right to do this. Open source does not mean the source is somehow 'open'." My code is on Github so that people can read it, reuse it and learn from it. "The freedom to study how the program works", as the FSF says. If some of the people reading it are machines, why would that matter? [1] http://steve-yegge.blogspot.com/2010/07/wikileaks-to-leak-5000-open-source-java.html http://steve-yegge.blogspot.com/2010/07/wikileaks-to-leak-50...
- happymellon 4y agoBecause a lot of this code would be put into closed source software, which is against the licence and would prevent people from exercising the right to study how a program works.
- FeepingCreature 4y agoBut I don't care if closed source programmers read my GPL code! The freedom to learn is not copyleft. So long as they put independent effort into their work, they're good in my book. Shared knowledge is a vital commons, and I'm honored if I can contribute to it. Maybe this goes back to that debunked paper that claimed that transformers were only remixing input samples?
- happymellon 4y agoThey aren't reading your code. This is a program copy/pasting code without attribution.
- FeepingCreature 4y agoAgain, the paper that said that transformers only copypasted input samples was highly misleading. It seems clear to me that Codex has true understanding. (Yes, I know that people have gotten secrets to appear in the output by prompting it in clever ways. That this happens doesn't prove that Codex doesn't understand what it's doing, it just shows that Codex doesn't understand everything.)
- albertzeyer 4y agoSo, how often does it actually happen? Does it happen more often than for a human? Does anyone actually have numbers on this? Of course, if you provide already a copyrighted prefix, and it has seen that code, the chances are high that it would complete the copyrighted code (because that is what you actually would also expect). So, for real use cases in the wild, where you write some own real novel code, how often would it suggest some copyrighted code? And how often would a human? I have used Copilot the last months and I have never ever seen such a case (I can be pretty sure because all the identifier names are really unique, and the code was very custom). However, I assume that I myself might have produced copyrighted code unknowingly because if you write common patterns (e.g. some tree or graph search, or some sort function, implement LSTM or Transformer, whatever), the chances are not so low.
- aaron695 4y ago
- amelius 4y ago"Good artists copy. Great artists steal." :)
- wolframhempel 4y agoWhen my last company got acquired, part of the due diligence process was a scan of our codebase for snippets from stack overflow. Every snippet found that wasn't posted with a clear license by the author was challenged and we rewrote it. Now, I'm not entirely sure how necessary this was from a legal perspective. But introducing an AI into the mix will bring up a lot of uncertainty when it comes to how much change is required for something to no longer be considered a copy/derivative.
- dmortin 4y agoDid the scan find the process if they changed the variable names, for example? Or is that considered a differing snippet then?
- wolframhempel 4y agoThis is exactly where it gets murky. We had the usual 1-4 line snippets. We went the extra mile to change them, rewriting them from scratch, partially with different implementations. Did we need to do that? Would it have been enough to just change a variable name or some spacing or similar? I don't think there's a clear standard. The music industry has struggled with this for a long time. When is a song derivative, when a copy, when is it "inspired by"...
- anonymoushn 4y agoThat sounds rough. Here's an 8-line snippet, please make sure you don't infringe my copyright: p = mmap( null, size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0, );
- redox99 4y agoIsn't all stack overflow content creative commons? https://stackoverflow.com/help/licensing https://stackoverflow.com/help/licensing
- 4y ago
- danamit 4y agoThe code Copilot suggest from any given project most of the time is not enough to credit such project, when I look up code in some GitHub repo, and copy it fully or part of it, I do not credit that project. I do not see Copilot as useful anyway.
- pornel 4y agoTough pill to swallow. Microsoft's actions don't seem fair, but fighting them with copyright could weaken fair use: https://felixreda.eu/2021/07/github-copilot-is-not-infringing-your-copyright/ https://felixreda.eu/2021/07/github-copilot-is-not-infringin... There's a good argument that demanding copyright protections on scraped datasets and short snippets is a double-edged sword. It could harm search engines, distribution of news, and non-commercial ML research too.
- presentation 4y agoGoogle just sells content other people wrote.
- ThereIsNoWorry 4y ago1. You most likely agreed to that by using GitHub. 2. Copy&Pasting Code by manual search exists. 3. This is just a smart tool so you don't have to figure out yourself what to copy&paste (in the best case) and save a lot of time. Sometimes I truly wonder how people can genuinely be upset about things like this. What is broken are copyright and patent laws in the 21st century.
- IdiocyInAction 4y agoI don't think that something like CoPilot is what most GH users had in mind when they published their code. Also, licenses exist (which CP demonstrably doesn't give a shit about).
- zufallsheld 4y agoAs to your first point, there are many repositories on github that the author of code did not upload there or where not all contributors to the code are on github or agreed to let their work be used in such a case.
- redox99 4y agoThat's really no different than somebody uploading proprietary code they don't own (stolen, leaked, whatever reason etc) on Github. Github has to assume that you are allowed to do so. What are they going to do otherwise, somehow manually verify that each repository is legit? Now you might say, what about GPL code you don't own. You are allowed to redistribute it (upload to github). But because you are not the owner you can't license it to Github under new terms (that allow them to use it for ML training). But the question still is, is there anything in the GPL that forbids it's code being used for ML training? Even if the generated model is proprietary, has no attributions, etc?
- megous 4y agoOk, takedown requests exists. Say Qualcomm finally wises up and asks github to takedown a copy of the millions lines of their super proprietary 4G modem firmware implementation from github. Will github retrain the model after each such takedown? :D If not, then it's kinda stupid to argue the point about the lack of knowledge, since lack or not lack of knowledge clearly doesn't matter. Github will happily continue using confidential code even from trigger happy companies like Qualcomm for copilot.
- oytis 4y agoCopilot sells the service of finding the code that makes sense for what you write. Would be better if it could correctly attribute the source(s) though, I hope they will solve this problem at some point.
- nl 4y agoThis isn't how a language model works. It's SO frustrating that even on HN people still fall for this naive and incorrect analysis. Pasting bits I've said before on this topic: Language models do not work like this. They can copy content but usually that's for something like the GPL language text. Generally they work on a character by character basis predicting what is the most likely character to appear next. This very rarely results in copying text, and almost never rare text. Mechanically it has learnt both syntax of language and how concepts relate. So when it starts generating it makes sentence that are syntactically valid but also make sense in terms of concepts. That's really different to just combining bits of sentences, and it gives rise to abilities you wouldn't expect in something just cutting and pasting bits of sentences. For example, few shot learning is mostly driven by its conceptual understanding and can't be done by something with no way to relate concepts.
- deleted 4y ago[deleted]
- tyingq 4y agoIf this were true, then they would have trained it on all of MS's proprietary source code too.
- nl 4y agoIt is true. And that doesn't follow at all.
- tyingq 4y agoThere's enough examples of it regurgitating longish verbatim code out there, and not just comments or GPL license text. If they are comfortable training it on code that isn't licensed for unrestricted copy/paste, I don't personally understand why they can't train it on their own code that's also not licensed for that. Edit: They even added 'q rsqrt,' to their banned word list to squelch an example of long verbatim code passages. Basically, it's not that I don't understand your explanation. It's that it does emit long passages of unchanged code in practice, for whatever real-world reason.
- lysecret 4y agoDon't we all.
- HumanReadable 4y agoSorry for the unproductive tone of this comment, but there's something about the attitude of this tweet that really grinds my gears. Any time someone invents something new and incredible, there's always a crowd of negative nancies eager to discredit and explain why the invention is nothing new and a detrement to society. I don't understand why someone would willingly share their code on github where it is publicly available just to complain when others make use of that knowledge. 'co-pilot just sells code other people wrote' is such a ridiculous understatement of what co-pilot does. Instead of marvelling at the human ingenuity that went into creating it, they sneer at the audacity of openAI to do something without first asking their permission.
- nerdponx 4y agoBoth things can be true. It's clear that it violates the licenses of many software projects. But I do agree that denigrating it as "just selling other peoples code" is missing the whole point of the product and of what you pay for when you subscribe to it.
- meheleventyone 4y agoThey own their code and it either has a license for use or is implicitly rights retained if not. If Copilot regurgitates their code, from a project that is public but with a non-permissive license they are having their IP rights violated so are totally correct in being unhappy about that. Just because you’ve made something cool doesn’t give you the right to harm others in the process. If MS or OpenAI don’t think this is the case then they should have also included their private repositories.
- Zambyte 4y ago> from a project that is public but with a non-permissive license Permissive or not doesn't matter. Public Domain or not is what matters. Permissive licenses still require you to propagate the copyright notice, which Copilot strips.
- core-utility 4y agoDo we have any evidence that copilot doesn't check/filter by license?
- Ciantic 4y agoI'm bit mixed on this, code Copilot usually autocompletes me is not particularly novel, it's just mundane stuff I would write anyway. Most of these snippets are not copyrightable in my opinion, because it was obvious in the first place. Like CSS nth-child odd / even logic, or one case it filled me ~10 lines JS logic of filtering rows by category stored in dataset, which I would have written anyway. Then there are cases where it amazes me completely, it wrote 10 lines of C++ code for rendering a monochrome glyphs with bits using Freetype library. It though had odd subtle bug, the glyphs came reversed and it worked with only certain font size which it seemed to pick up from different file all together.
- tremon 4y agoI might start considering Copilot if Microsoft were to train it on their own internal codebases (Windows, Office, SQL Server). Until they do, it's clearly a "tool for thee but not for me" type of situation.
- clircle 4y ago"tool for thee but not for me" <- what does this even mean?
- iptq 4y agoI know this isn't really related to the whole copying ethics debate, but I definitely feel like there's some sort of foul play happening here. For all of the unlicensed projects out there, the license that is automatically granted to Github includes: > the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time It's insane how vague this is. Is Copilot a "Service"? Sure, by its definition: > The “Service” refers to the applications, software, products, and services provided by GitHub, including any Beta Previews. And since much of the code was published before Copilot's inception, this means Github can just arbitrarily add more "services" and milk the code for whatever it wants. Automatically service-ify any public repository? Sure, pay us for quotas. It's like a legal loophole to let Github just bypass any license restrictions you put on it.
- AtNightWeCode 4y agoCopiliot will be that bandmate that plays a new riff and leave you wondering about where it was borrowed from.
- captainbland 4y agoIf we're all standing on the shoulders of giants (specifically code that other people wrote) then really what Copilot is selling is a ladder to get onto those shoulders faster. I think that's a legitimate aim, as such. However it should be careful about not including unlicensed code and should have a specific 'GPL' option for a model trained with GPL code included. I suppose it should also generate appropriate copyright notices to satisfy many open licenses. I'd be surprised if copilot could actually link back to the original code like that, though.
- bborud 4y agoWell, this does invite an interesting comparison. If we imagine something like Copilot applied to music I believe the chances of ending up in court would be pretty high. There are a lot of examples of plagiarism lawsuits in popular music and the outcome seems to be entirely random. One could argue that the information density in chord progressions, bass lines and beats is extremely small. And that any recognizable part of a musical idea that has been "borrowed" would necessarily make up a larger percentage of the complete work than would be the case for a typical application with borrowed snippets. That's not a bad argument, but it is unsatisfactory because it means that at some point someone has to make a judgement on how much you can borrow.
- bborud 4y agoMy personal reasons for not using copilot are a bit simpler. I believe the act of researching which solutions to use for a given problem is not so much about time, or the code you end up with, but about developing a better understanding of what you are doing. You may end up just cutting, pasting and modifying a piece of code you found, but hopefully, you were exposed to a few different ways to accomplish the same thing, and it made you aware of other choices that could have been made. You could think of the evolution of practical problem solving in software engineering like this: 1. I have to invent a solution (because nobody else in the world has a computer) 2. I have to know of a solution (education, word of mouth...) 3. I have to look up a solution in the books I have (commoditized knowledge) 4. I can look up solutions on the internet <-- (we are here) 5. The computer suggests something and I accept (some are here too) From 1 to 4 the amount of cleverness required to solve small problems drops a bit, but your productivity and exposure to knowledge probably goes up. I'm not quite sure what happens from 4 to 5. Personally I'm actually more interested in the context solutions are presented in than just the solution. In fact, I rarely copy and paste code from the Internet, but I often look at multiple suggestions/solutions and then borrow ideas or combine ideas from several sources.
- Yenrabbit 4y agoAt least the way I use it, it's not taking much away from my problem solving. It's just that instead of having to type `particlesGeometry.setAttribute('position', new THREE.BufferAttribute(positions, 3))` I just write `//Add as an attribute` and then hit TAB, since Copilot is smart enough to see that I've just prepared some geometry and populated an array of positions (both operations also sped up by not having to type the obvious bits). You're still having to think through the solutions (I'm not just typing '//make a cool particle sim') but no longer need to hit SO every few minutes for syntax examples when using a new library or something.
- ModernMech 4y agoThat sounds like a problem that could be better solved through language and library design rather than an AI that sucks up all the code in the world.
- janandonly 4y agoIsn't every programmer in history (except the gall who invents her own language and writes all her own code) simply an archeologist for other people's work? We all Duck/Google for code anyway. Why not admit and make it easier?
- eline43 4y agoYou don't understand the difference between many open source licenses or the concept of crediting open source code authors... it does not mean that the code is free for everyone to just use as they please... https://www.gnu.org/licenses/license-list.en.html https://www.gnu.org/licenses/license-list.en.html for a quick intro Also, are you okay with other people selling *your* work and *you* getting nothing out of it? Many people are not.
- pacifika 4y agoCopilot is doing this on an industrial scale. It’s the difference between copying sample code and outsourcing your work to a third party colkectively
- tiku 4y agoI'm using it for a day now and i'm really impressed. It is so aware of stuff in old code, that it is scary. I'm working in an old application with Zend Framework.
- Proven 4y ago
- eline43 4y agoThere needs to be an update to either licenses or GitHub (and other) software directly, or even software terms of services, that gives the user an opportunity to opt-out of their data being used to train proprietary AI models. 'I don't agree with having an AI trained on/with my data.' IMHO, all other problems with copilot stem from this.
- shireboy 4y agoI do feel these arguments are valid if a little overstated. Most devs have googled, found some code, and pasted it in without thinking about attribution. Doesn’t make it right, but it is a question of how much code is being copied and how specific. For example, I peruse open repos to learn - I learned about the spread operator in JavaScript that way- doesn’t mean every time I use it I need to attribute whatever repo I saw it in. But, yeah, if I copied a larger chunk and the owner wants attribution, probably wrong. I like the idea of having the bot automatically update a attribution file if it detects it’s used licensed code. Seems like it would be fairly trivial. Also a robots.txt for repo owners to control automated use. Also, they should totally pay back a portion of revenue to the community and support the repos used to train. That seems like it would be a good PR move if nothing else.
- kachhalimbu 4y agoI like this take. Copilot to me seems a glorified (very intelligent) auto-search-paste/autocomplete service. It is just mimicing what usual devs do which is to copy-paste code from StackOverflow/github for many mundane types of codes like for loops, mongo find queries, callback func definitions etc for JS devs for eg. The idea of auto-attribution if copilot surfaces licensed code is best because then it keeps the copilot user honest where the code is coming from and honor the original license.
- teakettle42 4y ago> It is just mimicing what usual devs do which is to copy-paste code from StackOverflow/github for many mundane types of codes like for loops, mongo find queries, callback func definitions etc for JS devs for eg. I’m genuinely disturbed to see how many people in this thread think that casual plagiarism is the norm for “usual devs”.
- ParetoOptimal 4y ago> I’m genuinely disturbed to see how many people in this thread think that casual plagiarism is the norm for “usual devs”. I'm disturbed it is likely the reality.
- zokier 4y agoSure, the concern is valid but I feel like this tweet adds absolutely no substance to the discussion and just repeats the same opinion that was already rehashed to death since copilot originally launched. As such, especially with the tone that the tweet has, I don't expect constructive discussion to raise here.
- Aeolun 4y ago> what github / microsoft is counting on here is that open source developers do not have enough collective power to do anything to stop this I think it much more likely that they count on everyone liking it way too much to give a shit about their MIT code not being attributed correctly. I certainly don’t. MIT just seems like the most convenient license for people that need licenses (corporations?), so that is what I use.
- marstall 4y agomost of the code I write is glue sticking together 8 proprietary systems nobody's ever heard of. how is copilot gonna help me with that?
- SMAAART 4y agoOnce again Innovation challenges IP.
- floor_ 4y agoI started self hosting when Microsoft bought github and with this mass theft of copyrighted material and then reselling it for money I'm even more happy with my decision.
- boomer_joe 4y agoWe need a licence that forbids use in ML and the people willing to sue github for it ASAP.
- ilikehurdles 4y agoBut using it in a GitHub project would be akin to those Facebook comments that demand the company not monetize them.
- sytelus 4y agoGoogle just sells content other people wrote.
- vbezhenar 4y agoI somewhat agree with that. Yesterday I edited some exotic configuration (Kubernetes CSI driver for Cinder) and Copilot suggested me config which looked like someone's config. There were no values, so it was good at filtering them out, but it definitely looked like cleaned part of code which resides in some project. I don't think that's bad though. Code sharing is good for overall productivity.
- blitz_skull 4y agoMan, people really do be angry that the public code they put on a public platform is being used publicly. Wild.
- namose 4y agoWild that people who draw up licenses for their code which up until this point have been reliably enforced expect them to continue to be enforced!
- dgb23 4y agoReading many of the comments here I feel like one important thing is being left out that is not related to legal, but to social issues: Who is on the side of open source? Where are the big, powerful institutions and companies that deeply care about authors and communities providing free software that so many of us rely on?
- janosdebugs 4y agoIt'd be nice to see some proof here. Copyright is not absolute and does not extend, for example, to things that have no creativity in them. There are only so many ways to write a for loop or an if condition. Training an ML model from a large body of code IMHO violates copyright no more than any of us reading code and learning from it, as long as GH Copilot doesn't spit out code that's exactly the same as something already existing.
- namose 4y agohttps://twitter.com/mitsuhiko/status/1410886329924194309?s=21&t=JQZgvz31RaND783DjQmlig https://twitter.com/mitsuhiko/status/1410886329924194309?s=2...
- VoodooJuJu 4y agoIt is now proven that copilot returns code from codebases with non-permissive licenses [1]. I'm curious - what are the legal implications of this going forward? I've so many questions. 1. Will Microsoft ever face lawsuits for these license violations? 2. If so, who/how? Class-action? 3. Will copilot be forced to open-source in the future? Under which license? Some open source licenses are incompatible with others, but copilot uses code from probably every OSS license conceived. 4. If Microsoft faces no justice, will we start seeing more OSS license violations? Will Google start using AGPL-licensed code? [1] https://news.ycombinator.com/item?id=27710287 https://news.ycombinator.com/item?id=27710287 | Copilot regurgitating Quake code
- pwdisswordfish9 4y agoIs there any leaked Microsoft code on GitHub? Someone should check if Copilot regurgitates that as well, then see how Microsoft reacts when someone slaps an AGPL license on that…
- q-big 4y ago> Is there any leaked Microsoft code on GitHub? There seems to be. Google 'windows nt source code leak github': https://www.google.com/search?q=windows+nt+source+code+leak+github https://www.google.com/search?q=windows+nt+source+code+leak+... First search results: Windows NT 4.0: > https://github.com/lianthony/NT4.0 https://github.com/lianthony/NT4.0 > https://github.com/ZoloZiak/WinNT4 https://github.com/ZoloZiak/WinNT4 Windows XP: > https://github.com/tongzx/nt5src https://github.com/tongzx/nt5src > https://github.com/onein528/NT5.1 https://github.com/onein528/NT5.1
- 542458 4y agoIANAL. My understanding is that the general legal precedent in the US is that a) datamining text has no copyright implications (in the same way that reading a book has no copyright implications) and b) it is not a copyright violation to use a small amount of copyrighted material provided the context is sufficiently transformative. This might seem silly or unfair to you, but that is the current legal reality. But even ignoring that, everybody uploading code to GitHub has given GitHub the right to analyze that code as per the GitHub ToS. This is the same mechanism by which you can't upload code to GitHub with a license that says "nobody is allowed to display this code on the internet" and then sue GitHub.
- Havoc 4y agoYes, though in a way so does stackoverflow & friends. Large chunk of dev ecosystem is copy paste and I don't think this is inherently problematic. It is always a case of standing on the shoulders of giants. Its more of a licensing issue to me. As far as I can tell it was train on a blend of licenses which to me makes it inherently non-compliant. At least some of it is going to be copyleft and find its way into closed source.
- c01n 4y agoMS and Github are thieves, all their code is closed source, yet they sell copyrighted code they don't own. If they told us years ago that our code will be automatically stolen by an "AI", most coders would not have created an account. The innovation here is that they have access to most of the worlds open source code and automated the stealing.
- LeonTheremin 4y agoAnd social media sells ideas other people thought. Copilot is limited to public code now, but it may easily be trained on non-public code - albeit this probably won't be for sale to the public.
- williamcotton 4y agoShould the snippets that Copilot is regurgitating be considered for copyright in the first place? It seems akin to trying to copyright a certain drum pattern or chord progression. Also, the history of the GPL, MIT, commercializing lisp machines, Symbolic, infighting, etc… seems a very different context than Copilot so I am having difficulty seeing the systemic problems that tools like this encourage. There is of course a surface level similarity in that a corporation is profiting from IP in the public domain but the devil is in the details.
- tsujp 4y agoCopilot produces verbatim GPL'd code. It's also a closed box. Source: https://twitter.com/mitsuhiko/status/1410886329924194309 https://twitter.com/mitsuhiko/status/1410886329924194309
- ewalk153 4y agoIf the portion of code that Copilot lifts is the "heart" of the original work, that would be much less likely to be considered fair use[1], regardless of the length. > For example, it would probably not be a fair use to copy the opening guitar riff and the words “I can’t get no satisfaction” from the song “Satisfaction.” I wonder how this could be integrated into the system? [1] https://fairuse.stanford.edu/overview/fair-use/four-factors/#the_amount_and_substantiality_of_the_portion_taken https://fairuse.stanford.edu/overview/fair-use/four-factors/...
- rictic 4y agoCopilot very rarely copies code verbatum, and when it does it's very short snippets. When Oracle sued Google over allegedly copying short and fairly trivial snippets of code they were justly derided. I can't speak to the legal side, but I just don't understand the moral outrage over very occasionally copying such short snippets of code. The key innovations and the actual value that licenses are intended to protect aren't in these short snippets. And what does copilot bring to the community? Free use by students, free use by open source maintainers, and a huge boost in productivity for a modest fee for professional devs, for a service that no doubt costs a lot to run, even on the margin.
- aetherspawn 4y agoCopilot is a fancy pattern bot. Humans make original patterns, but since Copilot cannot think, then Copilot does not. It squashes together a bunch of small individual patterns, each under their own license, but at no stage does it do anything more than pick a line from here, and a line from there. It doesn’t think, and it doesn’t create new IP. It is like making a picture out of small snippets of a thousand other pictures, and then selling it.. clearly not OK. You still ripped off the original artists. Or like plagiarising 100 of your class mates’ assignments. Are you less guilty because you went to the effort to steal just a few sentences from each? A criminal who steals a cent from every account at the bank is a more sophisticated thief than someone who holds up a petrol servo. If Copilot doesn’t create new IP (it doesn’t; we established this), then it uses existing IP. And in that case it is no different to any of the three analogies above.
- GuB-42 4y ago> Copilot just sells code other people wrote So what? Selling code other people wrote is the foundation of the free software movement. It is the entire business model of countless companies, and it is a good thing. Among them are most major linux distro vendors like Red Hat and Canonical. The value added by Copilot is that they sell you the lines "code other people wrote" you want out of billions. I still think it is derivative work, and that they should only process code under permissive licenses, or, if they want to include GPL code, make a GPL-only version, usable only for GPL projects. I thought it is what they did, there is so much code under permissive licenses that is should be enough to train their model, but apparently, they don't care, as long as it is public, it is included. For me, they are shooting themselves in the foot, several companies have already banned Copilot due to the potential issues with copyright.
- tiborsaas 4y agoMrDoob has an excellent point about this: https://twitter.com/mrdoob/status/1539740854956412929 https://twitter.com/mrdoob/status/1539740854956412929
- 0x_rs 4y agoI'm not a lawyer, nor very well versed in the vast world of licenses and their definitions in court contexts, but I've been wondering about something with the growing appeal ML-generated content has for the average person (and the "high" barrier for entry in the market) — are licenses in some form or another going to adapt to this phenomenon? From a brief search, I have not found any new license with a no-dataset-usage clause (assuming fair use does not apply, that's another big question). What are the chances anything of the sort will become an option for any "creative" work that's usually shared freely (such as artwork, code, et cetera) even despite copyright? What about the ownership of the dataset? It seemed to be questionable years ago already that possibly IP-protected content goes through the black box and resembling material gets on the other side, whose ownership is it really? I'm guessing some notable court cases in the future could define this in the following years if the popularity continues growing.
- olalonde 4y agoI'm going to make a bold prediction: no one will ever lose a copyright lawsuit due to usage of Github Copilot generated code. The code snippets it produces are too small or trivial to qualify for copyright infringement.
- ModernMech 4y agoCoPilot is a new technology, and smallish snippets of code are all it is capable of at this point. Microsoft will surely work to expand its capabilities to produce larger and more complex programs, don’t you think?
- thih9 4y agoIs github copilot using private repositories for the learning process? If yes, how do they mitigate the risk of exposing private data when something is quoted verbatim? If not, then why are repos with non permissive licenses ok?
- stakkur 4y agoAt every turn, in every instance, for decades, all stories involving Microsoft end in "...and then Microsoft fucked people over." I've witnessed this firsthand since the 80s.
- honkler 4y agolicense issues will save many thousand jobs.
- BiteCode_dev 4y agoIt is incredible to use though. I pasted the return value of an API call in comment, then started to write a schema class. Codepilot just created the entire class for me. wanted to extract a subset of the data, I typed get_<_name_of_the_subset>(), it wrote the code I would have written. So even without using someone else code, just the pattern understanding and the production of simple boiler plate code is great.
- noisy_boy 4y agoSay, I want to write a getter method like below: String getName() { return name; } Let us also assume that this snippet, unsurprisingly, has been in several copyrighted repos that didn't grant Github the right to share this code. So I start tying "getName" and copilot suggests the exact snippet above. If I use this snippet, is it plagiarism? Even though the above code is the most "obvious" way to write this getter and I would have written it this way even without copilot's suggestion? Or does the "uniqueness" or "non-trivial quantity" of the suggestions have any bearing in determining copyright violation? How/where do we draw the line?
- glouwbug 4y agoLucky for you if you, if you wrote a noise function that copilot returned as an implementation of Perlin noise you'd be breaching a _patent_! Said patent just expired a 20 year run, so you'll be okay this time!
- warkdarrior 4y agoClearly your code could be improved with some `Factory` objects and some dependency injection!
- nickjj 4y agoThis might be overreacting but is there a way to opt-out of Copilot using your code in open source repos? It feels morally wrong to me that I can spend thousands of hours working on projects on my own free will but then a company can sell the code I wrote to others in the form of snippet completion as a service. In fact they end up selling your code back to yourself if you plan to use the service. If the answer is no, that moves the needle pretty far in the direction where I'd at least consider the idea of moving all of my repos to Gitlab. I don't care much about stars or popularity. I open source things that are interesting and useful to me and if other folks want to use it they can but I don't gain motivation from others using the projects I release. I like Github and its UI and it's no doubt "the spot" for open source but selling code written by others rubs me the wrong way a lot. It stinks because it also means no longer contributing to other code bases too. It's moving us in the opposite direction of what open source is about.
- jaywalk 4y agoIf your code is using a license that allows it, how could you possibly opt-out aside from using a different license?
- okasaki 4y agoMicrosoft could provide an opt-out for projects or even contributors, regardless of licence.
- nickjj 4y ago> If your code is using a license that allows it, how could you possibly opt-out aside from using a different license? A repo setting that instructs Github not to use your code for Copilot, it could be a similar option as turning Discussions on / off. If they really want to win developers over they would even have Copilot scanning disabled by default but that'll never happen.
- jonny_eh 4y agoSounds like you want a new license that just prohibits use by one company for one purpose.
- lfrigodesouza 4y agoIt's as the saying go, "when a product is free to use, the real product is actually you". In this case, our code is the product. Just considering now on swapping to another git provider...
- Guid_NewGuid 4y agoI find this whole topic very annoying, this is like the 3rd variation to reach the front page today. But it has made me realize why I instinctively dislike Free Software as a movement. Copyright and licensing are bad, actually. Stop getting worked up about the idea of using courts to punish theft. Stop getting into a frenzy of arousal about the police kicking down doors to drag Billy Gates to jail because 80 characters of fast square root is theft but 79 isn't. Where on earth is the ambition and vision!? Knowledge is public domain. A commons of knowledge is a public good. The cost of code copying is zero. Sure in our day job we have to pretend to care about this stuff. But when did the ideological scope of what can be achieved become rules lawyering over license text. Copy my MIT licensed code without attribution? I don't give a shit, go ahead, I hope it helps, in fact I want a truly public domain license but copyright law is so hostage to corporate interests no such thing exists in many countries. Free the code.
- progman32 4y agoI see the free software movement as a variant on your ideals but rooted in practicality given the current environment.
- Guid_NewGuid 4y agoI think we share a lot of the same goals but they presuppose openness based on violence, if you don't do what their license says exactly then they're going to use lawyers and courts and the state's monopoly on violence to make you comply. I think at a fundamental level this abandons any vision of a true commons since as copilot discussions reveal the well is now polluted (to mix metaphors) and though in some frames the code is more free you certainly won't be if you fail to pay the penalty levied in a civil case for misusing it.
- imtringued 4y agoThat is true of any license.
- sirsinsalot 4y ago
- rosmax_1337 4y agoI think this problem has no good solution until IP laws around the world are properly reimagined from the ground up. I'm of the quite radical stance that code, music, art in terms of their intellectual existence should be free for anyone to take. (you can own a harddrive with code on it, and claim noone should steal it, but not the idea of the code itself) If you have ideas, code, music or art which you wish for noone to partake in, do your best to keep them secret. Certainly, breaking into secret areas should be illegal, but once the cat gets out of that bag it gets out of the bag. The creative people behind these ideas I believe will be able to find good compensation nonetheless in society, IP-laws nowadays only serve to protect megacorporations to the detriment of creativity and ideas.
- zzo38computer 4y agoI agree. This will fix it. I think that copyright and patent should be abolished, but that if it is secret then it is still secret (unless someone else manages to come up with the same thing (e.g. by decompiling a published computer program to reconstruct the source code), which case it can be public). And so then also the AI can copy the code too just as much as you may do so manually; if it is published then you can do it and it should not be illegal to write such things.
- sirsinsalot 4y agoJaron Lanier's book "Who Owns the Future?" Is all about AI and compensating those that input in training these very valuable models. I highly recommend everyone read it.
- shahar2k 4y agoand Dalle2 sells art other people created (I'm actually not being sarcastic, I think there needs to be some sort of pipeline for compensating the artists who are used to train these models
- borishn 4y agoCopilot is fair use, get over it! Copilot is not writing your code any more that Google search is writing your code. You are writing your code, and Copilot is just making suggestions. US constitution secures limited copyright to "To promote the progress of science and useful arts". Copilot is just that, get over it!
- Buttons840 4y agoA good and well argued opinion made hostile by saying "get over it" twice! Saying "get over it" discourages further discussion. Your comment would be better without it.
- nescioquid 4y agoNot an expert, but fair use generally covers education, criticism, parody, and satire. There is a test for meeting fair use and it includes things like amount copied and commercial or non-profit interest. The amount copied from any particular source might be small, but an aggregate strip-mining of many copyrighted sources is an interesting twist. Another might be, as you suggest, it might be a machine that itself does not violate copyright, but has the effect of causing users (who accept the suggestions) to violate copyright.
- collegeburner 4y agoGoogle does the same thing taking snippets out of pages or even completely caching them so you can see the entire page from their servers.
- brianmcc 4y ago
- HeavyStorm 4y ago
- andrelaszlo 4y agoThere are a few reasons why this could be considered ethical. First, open-source code is typically free to use, so the company would not be taking advantage of anyone by using it to train their AI. Second, the company would be providing a service that people are willing to pay for, so they would be generating value for society. Third, the company would be transparent about what they are doing and would not be hiding anything from the public. ...the above was generated by GPT-3 (text-davinci-002). Prompt: Write an argument for why using open-source code to train an AI and then sell the code generating service (without open-sourcing it) is ethical. The main argument against this is that it takes away from the open-source community that contributed to the development of the code in the first place. By selling a code-generating service without open-sourcing it, the company is profiting from the work of others without contributing back. This is unfair and takes away from the overall open-source ecosystem. Added two characters to the prompt :P
- madrox 4y agoI don't think any professional community is aligned on how to think about ML-generated content yet. We don't know how to apportion rights between the data owner, the model owner, and the end user, and I don't think existing copyright law is ready for it. At least for software, I think the way forward is for the next generation of software licenses to explicitly state whether the code can be used to train ML models and what those models can be used for. Without explicit language, we'll be squabbling over interpretations of fair use. There's going to be some big cases here. It's going to end up in the Supreme Court sooner or later, and if it were to go there today I think I know what they'd say.
- sirsinsalot 4y agoBeware geeks with gifts. This is Microsoft. The question isn't "is it good?" but "Why are Microsoft offering it and how is it undermining everyone else?"
- pen2l 4y agoBit of a stretch to fashion AI-derived/AI-coauthored works as other people's work. Are DALL-E portraits done Picasso-style unrightfully selling Picasso's works? Is an individual selling portraits done Picasso-style unrightfully selling Picasso's works? No, of course not. Joyce's literature was influenced by Ibsen, Mozart looked up to Haydn, Newton was humble enough that he openly professed he stood on the shoulders of his predecessors, Perelman refused the Millennium prize because it wasn't also offered to his colleague Hamilton. All human innovation is iterative, and derivative. https://www.youtube.com/watch?v=jcvd5JZkUXY https://www.youtube.com/watch?v=jcvd5JZkUXY Our skill doesn't grow in vacuums, without outside mentorship and guidance. There are areas where I am upset about the application of AI, but this is not one of them. Consider copilot a gentle guiding hand for those without access to a second pair of eyes nearby to give you reminders on what you may otherwise have on the tip of your tongue. But in the way that Led Zeppelin refused to recognize how heavily their music was influenced by delta blues artist was unbecoming, I can accept the argument that it is perhaps douchey of Github to sit on Copilot as squarely their creation.
- pvaldes 4y agoEach day sounding more as Zopilote, it seems.
- acuozzo 4y agoThis is, in part, why I will continue to use the original 4-clause BSD license for the code I write.
- thewoolleyman 4y agoArtificial Intelligence is causing us to revisit the difference between free as in beer and free as in speech (https://en.wikipedia.org/wiki/Gratis_versus_libre https://en.wikipedia.org/wiki/Gratis_versus_libre). It is putting a new spin on some traditional Open Source Lessons (https://en.wikipedia.org/wiki/The_Cathedral_and_the_Bazaar#Lessons_for_creating_good_open_source_software https://en.wikipedia.org/wiki/The_Cathedral_and_the_Bazaar#L...). People share and reuse snippets of unattributed snippets of MIT-licensed and GPL-licensed code on the internet all the time, StackOverflow, etc. StackOverflow is profiting from that activity indirectly by facilitating it. They profit passively through ad revenue, and actively through the Teams subscription offering. But nobody seem too upset about that. How is an AI which facilitates the same code sharing fundamentally any different? Because it’s scraping it itself, rather than humans contributing it? Seems like a tenuous argument at best.
- mawadev 4y agoWhat stops me from re-uploading copyrighted source, where I remove the notices and push it with an MIT license? If such a data set has been trained with, how do you get it out?
- HeavyStorm 4y agoSo much bullshit my head hurts.
- mullikine 4y agoTraditional 'real' (as opposed to 'imaginary') programming is like writing in assembly code; It's outmoded because of generative models, in a way similar to 'C' outmoding assembly code. The most important thing, I think, is that free (libre) software developers are able to work with the language models directly, so that libre software is allowed to continue progressing into what I call Imaginary Programming. That's because with a generative internet all you really need is blockchain + prompting. https://huggingface.co/spaces/mullikine/ilambda https://huggingface.co/spaces/mullikine/ilambda Language models are able to 'steal' the linguistic meaning-making 'essence' of the software, by modelling: - How the software is used (mimicing its function) - external meaning - How functions are 'inspired' - internal meaning (reflection) https://github.com/semiosis/imaginary-programming-thesis https://github.com/semiosis/imaginary-programming-thesis The models themselves should be clear about where the data came from. However, this is only possible in a fair world which we do not live in. Compromise must be made to protect national interests. Generative models are license blind and there's very little that could be done to prevent progress. Like what the invention of the camera has done for art. Large language models including Codex are a transformative technology. Bi-directional fair-use is probably the best result we can hope for. So long as Microsoft and OpenAI are not selling back usage of the model to the open-source community, I think it's OK, though it's the bare minimum obligation.
- powerapple 4y agoWhy is it a bad thing? You either have people spending time reading code and learn every little thing and produce the same work in days, or have Copilot saves human life time for hours. Coding would be more efficient, it is a win-win for everyone in this industry, right? I know people attach to the code they write, but we all learn from books, and the result is common enough.
- thfuran 4y agoWhat's good about producing a bunch of code that no one understands and only probably does the right thing?
- stefanos82 4y agoSeems like my original questions [1] are more relevant than ever! [1] https://news.ycombinator.com/item?id=27677598 https://news.ycombinator.com/item?id=27677598