26 ms·
Hi HN, we've been building GitHub Copilot together with the incredibly talented team at OpenAI for the last year, and we're so excited to be able to show it off
by natfriedman 5y ago
Hi HN, we've been building GitHub Copilot together with the incredibly talented team at OpenAI for the last year, and we're so excited to be able to show it off today.
Hundreds of developers are using it every day internally, and the most common reaction has been the head exploding emoji. If the technical preview goes well, we'll plan to scale this up as a paid product at some point in the future.
- fay59 5y agoHave there yet been reports of the AI writing code that has security bugs? Is that something folks are on the lookout for?
- natfriedman 5y agoI haven't seen any reports of this, but it's certainly something we want to guard against: https://copilot.github.com/#faq-can-github-copilot-introduce-insecure-code-in-its-suggestions https://copilot.github.com/#faq-can-github-copilot-introduce...
- peterburkimsher 5y agoHas there been an attempt to train a similar ML model on a smaller dataset of standards-compliant code? e.g. MISRA C. I started working at a healthcare company earlier this year, and my whole approach to software has needed to change. It's not about implementing features any more - every change to our embedded code requires significant unit testing, code review, and V&V. Having a standards-compliant Copilot would be wonderful. If it could catch some of my mistakes before I embarrass myself to code-reviewing colleagues, the codebase would be better off for it and I'd be less discouraged to hear those corrections from a machine than a person.
- enriquto 5y ago> Hundreds of developers are using it every day internally, and the most common reaction has been the head exploding emoji Apart from developing this "head exploding" stuff, couldn't some of these incredibly talented hundreds of developers fix the github code search?
- Asmod4n 5y agoMight this end up putting GPL code into projects with an incompatible license?
- natfriedman 5y agoIt shouldn't do that, and we are taking steps to avoid reciting training data in the output: https://copilot.github.com/#faq-does-github-copilot-recite-code-from-the-training-set https://copilot.github.com/#faq-does-github-copilot-recite-c... https://docs.github.com/en/early-access/github/copilot/research-recitation https://docs.github.com/en/early-access/github/copilot/resea... In terms of the permissibility of training on public code, the jurisprudence here – broadly relied upon by the machine learning community – is that training ML models is fair use. We are certain this will be an area of discussion in the US and around the world and we're eager to participate.
- Hamuko 5y ago>training ML models is fair use How does that apply to countries where Fair Use is not a thing? As in, if you train a model on a fair use basis in the US and I start using the model somewhere else?
- KMnO4 5y agoI don’t think it’s fair to ask a US company to comment on legalities outside of the US.
- detaro 5y agoIt's fair to expect a international company pushing its products all over the world to be prepared to comment on non-US jurisdictions. (I have some sympathy for "we have a local market, and that's what we are solely targeting and preparing for" in companies where that is actually the case, but that's really not what we are dealing with in the case of Microsoft/GitHub)
- 5y ago
- yewenjie 5y agoThis looks really cool. Do you plan to release this in some other form like a language server so that it can be easily integrated to other editors?
- jonas_kgomo 5y agoThis is obviously controversial, since we are thinking about how this could displace a large portion of developers. How do you see Copilot being more augmentative than disruptive to the developer ecosystem? Also, how you see it different from regular code completion tools like tabnine.
- tux1968 5y agoHow many jobs have developers helped displace in business and industry? I don't think it's controversial that we become fair game for that same automation process we've been leading.
- AnIdiotOnTheNet 5y agoIndeed. It should be the goal of society to automate away as much work as possible. If there are perverse incentives working against this then we should correct them.
- gugagore 5y ago1. How do you define work differently from "that which should be automated"? 2. While I agree with your stance, it is not by itself sufficient. If you provide the automation but you do not correct the perverse incentives (or you worry about correcting them only later) that you mention, then you are contributing to widening the disparity between a category of workers (who have now lost their leverage) and those with assets and capital (who have a reduced need for workers).
- AnIdiotOnTheNet 5y agoI agree, the fact we're even talking about this is evidence that our society has the perverse incentive baked in and we should be aware of and seek to address that. Regardless, programmers would be hypocritical to decry having their jobs automated away.
- mkr-hn 5y ago
- pera 5y agoCool project! Have you seen any interesting correlations between languages, paradigms and the accuracy of your suggestions?
- deleted 5y ago[deleted]
- kif 5y agoI visited https://copilot.github.com/ https://copilot.github.com/, and I don't know how to feel. Obviously it's a nice achievement, not gonna lie. But I have a feeling it will end up causing more work. e.g. the `averageRuntimeInSeconds` example, I had to spend a bit of time to see if it was actually correct. It has to be, since it's on the front page, but then I realized I'd need to spend time reviewing the AI's code. It's cool as a toy, but I'd like to see where it is one year from now when the wow factor has cooled down a bit.
- ec109685 5y agoIt has the ability to generate unit tests as well, which will help cut down some on the verification side if you feed it enough cases.
- CloselyChunky 5y agoWell then you have to check the generated tests. That's just one more layer, isn't it?
- freedomben 5y agoI think I'd love to use this to generate tests and then write the functions myself. Test generation seems like a killer feature.
- taftster 5y agoYes!! Totally agree. Imagine writing a method and then telling an AI to write your unit tests for it. The AI would likely be able to come up with the edge cases and such that you would not normally take the time to write. While I think the AI generating your mainline code is interesting, I must certainly agree that generating test code would be the killer feature. I would like to see this showcased a little more on the copilot page.
- kimburgess 5y agoYou don’t need AI for that. While example based testing is familiar to most, other approaches exist that can achieve this with less complexity. See: property based testing.
- verst 5y agoI have been using this - for example working in Go on Dapr (dapr.io) or in Python on one of its SDKs. I love it. So often the code suggestions accurately anticipate what I planned to do next. It's especially fun to write a comment or doc string and then see Copilot create a block of code perfectly matching your comment.
- ArtWomb 5y agoHi Nat! Just signed up for the preview (even though I'm the type to turn off intellisense and smart indent). I was wondering if WebGL shader code (glsl) was included in the training set? Translating human understandable graphics effects from natural language is a real challenge ;)
- foobarbazetc 5y agoAre those developers worried about having their jobs replaced by a code-writing AI? :) I mean... why would 95% of developer jobs exist with this tech available? You just need that 5% of devs who actually write novel code for this thing to learn from.
- mkr-hn 5y agohttps://en.wikipedia.org/wiki/Profession_(novella) https://en.wikipedia.org/wiki/Profession_(novella) An Isaac Asimov story about someone who didn't take to the program and, as a result, got picked to create new things because someone has to make them.
- mooreds 5y agoI love this story. If you want to read the whole thing, it's here: https://www.abelard.org/asimov.php https://www.abelard.org/asimov.php
- mdellavo 5y agoWhat do you think about this being overall detrimental to code quality as it allows people to just blindly accept completions without really understanding the generated code. Similar to copy-and-paste coding. The first example parse_expenses.py uses a float for currency - that seems to be a pretty big error that's being overlooked along with other minor issues around no error handling. I would say the quality of the generated code in parse_expenses.py is not very high, certainly not for the banner example. EDIT - I just noticed Github reordered the examples on copilot.github.com in order to bury the issues with parse_expenses.py for now. I guess I got my answer.
- as300 5y agoWhy would you say it's an error to use a float for currency? I would imagine it's better to use a float for calculations then round when you need to report a value rather than accumulate a bunch of rounding errors while doing computations.
- joquarky 5y agoYou don't want to kick the can down to the floating point standard. Design for deterministic behavior. Find the edge cases, go over it with others and explicitly address the edge case issues so that they always behave as expected.
- mdellavo 5y agohttps://stackoverflow.com/questions/3730019/why-not-use-double-or-float-to-represent-currency https://stackoverflow.com/questions/3730019/why-not-use-doub...
- RobLach 5y agoStandard practice is to use a signed decimal number with an appropriate precision that you scale around.
- Tainnor 5y agoIt is widely accepted that using floats for money[1] is wrong because floating point numbers cannot guarantee precision. The fact that you ask is a very good case in point though: Many programmers are not aware of this issue and would maybe not question the "wisdom" of the AI code generator. In that sense, it could have a similar effect to blindly copy-pasted answers from SO, just with even less friction. [1] Exceptions may apply to e.g. finance mathematics where you need to work with statistics and you're not going to expect exact results anyway.
- stephen82 5y agoLots of questions: - the generated code by AI belongs to me or GitHub? - under what license the generated code falls under? - if generated code becomes the reason for infringment, who gets the blame or legal action? - how can anyone prove the code was actually generated by Copilot and not the project owner? - if a project member does not agree with the usage of Copilot, what should we do as a team? - can Copilot copy code from other projects and use that excerpt code? - if yes, *WHY* ?! - who is going to deal with legalese for something he or she was not responsible in the first place? - what about conflicts of interest? - can GitHub guarantee that Copilot won't use proprietary code excerpts in FOSS-ed projects that could lead to new "Google vs Oracle" API cases?
- natfriedman 5y agoYou should read the FAQ at the bottom of the page; I think it answers all of your questions: https://copilot.github.com/#faqs https://copilot.github.com/#faqs
- rozab 5y agoThis page has a looping back button hijack for me
- samtheprogram 5y agoThe most important question, whether you own the code, is sort of maybe vaguely answered under “How will GitHub Copilot get better over time?” > You can use the code anywhere, but you do so at your own risk. Something more explicit than this would be nice. Is there a specific license? EDIT: also, there’s multiple sections to a FAQ, notice the drop down... under “Do I need to credit GitHub Copilot for helping me write code?”, the answer is also no. Until a specific license (or explicit lack there-of) is provided, I can’t use this except to mess around.
- netcraft 5y agoI dont see the answer to a single one of their questions on that page - did you link to where you intended? Edit: you have to click the things on the left, I didn't realize they were tabs.
- mwcampbell 5y agoHas this been tested for accessibility yet, particularly with a screen reader?
- r3trohack3r 5y agoIs there a public API? Will it be documented? Are you open to folks porting the VSCode plugin to other editors (I.e. kakoune’s autocomplete)?
- nightski 5y agoI'm glad they find it head exploding but my concern is that it would be most head exploding to newbies who don't have the skill to discern if AI code is how it should be written. For a seasoned veteran writing the code was never really the hard part in the first place.
- amelius 5y ago> For a seasoned veteran writing the code was never really the hard part in the first place. Yes, to most coders this Copilot software is just a fancy keyboard.
- williamdclt 5y agoSounds great. I'm a bad typist, anything that makes me type less (vim, voice assistant, completion) is a big win to me
- amelius 5y agoVim is great because it doesn't try to be smart.
- intricatedetail 5y agoAnd you can't quit it so you are forced to learn it well.
- dang 5y agoOk, we changed the URL to that from https://github.blog/2021-06-29-introducing-github-copilot-ai-pair-programmer/ https://github.blog/2021-06-29-introducing-github-copilot-ai....
- 6gvONxR4sf7o 5y agoIf I put a section in my LICENSE.txt prohibiting use as training data in commercial models, would that be sufficient to keep my code out of models like this?
- mdaniel 5y agoOnly if they trained a model to be able to read and understand LICENSE.txt files -- wowzers what a monster improvement that would be for the world Or, I guess a sentinel phrase that the scraper could explicitly check: `github-copilot-optout: true`
- 6gvONxR4sf7o 5y agoOr it could explicitly check for known standard licenses that permit it, if it were opt in instead of opt out, the way most everything else in software licensing is opt-in for letting others use.
- dragonwriter 5y ago> If I put a section in my LICENSE.txt prohibiting use as training data in commercial models, would that be sufficient to keep my code out of models like this? Neither in practice (because it doesn't look for it) nor legally in the US, if Microsoft’s contention that such use is “fair use” under US copyright law. That “fair use” is an Americanism and not a general feature of copyright law might create some interesting international wrinkles, though.
- 6gvONxR4sf7o 5y agoTheir contention is > Why was GitHub Copilot trained on data from publicly available sources? > Training machine learning models on publicly available data is now common practice across the machine learning community. The models gain insight and accuracy from the public collective intelligence. But this is a new space, and we are keen to engage in a discussion with developers on these topics and lead the industry in setting appropriate standards for training AI models. Personally, I'd prefer this to be like any other software license. If you want to use my IP for training, you need a license. If I use MIT license or something that lets you use my code however you want, then have at it. If I don't, then you can't just use it because it's public. Then you'd see a lot more open models. Like a GPL model whose code and weights must be shared because the bulk of the easily accessible training data says it has to be open, or something like that. I realize, however, that I'm in the minority of the ML community feeling this way, and that it certainly is standard practice to just use data wherever you can get it.
- deleted 5y ago[deleted]
- ggsp 5y agoOne question: how long is the waitlist? Very excited to try this!
- Celenduin 5y agoThis is impressive. And scary. How long has your team been working on this first release?
- shanebrunette 5y agoIs there anyway to port this into emacs?
- fighterpilot 5y agoI assume this will also work alright on other languages that aren't in the demo, e.g. C++?
- jereees 5y agoThe technical preview document points out that editor context is sent back to the server not just for code generation but as feedback for improvement. Are you (or OpenAI) improving the ML models based on the usage of this extension? It is interesting what the pricing will look like given that the model was originally trained on FOSS and then you go and harvest test cases from real users. If that’s the case I think that should be clearly explained upfront.