6 ms·
I find GitHub Copilot close to useless for production code. The worst, most obscure bugs I've had to debug in the last year were all in Copilot-written code. It
by klauserc 2y ago
I find GitHub Copilot close to useless for production code. The worst, most obscure bugs I've had to debug in the last year were all in Copilot-written code. It _looks_ plausible, but it makes extremely subtle mistakes. Occasionally, you have repetitive sections of code where it can copy&adapt lines from the context, but that's about it.
It's a different story for test code. Test code is often formulaic and "standardized" (given/when/then). For instance, I find myself writing the first test case and Copilot can come up with additional test cases. Or I might write the method name ( FeatureUnderTest_Scenario_ExpectedOutcome) and Copilot provides the implementation.
I have not found any value in Copilot chat.
- yuppiepuppie 2y agoAssuming you work on a team with pull requests and code review, how much do you also put blame on that process?
- shaky-carrousel 2y agoCopilot for me is very useful with all the scaffold boring code. It sometimes helps with problems, but I have to guide it, and be very precise with my request.
- zahrc 2y agoAnd even then it happens to ignore context or queries from start or halfway through. I'd rather spend the time coding then trying to bruteforce it to give me the answer I need.
- devjab 2y agoI think automated tests are the one area that LLMs will truly improve productivity (and overall code quality). It’ll likely also lead to a lot of tests that actually tests nothing, but as a whole, it’ll hopefully be capable of both generating and updating tests if you give it some good inputs to do it on. Documentation is another area where I have high hopes. In the ideal world people update it as they change things. In reality, however, well… Then there is the design side of things. I really feel bad for designers of Icons now that you can get some really good one really fast by tasking one of the image generating AIs. I’m not sure LLMs will ever really be capable helpers as far as programming goes. Well I guess it’s two part, they can help with trivial tasks, but they can’t help with anything related to the actual work of generating business value with code. It’s two-sided of course. They certainly allow a lot of people write functioning, though really shitty, code. Which is a huge benefit for a lot of programming tasks where it doesn’t really matter that it’s inefficient and well terrible. We’ve already seen our more digitally inclined employees make great things with power apps, most of which are eventually replaced by more robust software as they scale. But we also see small Python programs helping out with tiny personal tasks around our offices, and while IT operations aren’t too happy it’s generating a lot of individual value that wasn’t there before.
- portaouflop 2y agoI think few if any designers rely on icons for income. There have been thousands of free icons around before genai.
- kqr 2y ago> It’ll likely also lead to a lot of tests that actually tests nothing So pair it with mutation testing!
- jtbayly 2y agoYes, I just used ChatGPT to write me some code to iterate through a CSV and add each row to a system via its API. It wrote a python app. It hard coded the API key and the CSV file. And then it told me to pass the file name as an argument. lol. I just asked it to fix that and tested with a two line csv. Worked like a charm and saved me quite a bit of time trying to figure a few new things out. But a proper programmer would have been slowed down by this, for sure.
- datavirtue 2y agoA test that tests nothing is redundant and therefore is not a test. I have seen people make claims about "useless tests" when they are not able to reason about the coverage. You should be using a tool to gauge test coverage. Tests should be proving accuracy and precision. It's easy to conflate those or lose sight of one.
- collyw 2y agoIf the code isn't doing anything special, it spits out decent enough code (I am using the paid version of ChatGPT with the various customization). As someone who spends 80% of his time in the backend, I find it great for JavaScript whereas it's not so good for Django which I know pretty well.It can still be useful though and is often faster than looking up docs for specific things.
- howtofly 2y ago[dead]
- couchand 2y agoTest code is code. It's as much of a burden as every other piece of code you are troubled with, so you must make it count. If you're finding it repetitive and formulaic, take that opportunity to identify the next refactoring. Just churning out more near copies is not a good answer.
- lolinder 2y ago> If you're finding it repetitive and formulaic, take that opportunity to identify the next refactoring. It doesn't really matter how many helper functions you extract from your test code, in the end you have to string them together and then make assertions, and that part will always be repetitive and formulaic. If you've extracted a lot of shared code, then it might look something like "do this high-level business thing and then check that this other high-level business thing is true". But that is still going to need to be written a dozen times to cover all the test cases, and you're still going to want test names that match the test content. There's a certain amount of repetition and formulaism that will never go away and that copilot is very good at.
- causal 2y agoLLMs are pretty good at anything that follows a pattern, even a really complex pattern. So unit tests often take a form similar to the n-shot testing we do with LLMs, a series of statements and their answers (or in the case of unit tests, a series of test names and their tests). It makes sense to me that LLMs would excel here and my own experience is that they are great at taking care of the low-hanging fruit when it comes to testing.
- nasmorn 2y agoI agree. A very high impact change I made for an application my team is working on was allowing easy creation of test cases from production data. We deal with almost unknowable upstream data and cheaply testing something that was not working out has reduced the time to find bugs tremendously
- AlexandrB 2y agoThe problem with refactoring test code is twofold: 1. It can make it harder to see what's actually being tested if there are too many layers of abstraction in the test. 2. Complex test code can have significant bugs of its own that can result in false passes. What tests the test code? Thus I generally see repetitive or copy/pasted test code as a necessary evil a lot of the time.
- packetlost 2y agoI've found Supermaven to be substantially better than Copilot. The latency is near instant and the results are mostly confined to a line or 2 where the success rate is higher. Meanwhile I agree that Copilot was less than useless for me. Actively hurt my workflow and made things harder.
- zarathustreal 2y agoIf you’re not using a language that can properly support algebraic structures and randomized property-based testing you’re essentially getting no guarantees about your code from tests. You wrote the code, you wrote the tests, they’re equally likely to be incorrect.
- munksbeer 2y agoI find this to be a bit of a meaningless point. What are you actually trying to say?
- zarathustreal 2y agoMy statement is clear and straightforward, I’m not sure how to put it any other way. LLM-generated tests don’t make sense as a concept because there are only roughly five properties you actually need to write tests for if you’re writing tests that actually provide any guarantees.
- munksbeer 2y agoApologies, but I understand the English words you're typing but I'm still not sure of the intent you're trying to convey to everyone. You're conversing in a very rigid style which isn't sympathetic to how people typically interact. I could just leave the discussion I guess, but in the interest of discourse, I don't find your statement meaningful because we're not all working in languages that I think you refer to. Our unit tests are absolutely not perfect and don't offer perfect guarantees, as we're fallible and will write fallible code. And as such, I just don't understand what point you're trying to make by saying that LLM generated tests are no good because they can't offer perfect guarantees.
- zarathustreal 2y agoAh that makes sense to me, I see where I misunderstood you. When you say you don’t understand what you mean is that you do understand but you disagree with the point. I’m on mobile so it’s hard to reference what I previously said but I’m assuming my statement needs to be weakened a bit to be correct. What I meant to say was that unit tests provide essentially no value because they can’t offer perfect guarantees, which is probably different than what I originally said. I’m assuming I just said “they offer no value” which is probably false in some cases for some people and some teams depending on their definition of value. My point was that unit tests do not make sense insofar as their purpose is to provide guarantees about the behavior of code because the information they provide does not meet the standard definition of “a guarantee”. For the above mentioned people/teams/situations/value definitions, they may make sense. Hope that clarifies what I was trying to say. Regarding languages, algebraic structures can be implemented in any Turing complete language. Likewise with property-based testing (with, eg randomized inputs across the domain). I’d be willing to guess it’s just a matter of education and/or desire keeping most developers from using it.
- tracker1 2y agoI've found Github Copilot to be pretty great for boilerplate code... It's a tossup for anything much more complex.