13 ms·
I worry our Copilot is leaving some passengers behind
- nunez 3y ago> Why do we accept a product that not only misfires regularly, but sometimes catastrophically? RIGHT?!! ChatGPT went completely bonkers today, and people are basically ":shruggie: it happens" THAT IS WILD TO ME. People are using ChatGPT for medical advice. People are using ChatGPT for activities that have financial implications. How is this okay?!
- deleted 3y ago[deleted]
- jeffbee 3y agoI work with a guy who is absolutely dedicated to using LLMs to generate C++ code. If I ask him for a specific small thing I'll get back a PR with hundreds of lines of irrelevant crap and when I ask why it has this move constructor or whatever, they won't have a good reason. Even though my colleague is an industry veteran, their new habit has made it feel like they are delegating all their work to the stupidest teammate I've ever had. I feel like we are going to need to work out some norms and customs in this industry for using code-generating systems in a way that respects the time and attention of coworkers.
- stouset 3y agoRequest changes on the PR with the exact same reasoning you would use with any other developer who works like that?
- halfmatthalfcat 3y ago> absolutely dedicated Hard to reason with developers like this, especially if they're more senior than you.
- convolvatron 3y agoit really doesn't matter how many years they have been working, or how old they are, or how long they have been at the organization. we should all agree that someone who offloads the error correction of llms to their teammates isn't someone that's really 'senior'
- stouset 3y agoThen don’t approve their PR since it’s unreviewable in its current form.
- bluefirebrand 3y agoOffloading all of the actual code reasoning onto your team because you cannot be bothered to write the code yourself and are trusting an LLM should get you fired on the spot. I cannot imagine a worse teammate or a worse developer.
- lpapez 3y agoI had one such coworker until recently, and he was actually fired because nobody on the team felt he was pulling his own weight. He produced massive amounts of code which did not fit the style of the codebase at all, and when questioned point-blank if it was LLM-generated he denied it (even thought it was undeniable). I'm all for using tools to boost your productivity, but IMO when you offload generated junk to be reviewed by your team it's a sign of disrespect.
- rvnx 3y agoAsk the PR to be reviewed by an LLM. Enjoy your new life with lot of free time.
- johnny22 3y agountil you get tasked with fixing the buggy code
- ptero 3y agoThis. Good code is clean and has a well thought through internal architecture. LLM-ifying the code and treating it as a black box (if it passes the tests, it is acceptable) is tempting, but it works until it does not and the "does not" might come pretty quickly: once a human cannot easily untangle the logic the only fix is a rewrite. I think there is a way to extend the useful life of such an approach by setting up a good architecture with lean, strict interfaces and thorough tests. Then one can treat any module that is compliant as a black box and give a computer the power to insert as much crap as it can generate. You then should be ready and willing to rewrite any box that has become so convoluted that LLM can no longer fix, likely by splitting it into smaller externally observable and testable elements. I doubt that this is a long-term viable approach, but this is just a personal hunch. It would be interesting to see how such approaches develop. My 2c.
- flappyeagle 3y agoMaybe he needs to get a stern talking to by his manager? Has that happened?
- timeon 3y agoHope manager won't be like: "According to chatGPT..."
- blibble 3y agoif he's doing that the company might as well save his salary and get an intern with a ChatGPT subscription
- halfmatthalfcat 3y agoCopilot is a decent tool for experienced developers (though, hasn't replaced Google-foo by any stretch) and a trap for inexperienced ones. Sure it may be able to speed things up in the beginning but it's a crutch for long-term sustainability in the industry. You inevitably have to understand the paradigms and patterns that LLMs regurgitate; taking them at face value (which I suspect is what most LLM users do), is a recipe for disaster and unfounded confidence.
- skissane 3y ago> Copilot loves suggesting about 25 nested divs as a starting point. > I assume this is because of a flaw in how LLMs work. I know with some LLM implementations, you can configure the sampling to penalise repetitions – this is making me wonder if Copilot might benefit from that? > What does it say about Copilot’s knowledge of accessibility when it will hand us code even basic checking tools would flag? Maybe it could do with some fine-tuning based on those checking tools? e.g. sample many answers to same prompt, run them through checking tool, and then fine-tune it to prefer the answers which caused the least warnings? Or: run the suggestion through checking tools, and if it triggers warnings, sample a new suggestion, and see if the new one doesn't. This could be done on the client side in a loop – run suggestion through checks, if it fails, ask the LLM for a new and different suggestion, repeat until we get one which passes checks, or we give up.
- teaearlgraycold 3y agoThe issue is that often you do want heavy repetition when programming. Think about a list of strings where they mostly have a common prefix. Or JSON, or a bunch of imports and exports. Good code often has these low entropy sections.
- skissane 3y ago> The issue is that often you do want heavy repetition when programming. Up to a certain number of tokens, yes. But, I doubt any high quality code would have the exact same sequence of N tokens repeated 25 times consecutively. There's a certain heaviness of repetition at which it is unlikely to be genuinely useful. > Or JSON, or a bunch of imports and exports A human programmer, when evaluating whether code is repetitive, doesn't treat all tokens as equal – they ignore "expected"/"necessary" repetitions, and focus on the "unexpected"/"unnecessary" ones. So, penalising repetitions in sampling doesn't have to treat all tokens equally either. For example, in a JSON document, one might choose to ignore the tokens required by JSON syntax. In Java, one might penalise repetitions less in the import block than in a method body. Of course, this means the sampling actually has to be aware of the syntax of the language being generated – which is possible, and can have some other advantages (e.g. if sampling only samples tokens which are allowed by the language grammar, you can eliminate many possibilities of generating syntactically invalid code.)
- Doches 3y agoThis is one of the more thoughtful, nuanced criticisms of the current LLM fad that I've read, and I'm delighted to see it make it show up on HN. The author starts off with a series of well-thought experiments that show Github Copilot generating _pretty valid_ frontend code, code that works and fulfills the prompt: but code that ignores every web accessibility rule of thumb in the most egregious ways. Sure, yes, bad web devs write bad code, and Copilot is -- on its best day -- a perfectly cromulent bad developer. Yawn, news at 11, etc. But where he takes those examples and where his thoughts end up is where this essay really hit home for me: > As more and more of the internet is generated by LLMs, more and more of it will reinforce biases. Then more and more LLMs will consume that biased content, use it for their own training, and the cycle will accelerate exponentially. And 'biases' here isn't the usual "models are woke-lobotomized!" yammering, but rather a thoughtful take on how the use of LLMs for code generation may, at least for the current state of LLMs, slowly normalize _writing worse code_.
- rebolek 3y agoSo that’s nothing new. Code is getting worse for decades. Moore’s law is making worse code acceptable. In the meantime, some people write better code and do care about it and LLMs aren’t going to change that. So there will be worse code and there will be better code as always. LLM is just a tool.
- timeon 3y agoThis is true but let's not forger that it is race to the bottom. Even this blog did reshaped itself while I was reading it. Moore’s law is lagging here.
- here4U 3y agoIt is clear that despite these tools having flaws on the whole they save a lot of time. It is not clear what the tradeoff with introducing poorly understood or faulty code will bring, but given the utility we're never going back.
- callamdelaney 3y agoHalf the time copilot doesn’t even return a solution
- Smaug123 3y agoThis is a good thing in the context of the article (it even explicitly says "They’re not made to give you verifiable facts or to say 'I don’t know'" in a context which suggests this is in fact a bad trait). Better to return no code than to return crap code.
- callamdelaney 3y agoIt would be a good thing if the requested solution wasn’t extremely simple
- timeon 3y agoBetter no-solution than cognitive overload with bad solutions.
- plondon514 3y agoIs it just me or has copilot gotten progressively worse lately? It used to feel like it was making well informed guesses, now they feel like literal guesses with no context at all. For example in my phoenix live view (elixir) app it guesses “xxx@xxxxx” for _any_ attribute I pass in to a component.
- stanleydrew 3y ago> Shouldn’t the results I get from a paid service at least be better than a bad StackOverflow suggestion that got down-voted to the bottom of the page (and which would probably come with additional comments and suggestions letting me know why it was ranked lower)? I don't know why you would expect this, when the model is likely trained on StackOverflow material (or similar publicly available code examples).
- godelski 3y agoI think the problem is that any tool like this (even one theoretically much more powerful) is most beneficial to those that need it the least and least beneficial to those that need it the most. If you're an expert you can identify the mistakes and they are not generally a roadblock. But if you're a novice you can't and you'll simply be unaware of any hallucinations. The benefit SO has over this is just the extra friction of needing to copy paste or retype because it slows you down and forces an opportunity to think. My worry is that we become too reliant on tools and outsource our thinking to them before they are ready to take on that task. This will only accelerate the shitification of things we have. More apps that use far too many resources. Things that are security nightmares. Interfaces with more friction. All of it. The problem is pareto efficiency. 80% of your code is written in 20% of your time but 80% of your time is is required for 20% of your code. The problem is that the devil is in the details. So even a 95% or 99% accurate code generator is going to make for hard work. That's 1 in every hundred lines of code. I hope the compilers people are writing good error messages.
- jimbob45 3y agoMy worry is that we become too reliant on tools and outsource our thinking to them before they are ready to take on that task. Industry greats like Spolsky have been beating this drum for decades [0] with no success. Those with natural curiosity will gravitate towards understanding the low-level mechanisms of things, just as they always have. Others won’t. [0] https://www.joelonsoftware.com/2001/12/11/back-to-basics/ https://www.joelonsoftware.com/2001/12/11/back-to-basics/
- heisenbit 3y agoThe key issue is not the power of the tool but the tool powerfully amplifying practices that ought to be resisted but exist in the majority of the code in the wild.
- godelski 3y agoThe problem with move fast and break things is that you need to at some point slow down and fix things. But we've developed systems that incentivize never stopping and so just enshitify everything. You win by not having no shit, but by being ankle deep in shit rather than waist deep. By being less shitty.
- bluefirebrand 3y agoI agree with most of the stuff in this article but I'm a bit puzzled by some of the attitudes of the author. They seem to care more about the LLM delivering code with poor accessibility than they care about the LLM delivering completely wrong answers. The "any good developer would realize this is bad code" rings strongly as "no true developer would think the LLM's bad answer was correct". Seems like a short sighted opinion to me. I also think you can replace "accessibility" with any number of programming meta concepts and find problems too. How about "internationalization"? Are LLMs any good at producing code that is nicely internationalized? Or more importantly "security". Are LLMs going to produce millions of lines of poorly secured code that people never double check? Almost assuredly. The fact is that LLMs are prediction engines. They run off of probabilities based on the prompt and the training model. Thus, unless the training model is weighted towards cherry picked examples of excellent code, it's going to follow the masses. And the masses write bad-to-average code mostly.
- wlesieutre 3y ago> They seem to care more about the LLM delivering code with poor accessibility than they care about the LLM delivering completely wrong answers. I think the idea with this is that if it gives you completely wrong answers and the code doesn't work, it will obviously not work and have to figure out how to fix it. Meanwhile when it gives you code that appears to do what you wanted except the accessibility is broken, you'll ship it because you don't realize there's anything wrong with it.
- bluefirebrand 3y agoThe problem is that often it will give answers that are only subtly wrong, and those will get shipped too. I think my puzzlement is with the focus on accessibility as though it was a high priority item. In my experience it's usually an afterthought, if it's a thought at all. Personally I've never worked on a codebase where accessibility was in the top 5 priorities. No one would ever block a prod release for an accessibility mistake. But like I said, you could take this whole argument, find+replace "accessibility" with "security" and you would have a much more compelling argument imo. Given time constraints, code should prioritize security over accessibility basically always.
- wesleyyue 3y agoI think a lot of these are actually solvable problems today, Copilot just hasn't prioritized actually improving the product (don't need to improve the product for breakneck growth when you have github.com as a distribution channel!) It feels like there are a lot of well-intended AI coding products that just don't pay attention to getting the details right. I actually started building my own extension recently, with an emphasis on getting all the little things right, because I got so frustrated at Copilot. Things like closing brackets properly, not interrupting me and destroying my train of thought when writing comments, not suggesting imports unless it's highly certain (or verified with intellisense), etc. Like why am I wasting my precious time talking to copilot chat with gpt3 when gpt4 exists? It's still a pretty early version, and ultimately we're using the same underlying model for completion, but I think getting these details right make a huge difference (at least to my biased self). If you want to try it: https://marketplace.visualstudio.com/items?itemName=doublebot.doublebot https://marketplace.visualstudio.com/items?itemName=doublebo... You'll need to install the pre-release version for auto-complete.
- throwanem 3y agoI'll use a locally hosted Llama 2 or CodeLlama instance as a 'consultant', via a chat window. These models can be great for that! A well-formulated question often elicits a precise and accurate answer, even from the unspecialized model. I won't use Copilot or anything else that integrates that tightly into my workflow, even though it is now possible to do so without losing the incremental-cost and customizability benefits of selfhosting. The context switch is important. To a very good first approximation, our task as engineers is to think before we assume, and I have found Copilot recklessly encourages the latter at the expense of the former.
- _flux 3y agoI wonder though if Copilot had fared better here had it been told to pay attention to accessibility. I mean, maybe it should do it by default (and maybe it could be part of its system prompt or otherwise in its material), but it's still a tool that needs some expertise for using, even if it's trying its best to trick people into believing otherwise. Ultimately I don't think there's a solution to people misusing tools. Paraphrasing sentiment I don't quite recall exactly: "If anyone can do it, then anyone will."
- rafram 3y ago> In a lot of ways, in fact, “AI” is just the newest iteration of a very old form of colonial capitalism; build a wall around something you didn’t create, call it yours, and charge for access. (And when the natives complain, call them primitive and argue they’re blocking inevitable progress.) This is pithy, but the dynamic between OSS devs and Microsoft/OpenAI is not exactly comparable to the dynamic between a colonial government and an indigenous population. I don’t think it really needs to be said, but open-source maintainers are not colonized natives. Even overlooking the very questionable metaphor, they’re not building a wall around existing repositories of code and selling them back to us. They spent a lot of money training an AI model on that code, and now they’re selling access to that model. You don’t need to pay Microsoft for access to the GitHub repos or Stack Overflow answers that they trained on.
- g-b-r 3y agoThese code generation systems should probably prepend a hidden "Generate accessible, secure, maintainable etc code" prompt Of course that doesn't provide any guarantee, and no developer should rely on it, but the average results would probably be a little better
- joenot443 3y agoI've come to largely agree with the author. These days, I keep it off by default, but is a Cmd+' away from being flipped on and filling in what I _know_ to be boilerplate that's well suited. If I was younger with less money, I probably couldn't justify the price, but these days if it can save me a half hour of busywork per month on my personal projects the $10 is more than worth it. Leaving it on while doing any thoughtful or challenging coding is super distracting for me.
- input_sh 3y agoAs someone younger with less money... well, I'm "lucky" enough to get it for free (fun fact: GitHub just gives it away in perpetuity to accounts above certain threshold of "karma"), so I use it. If I didn't get it for free... well I'd be lying if I said I don't get any value out of it, but you're spot on, definitely not enough to justify its perpetual subscription.
- samatman 3y agoThe vacuum cleaner analogy didn't land for me. I've bought that vacuum cleaner, and I didn't return it. I'm referring to one of the countless models of robot vacuum, of course. They clean the floor, most of it, most of the time, but they miss spots, they get stuck on things, and they don't have the suction of a full-size vacuum. I wish none of those things were true, but it saves me labor nonetheless, so I kept it. I can detail corners and pull the thing off the corner of the rug, and still get a mostly-clean floor, automatically. Sure, it doesn't get all the schmutz out of carpets, but it gets enough that I can go over them monthly instead of weekly. Yes, I'm talking about LLM code assistants. They have embarrassing failure modes, but experienced developers get a sense of what they can and can't do, and the result is something which saves time. I've found they're particularly good at "dumb debugging", where there's some fat-fingered error in the code and I can't spot it just by looking. I can copypasta the function into ChatGPT in seconds, and it gives a step-by-step description of what the code does, which routinely points out exactly where the bug is. I have my concerns about what these tools will do to the up-and-coming generation of developers, it's easy to imagine them as a crutch, training wheels which never come off. But that's a separate matter, and I trust that the more natively talented juniors will recognize the hazard there, and understand that a chatbot can't substitute for becoming a skilled programmer.
- coffeebeqn 3y agoI use them quite a bit. Write these tests for this function in format x, transform this struct with lots of fields in the manner y or just good old rubber ducking about a problem I’m having trouble debugging
- throwawaysleep 3y ago> but it saves me labor nonetheless, so I kept it. Yep. Trading accessibility and usefulness to the people who disable JavaScript for 20% more productivity is a bargain. Most companies make that trade for far less every day. I’ve never worked in a place that gave much thought to those. At most there was a contracted dev in some low cost country to slap aria tags around.
- deleted 3y ago
- throwuxiytayq 3y agoI am somewhat amused by all of the "copeelot bad" articles, and I dearly hope they keep proliferating, so that those of us who enjoy its frankly insane productivity boost get to stay ahead of the competition. I perceive no quality/reliability drawbacks in my own code. If anything, the ability to iterate more quickly makes my code better than ever. It's a skill issue. (You had it coming.)
- Karellen 3y ago> I perceive no quality/reliability drawbacks in my own code. How can you be sure that doesn't say more about you than it does about copilot?
- throwuxiytayq 3y agoI'm pretty sure, as I constantly judge and monitor the quality of my code. But thanks for immediately disregarding my personal experience and inserting your own uninformed prejudged assessment, random internet guy.
- notpachet 3y ago> thanks for immediately disregarding my personal experience and inserting your own uninformed prejudged assessment Isn't that exactly what your toplevel post is doing? Physician, heal thyself!
- throwuxiytayq 3y agoI hesitate to engage in this hopelessly fruitless discussion, but the answer is no. I don't even express my opinion of the article, arguably barring one humorous phrase that refers to the currently-fashionable wave of Copilot criticism. I don't mind the article. It's actually pretty well-written. None of this is incompatible with my statement that in my experience, Copilot lets me do my job better. Time to get off the internet, physician.
- Ologn 3y agoRedmonk says Kotlin is the 17th most popular programming language ( https://redmonk.com/sogrady/2023/05/16/language-rankings-1-23 https://redmonk.com/sogrady/2023/05/16/language-rankings-1-2... ). So can any of these LLMs and whatnot, even the ones supposedly geared toward programming do something like this: "Write a function in Kotlin that take a Long as a parameter, and sends back a List containing Long types. The parameter is a number, and the return is a list of prime numbers less than that number. All in one function." It seems it should be pretty simple, in fact I have written this program a number of times. If you think a list of prime numbers might take up too much memory, I have also done prompts only asking it to just give the largest prime under the input parameter. It is not a difficult task, and Kotlin is between Objective-C and Rust in popularity. Have any neural network programming tools been able to complete this? No. Some can, if the number input is 18L or the like. None have been able to handle 600851475143L (taken from the third Project Euler). If the program runs at all I get "java.lang.OutOfMemoryError: Java heap space". Even if I warn it to watch heap memory, it still is the same result. As I said, this is a prompt for a list, but even if I ask for only the largest prime number before 600851475143L, or any long such as that number, I have not seen any LLM or the like that can write that function. Especially ChatGPT 4, which I have tried it on extensively. I'm not saying LLMs will not get there, but this part of the third question on the Project Euler site, from a fairly popular language. It's a pretty simple question - a straightforward function to write. They can't do it yet. I see people worrying about AI being on the verge of taking programmers jobs. Until it can do something incredibly specified and simple as this, I am not worried at all.
- Smaug123 3y agoThe list of primes below 600851475143 contains 23038900221 elements. If each element is a long, that takes a little over 184GB (decimal) of storage. May I ask how you managed it without running out of memory? (Project Euler 3 asks for a factorisation, which using the most memory-hungry but reasonable algorithm would require a list of merely sqrt=775146 in length, which is much more manageable.)
- pavlov 3y agoPersonally I think of LLM code helpers as a warning smell. If I’m working on something where I’m tempted to generate a bunch of boilerplate from a bot that knows very little about the context of the project, am I really spending my time on the right thing? Either I should be working on something higher level, or the amount of boilerplate should be so low that I can write it myself. Anything else suggests that there’s a problem and the LLM bloat band-aid isn’t the solution.
- coffeebeqn 3y agoDepends. We have a large CRUDv service - lots of endpoints that has a lot of boilerplate and not much business logic but it doesn’t change often. It’s annoying to add a new endpoint but it’s not common enough that I want someone to spend a month+ refactoring it
- lostmsu 3y agoI'm sure you generate all your serialization code by hand. Not to mention the object model. Why let compiler make a virtual call if you can load the class pointer and look up the right function yourself, MIRIGHT?
- jononor 3y agoSerialization code should be generic, no boilerplate needed. Object models must reflect the problem domain well, something to think carefully about. Current code assistants is only able to do that automatically for very generic classes.
- hn92726819 3y ago> I'm sure you generate all your serialization code by hand. Is this sarcasm? Who generates serialization/deserialization by hand now? Even Java has mature annotation libraries so it can be done in a single line. Maybe legacy code, but I'd argue using IDE method generation would be much preferred on a legacy codebase than AI.
- pavlov 3y agoBut that’s the whole point of what I meant by “I should be working on something higher level.” Nobody (hopefully) uses an LLM to write mounds of x86 assembly code instead of letting the compiler do it. An AI model is a very blunt tool for most jobs. If I need code generation, most often I should write my own tool for it rather than copy-paste reams of barely inspected code from a language model.
- dexwiz 3y agoCopilot is a competent coder. If I tell it to generate a function with certain parameters, a class that follows a Gang of Four pattern, mass rename variables, or refactor loops into maps, then it does a pretty good job. Copilot is a bad engineer. If I tell it to build something, unless exceedingly simple, it usually fails. The ability for it to create something seems correlated to how many 5 minute tutorials for that exist on the internet. Which given its training, makes perfect sense. So if I think I could find an answer on Stack Overflow, then I will just ask Copilot instead. 10 Years ago everyone was afraid of the Stack Overflow developer, now its the GPT developer. I think its a combination of actual worry and hurt pride that your job can be accomplished by someone copy/pasting. But like usual, good engineers will learn to think for themselves when leveraging tools. And an exceedingly amount of code produced works but is bad by some arbitrary metric. I think the footer example is hilarious, because its exactly inline with web development trends of the last decade. Why use native elements when I can script my own behavior in Javascript on a div? And in a rush to "not use tables for formatting," I bet there are some 25 nested div websites out there. Even on Google sites, I have see grids built using absolutely positioned boxes with Javascript layout logic. The web is a wild place once you start looking past the tutorials and best practices.
- rikafurude21 3y agoultimately these are the kind of things programmers care about. code debt is real and anyone who has any experience having to pay the debt off usually learns their lesson and does a better job on the next try
- robocat 3y ago> Even on Google sites, I have see grids built using absolutely positioned boxes with Javascript layout logic Possibly a side-effect of the framework used. Some frameworks used HTML like a CANVAS and drew/layout everything in HTML using absolute positioning. Ugggh.
- CarefreeCrayon 3y agoOne concern that I have is that copilot is inherently additive in nature. It is unable to suggest that blocks of code be deleted which creates a bias that adding more code is always the solution and a lot of code that shouldn't be written ends up in the codebase. I believe this is a problem works against less experienced engineers because more senior engineers are better at recognizing that problem. In my experience the most senior engineers respond to that by just turning the tool off.
- idempotent_ 3y agoWould be very cool to be able to highlight sections or even an entire file and then have a right-click option to "Refactor Code" which, rather than being additive, would clean up and condense the code according to the idioms of the language.
- Nevermark 3y agocd /source/commercialOS/ lintbug -fix **/*.code refactor **/*.code echo -e "\a"
- perryizgr8 3y agoYou can do this with github copilot. It works exactly like you describe.
- 5Qn8mNbc2FNCiVV 3y agocody.dev can do that and it's pretty good. Running it on some late night coding sessions and I like the output around 90% of the time
- al_borland 3y agoTo be fair, this is a common human bias as well. I’m generally the lone voice advocating for solving problems by subtraction, rather than addition. Though I will admit, automating this addition, and thinking that copilot suggesting it is permission, does make the problem worse.
- PH95VuimJjqBqy 3y ago
- runarberg 3y ago> In a lot of ways, in fact, “AI” is just the newest iteration of a very old form of colonial capitalism; build a wall around something you didn’t create, call it yours, and charge for access. (And when the natives complain, call them primitive and argue they’re blocking inevitable progress.) What a wonderful analogy. LLMs also feel very pythagoran, where a secret cult (of capital owners; the bourgeoisie) guards the secrets of forbidden math, using it for their own benefits, and denying it to the masses. The amount of data and computing power needed to train a good model means it is pretty much inaccessible to the masses, the public can only ever hope to use an already trained model which is provided to us by this secret cult.
- cwkoss 3y agoDoes anyone who has found a good workflow with copilot have a good resource to share that demonstrates how to get the most out of it? I really want it to be more useful, but rarely find that it's helpful for completing more than a single line or two. Do you write out comments for everything you're going to do and then just write it yourself if the suggestion isn't useful? Is there a trick to getting it to read your code itself across files?
- layer8 3y agoIt’s appalling that most website/app developers even have to deal with those kinds of low-level considerations, after decades of web-tech evolution, instead of using a UI builder tool (or a better UI modeling language) that provides all the building blocks for the most common 97% of use cases, and where you would need to go out of your way to create a non-accessible link. TFA is right about LLMs, but it’s also an indictment of the web UI stack.
- taway_6PplYu5 3y agoBy a UI modelling language, do you mean HTML or Javascript or CSS?
- layer8 3y agoI mean something more suitable than the HTML+JS+CSS combination. HTML is a document markup language, not a UI definition language. CSS mixes layout with styling, which are largely orthogonal. (One should be able to specify a UI independently from styling/theming.) A programming language like Javascript shouldn’t be needed for building a UI, in the majority of cases. Most UI components and behaviors should be standard (built into browsers, or whatever is used as the UI runtime) and declarative.
- dilyevsky 3y agoSo basically you’re saying someone should make mui or any other component library except built into browser. Why is that an improvement?
- layer8 3y agoIt would be more than a component library, but yes, more akin to traditional GUI frameworks in spirits than the HTML+CSS approach. The improvement is that it would dramatically reduce the effort needed to implement and maintain web application, make app behaviors more consistent for users, and make adherence to accessibility and other best practices the default, rather than something developers have to actively put significant effort into.
- simonw 3y agoCopilot is bad at accessibility because web engineers are bad at accessibility. All of the bad habits in this post were learned from its training data. That's not to say this can't be fixed: a recurring lesson of LLMs is that the quality of the training data is /everything/. OpenAI made their models better at chess by feeding in higher quality chess data - they could absolutely make it better at accessible frontend code by curating and boosting better code examples. I doubt they'll do that any time soon, purely because there are so many other training data projects they could take on. Thankfully we aren't nearly as dependent on a few closed research labs as we used to be. It would be very exciting to see fine-tuned openly licensed models that target exactly this kind of improvement.
- EscargotCult 3y ago> Copilot is bad at accessibility because web engineers are bad at accessibility. All of the bad habits in this post were learned from its training data. 100%, and this is why Copilot is damn-near unusable for Bash scripting (yeah, the real problem is Bash scripting, use a better scripting language etc etc, but I do it, you've probably done it, and we've all definitely worked with codebases with Bash script linchpins) - there's a lot of bad Bash out there.
- d_sem 3y agoI don't know if its mindset or my owner ignorance, but I find myself using Copilot and other language model tools as a teacher, a debugger, a reviewer, and and idea brainstormer. I find each use case to enhance my ability to think more deeply about my code and helps keep me more engaged in problem solving. For some reason I find a inference from a compressed model which contains almost every notable open source program written in the history of humanity to be a decent sidekick. My experience tells me no software engineer is an expert at everything. Having a tool which allows us to try new things faster is a good thing.
- skybrian 3y agoNo mention of testing in the article. It seems odd how often accessibility advocates talk about following rules rather than testing. Shouldn’t we be testing with screen readers or something? If a website doesn’t work in Firefox, we fault the developer for not testing it in Firefox. Similarly for mobile browsers. If testing is in place, LLM’s are much safer to use. You’ll notice when they give you code that doesn’t work.
- asadotzler 3y ago100% failure isn't really a useful test, is it?
- skybrian 3y agoI don't know what you mean. With test-first development, you write a failing test and then you fix it.
- tydunn 3y ago> Copilot is encouraging us to block users unnecessarily, by suggesting obviously flawed code, which is wrong on every level: wrong ethically, wrong legally, and the wrong way to build software. I share many of the same worries as the author. This is why I think teams need to build and run their own Copilot-like systems, so that they can guide the suggestions they receive. Each developer and team has their own way of building software, and they need to be able to shape and evolve the suggestions they receive to fit their definition of the "right" way: https://blog.continue.dev/its-time-to-collect-data-on-how-you-build-software/ https://blog.continue.dev/its-time-to-collect-data-on-how-yo...
- htfu 3y agoIt's a very powerful autocomplete. "It doesn't generate all the code I need in full and if it does I have to poke at it" is just poor criticism. You don't have to press tab and insert everything it suggests. It will usually generate me half a line after typing the first half - that's pretty awesome in my opinion. If you stick to using it to merely speed-spell out what you were in fact already in the process of writing, and ignore 90% of the terrible crap it proposes, it's a nice productivity boost and has no way to make code worse by itself. Basically, instead of writing a big comment and then a function signature and expect it to do the rest, just start writing out the function, tab when it gets it, don't when it doesn't, or (most of the time) tab then delete half of it and keep the lines you intended, likely with some small tweak. Surely LLMs will be able to go so much more and without constant supervision in the future, but we're not there. That doesn't mean they're bad. Especially copilot since it's just there with its suggestions and doesn't require breaking flow to start spelling out in regular text what you're doing.
- khalilravanna 3y agoThis sounds like it mirrors my usage. Basically treat it like pairing with a really junior dev: assume everything it writes will be wrong and then go from there. If you do that then best case it speeds you up and worst case you waste a little time reading what it wrote that was wrong and ignoring the suggestion and moving on.
- horns4lyfe 3y agoThat’s fine, but it already exists, i.e. resharper
- Qwero 3y agoWhat we will see is that llms become so good in writing code that LLM first will emerge. LLM first means we will test it against our libraries, best practices and potentially even create a new language for it. Then programming in the classical sense won't exist anymore. The ara of code will end when we will deploy the first code written with LLM to write new code. Javallm or #llm. It might be full of examples for a LLM, it might focus on analyzing logic and fixing it on a higher level, until the AI is good enough to self write, evaluate and deploy it. After that it will become no longer understandable by us and researchers will start analyzing it after it was written. Historians will start tracking when ai started to create more efficient abstractions etc.
- Barrin92 3y agoThe biggest problem with Copilot/LLMs is that they effectively operate against anything that programming languages were designed for. What makes programming languages special is that they're well defined, semantically and syntactically rigorous and intended for machine execution. They give us the capacity to formally reason. Instead what we've got now is tools that literally argue with us, rather than anything that actually augments my capacity to reason about, inspect and understand the real performance and hardware of a system my code runs on. What I need is more Coq and less of something that just makes natural language suggestions. What makes a good engineering tool is something that can look at the code right there as it is, use the formal guarantees that programming languages were designed for and give me some verifiably correct suggestions. Not average out 90% of Stackoverflow answers and then hallucinate up some statistical response. Contrast Copilot with tree-sitter. What makes tree-sitter so good as a tool is that it leverages the regularity of programming languages. It can parse and correctly reason about code, instead of relying on some random regex collections and prayers. We've had so many good advances in recent years like the borrow checker in Rust. Why are we going back now and introducing tools that are by design incapable of ensuring correctness? Just to type a little bit faster?
- kromem 3y agoIt's only a cause for concern if its capabilities are going to plateau. More likely, advances in the field will mean that we end up in a more accessible world, where developers who don't normally think about accessibility have a generation engine doing a pass over their work adding appropriate labeling, fixing elements to work with screen readers, etc. We just had a big paper about using genAI to improve test coverage. And we haven't even really hooked LLM code generators up to linters and test suites broadly yet. I can foresee a future where language specific Copilot features might include running suggested generations for HTML though an ARIA checker while running Python generations through a linter, etc. Especially when costs decrease and speed increases such that we see multiple passes of generation, this stuff is going to be really neat. I still mostly consider the tech (despite its branding) in the "technical preview" stage moreso than a "finished product," and given the capabilities at this stage plus the recent research trends and the pace of acceleration, it's a very promising future even if there's valid and significant present shortcomings.
- nailer 3y ago> Copilot loves suggesting about 25 nested divs as a starting point. To be fair it costs a huge amount of money to hire a React/Tailwind person to create 25 nested divs as a starting point.
- heelix 3y agoOn a personal note, I'm in this picture. I've been doing back end code for decades, and beyond a bit of 2003 style ajax/css, my front end skills were non-existent. One of those itches - standing up a blog - as I'm working on my Rust skills, I looked at the tailwind components and started using them. And down the rabbit hole I went to understand what was there and how to use it. On a whim, I asked copilot, interpret this style - and I'll be damned if it it did not produce useful results. I also understand there is an entire ecosystem I'm oblivious to. I'll grumble about the Java code it generates. Mostly meh. As I look at the Rust suggestions, it seems fine. Is it Rust has better training data, or is it that I'm a weak Rust coder still and don't know right from working. My money is on the latter. Anyhow... considering the hours I spent yesterday trying to implement light/dark mode on some simple pages, your Tailwind comment resonates.