6 ms·
Spookily good at writing code? LLMs frequently hallucinate broken nonsense shit when I use them. Recognize what they do well (generate simple code in popular l
by henning 2y ago
Spookily good at writing code? LLMs frequently hallucinate broken nonsense shit when I use them.
Recognize what they do well (generate simple code in popular languages) while acknowledging where they are weak (non-trivial algorithms, any novel code situation the LLM hasn't seen before, less popular languages).
- simonw 2y agoDid you try learning HOW to get good code out of them? As with all things LLM there's a whole lot of undocumented and under appreciated depth to getting decent results. Code hallucinations are also the least damaging type of hallucinations, because you get fact checking for free: if you run the code and get an error you know there's a problem. A lot of the time I find pasting that error message back into the LLM gets me a revision that fixes the problem.
- lolinder 2y ago> Code hallucinations are also the least damaging type of hallucinations, because you get fact checking for free: if you run the code and get an error you know there's a problem. This is great when the error is a thrown exception, but less great when the error is a subtle logic bug that only strikes in some subset of cases. For trivial code that only you will ever run this is probably not a big deal—you'll just fix it later when you see it—but for code that must run unattended in business-critical cases it's a totally different story. I've personally seen a dramatic increase in sloppy logic that looks right coming from previously-reliable programmers as they've adopted LLMs. This isn't an imaginary threat, it's something I now have to actively think about in code reviews.
- simonw 2y agoYeah, the other skill you need to develop to make the most of AI-assisted programming is really good manual QA.
- lolinder 2y agoHave you found that to be a good trade-off for large-scale projects? Where I'm at right now with LLMs is that I find them to be very helpful for greenfield personal projects. Eliminating the blank canvas problem is huge for my productivity on side projects, and they excel at getting projects scaffolded and off the ground. But as one of the lead engineers working on a million+ line, 10+ year-old codebase, I've yet to see any substantial benefit come from myself or anyone else using LLMs to generate code. For every story where someone found time saved, we have a near miss where flawed code almost made it in or (more commonly) someone eventually deciding it was a waste of time to try because the model just wasn't getting it. Getting better at manual QA would help, but given the number of times where we just give up in the end I'm not sure that would be worth the trade-off over just discouraging the use of LLMs altogether. Have you found these things to actually work on large, old codebases given the right context? Or has your success likewise been mostly on small things?
- simonw 2y agoI use them successfully on larger project all the time. "Here's some example JavaScript code that sends an email through the SendGrid REST API. Write me a python function for sending an email that accepts an email address, subject, path to a Jinja template and a dictionary of template context. It should return true or false for if the email was sent without errors, and log any error messages to stderr" That prompt is equally effective for a project that's 500 lines or 5,000,000 lines of code. I also use them for code spelunking - you can pipe quite a lot of code into Gemini and ask questions like "which modules handle incoming API request validation?" - that's why I built https://github.com/simonw/files-to-prompt https://github.com/simonw/files-to-prompt
- gre 2y agoI had some success converting a react app with classes to use hooks instead. Also asking it to handle edge cases, like spaces in a filename in a bash script--this fixes some easy problems that might have come up. The corollary here is that pointing out specific problems or mentioning the right jargon will produce better code than just asking for the basic task. It's very bad at Factor but pretty good at naming things, sometimes requiring some extra prompting. [generate 25 possible names for this variable...]
- polishdude20 2y agoWhen they spit out these subtle bugs, are you promoting the LLM to watch our for that particular bug? I wonder if it just needs a vir more guidance in more explicit terms
- lolinder 2y agoAt a certain point it becomes more work to prompt the LLM with each and every edge case than it is to just write the dang code. I work out what the edge cases are by writing and rewriting the code. It's in the process of shaping it that I see where things might go wrong. If an LLM can't do that on its own it isn't of much value for anything complicated.
- joelanman 2y ago> if you run the code and get an error you know there's a problem. well, sometimes - other times it'll be wrong with no error, or insecure, or inaccessible, and so on
- xyzsparetimexyz 2y agoIs there more to getting 'good' at them then just copying error messages back in? Like, how do I get them to reason about e.g. whether a data structure compression method makes sense?
- AnimalMuppet 2y ago> Did you try learning HOW to get good code out of them? That is at least somewhat a valid point. Good workers know how to get the best out of their tools. And yet, good tools accommodate how their users work, instead of expecting the user to accommodate how the tool works. One could also say that programmers were sold a misleading bill of goods about how LLMs would work. From what they were told, they shouldn't have to learn how to get the best out of LLMs - LLMs were AI, on the way to AGI, and would just give you everything you needed from a simple prompt.
- simonw 2y agoYeah, that's one of the biggest misconceptions I've been trying to push back against. LLMs are power-user tools. They're nowhere near as easy to use as they look (or as their marketing would have you believe). Learning to get great results out of them takes a significant amount of work.
- henning 2y agoLike all AI simps, your blanket response to pointing out flaws is to tell me to do more prompt engineering and then dismiss the issue entirely. In the time it takes me to coax the model to do the thing I was told it knows how to do, I could just do the task myself. Your examples of LLM code generation are simple, easy to specify, self-contained applications that are not representative of software you can actually build a business on. Please do something your beloved LLMs can't and come up with an original idea.
- minimaxir 2y ago> not representative of software you can actually build a business on The only people pushing that you can BUILD AN APP WITHOUT WRITING A LINE OF CODE are the Twitter AI hypesters. Simon doesn't assert anything of the sort. LLMs are more-than-sufficient for code snippets and small self-contained apps, but they are indeed far from replacing software engineers.
- phantompeace 2y agoLike all stubborn anti-AI know-it-alls, you sound like you’ve tried a couple of times to do something and have decided to label all LLMs with the same brush. What models have you tried, and what are you trying to do with them? Give us an example prompt too so we can see how you’re coaxing it so we can rule out skill issue. And a big strength LLMs have is summarizing things - I’d like to see you summarize the latest 10 arxiv papers relating to prompt engineering and produce a report geared towards non-techies. And do this every 30 mins please. Also produce social media threads with that info. Is this a task you could do yourself, better than LLMs?
- henning 2y agoDue to unexpected capacity constraints, Claude is unable to reply to this message.
- phantompeace 2y agoJust as I thought, just snark and no real meaningful engagement. P.S my script uses local models - no capacity constraints (apart from VRAM!)
- CRConrad 2y ago> Did you try learning HOW to get good code out of them? Isn't that a bit "You're holding it wrong"? I mean, why isn't that the default; did anyone really think one would mainly want bad results out of them?
- simonw 2y agoWhen I say that these things are deceptively difficult to use I don't intend that as a ringing endorsement of the technology.
- th0ma5 2y agoSimon gets one thing working for one task and assumes everyone can do the same for everything. That's the trick is that he has no idea how the failures happen or how to maintain actual working systems.