5 ms·
Based on the list, LLMs are at a "very smart junior programmer" level of coding - though with a much broader knowledge base than you'd expect from even a senior
by taberiand 2y ago
Based on the list, LLMs are at a "very smart junior programmer" level of coding - though with a much broader knowledge base than you'd expect from even a senior. They lack bigger-picture thinking, and default to doing what is asked of them instead of what needs to be done.
I expect the models will continue improving though, I feel like most of it comes down to the ephemeral nature of their context window / the ability to recall and attach relevant information to the working context when prompted.
- threeseed 2y agoI wonder if people who say LLMs are a smart junior programmer have ever used LLMs for coding or actually worked with a junior programmer before. Because for me the two are not even remotely comparable. If I ask Claude to do a basic operation on all files in my codebase it won't do it. Half way through it will get distracted and do something else or simply change the operation. No junior programmer will ever do this. And similar for the other examples in the blog.
- zarathustreal 2y agoSince when is “do something on every file in my codebase” considered coding?
- curious_cat_163 2y ago> If I ask Claude to do a basic operation on all files in my codebase it won't do it. Not sure exactly how you used Claude for this, but maybe try doing this in Cursor (which also uses Claude by default)? I have had pretty good luck with it "reasoning" about the entire codebase of a small-ish webapp.
- ohgr 2y agoYep. My usual sort of conversation with an LLM is MUCH worse than a junior developer... Write me a parser in R for nginx logs for kubernetes that loads a log file into a tibble. Fucks sake not normal nginx logs. nginx-ingress. Use tidyverse. Why are you using base R? No one does that any more. Why the hell are you writing a regex? It doesn't handle square brackets and the format you're using is wrong. Use the function read_log instead. No don't write a function called read_log. Use the one from readr you drunk ass piece of shit. Ok now we're getting somewhere. Now label all the columns by the fields in original nginx format properly. What the fuck? What have you done! Fuck you I'm going to just do it myself. ... 5 minutes later I did a better job ...
- fn-mote 2y agoExcept for the first paragraph, I couldn’t tell if you were talking to an incompetent junior or an LLM. I expected the lack of breadth from the junior, actually.
- ohgr 2y agoI swear at the junior programmers less. To be fair the guys I get are pretty good and actually learn. The model doesn't. I have to have the same arguments over and over again with the model. Then I have to retain what arguments I had last time. Then when they update the model it comes up with new stupid things I have to argue with it on. Net loss for me. I have no idea how people are finding these things productive unless they really don't know or care what garbage comes out.
- groby_b 2y ago> the guys I get are pretty good and actually learn. The model doesn't. Core issue. LLMs never ever leave their base level unless you actively modify the prompt. I suppose you _could_ use finetuning to whip it into a useful shape, but that's a lot of work. (https://arxiv.org/pdf/2308.09895 https://arxiv.org/pdf/2308.09895 is a good read) But the flip side of that core issue is that if the base level is high, they're good. Which means for Python & JS, they're pretty darn good. Making pandas garbage work? Just the task for an LLM. But yeah, R & nginx is not a major part of their original training data, and so they're stuck at "no clue, whatever stackoverflow on similar keywords said".
- taberiand 2y agoRight, that is their main limitation currently - unable to consider the full system context when operating on a specific feature. But you must work with excellent juniors (or I work with very poor ones) because getting them to think about changes in the context of the bigger picture is a challenge.
- qingcharles 2y agoThis is definitely a huge factor I see in the mistakes. If I hand an LLM some other parts of the codebase along with my request so that it has more context, it makes less mistakes. These problems are getting solved as LLMs improve in terms of context length and having the tools send the LLM all the information it needs.
- qingcharles 2y agoBut at the same time it'll write me 2000 lines of really gnarly text parsing code in a very optimized fashion that would have taken a senior dev all day to crank out. We have to stop trying to compare them to a human, because they are alien. They make mistakes humans wouldn't, and they complete very difficult tasks that would be tedious and difficult for humans. All in the same output. I'm net-positive from using AI, though. It can definitely remove a lot of tedium.
- nomel 2y ago> and default to doing what is asked of them instead of what needs to be done. I don't think it's that simple. From what I've found, there are "attractors" in the statistics. If a part of your problem is too similar to a very common problem, that the LLM saw a million times, the output will be attracted to those overwhelming statistical next-words, which is understandable. That is the problem I run into most often.
- Groxx 2y agoIt's a constant struggle for me too, both "in the large" and small situations. Using a library which provides special-cased versions of common concepts, like "futures"? You'll get non-stop mistakes and misuses, even if you've got correct ones right next to it, or feed it reams of careful documentation. Got a variable with a name that sounds like it might be a dictionary (e.g. `storesByCity`), but it's actually a list? It'll try to iterate over it like a dictionary, point out "bugs" related to unsorted iteration, and will return `var.Values()` instead of `var` when your func returns a list. Practically every single time, even after multiple rounds of "that's a list"-like feedback or giving it the compilation errors. Got a Clean-Code-like structure in some things but not others? Watch as it assumes everything follows it all the time despite massive evidence to the contrary. They're rather impressive when building common things in common ways, and a LOT of programming does fit that. But once you step outside that they feel like a pretty strong net negative - some occasional positive surprises, but lots of easy-to-miss mistakes.
- taberiand 2y agoOh sure, the flip side of doing what was asked is doing what is known - choosing a solution based on familiarity rather than applicability. Also a common trait in juniors in my experience
- nomel 2y agoRelated, the similarities I see when a human runs out of context window are really interesting. I do a lot of interviews, and the poor performers usually end up running out of working memory and start behaving very similar to an LLM. Corrections/input from me will go into one ear and fall out the other, they'll start hallucinating aspects of the problem statement in an attractor sort of way, they'll get stuck in loops, etc. I write down when this happens in my notes, and it's very consistently 15 minutes. For all of them, it seems to be the lack of familiarity doesn't allow them to compress/compartmentalize the problem into something that fits in their head. I suspect it's similar for the LLM.
- lelanthran 2y ago> I expect the models will continue improving though, How? They've already been trained on all the code in the world at this point, so that's a dead end. The only other option I see is increasing the context window, which has diminishing returns already (double the window for a 10% increase in accuracy, for example). We're in a local maxima here.
- dcre 2y agoThis makes no sense. Claude 3.7 Sonnet is better than Claude 3.5 Sonnet and it’s not because it’s trained on more of the world’s code. The models are improving in a variety of ways, whether by being larger, faster, using the same number of parameters more effectively, better RLHF techniques, better inference-time compute techniques, etc.
- lelanthran 2y ago> The models are improving in a variety of ways, whether by being larger, faster, using the same number of parameters more effectively, better RLHF techniques, better inference-time compute techniques, etc. I didn't say they weren't improving. I said there's diminishing returns. There's been more effort put into LLMs in the last two years than in the two years prior, but the gains in the last two years have been much much smaller than in the two years prior. That's what I meant by diminishing returns: the gains we see are not proportional to the effort invested.
- dcre 2y agoYou said we're in a local maximum. Your comment was at odds with itself.
- taberiand 2y agoOne way is mentioned in the article, expanding and improving MCP integrations - give the models the tools to work more effectively within their limitations on problems in the context of the full system.
- DanHulton 2y agoThis was my thought when browsing this list, too, and it helped crystalize one of the feelings I had when trying to with with LLMs for coding: I'm a senior developer, and I want to develop as a senior developer does, and turn in senior developer-quality code. I _don't_ want to spend the rest of my career in development simply pairing with/babysitting a junior developer who will never learn from their mistakes. It may be quicker in the short run in some cases, but the code won't be as good and I'm likely to burn out, further amplifying the quality issue. > I expect the models will continue improving though I try to push back on this every time I see it as an excuse for current model behaviour, because what if they don't? Like, not appreciably enough to make a real difference? What if this is just a fundamental problem that remains with this class of AI? Sure, we've seen incredible improvements over a short period of time in model capability, but those improvements have been visibly slowing down, and models have gotten much more expensive to train. Not to mention that a lot of the problem issues mentioned in this list are problems that these models have had for several generations now, and haven't gotten appreciably better, even while other model capabilities have. I'm saying this not to criticize you, but more to draw attention to our tendency to handwave away LLM problems with a nebulous "but they'll get better so they won't be a problem." We don't actually know that, so we should factor that uncertainly into our analysis, not dismiss it as is commonly done.
- ezyang 2y agoI definitely agree that for current models, the problem is finding where the LLM has comparative advantage. Usually it's something like (1) something boring, (2) something where you don't have any of the low level syntax or domain knowledge, or (3) you are on manager schedule and you need to delegate actual coding.