8 ms·
I still don’t understand how people are getting value out of AI coders. I’ve tried really hard and the commits produced are just a step up from garbage. Writing
by fovc 2y ago
I still don’t understand how people are getting value out of AI coders. I’ve tried really hard and the commits produced are just a step up from garbage. Writing code from scratch is generally decent. But after a few rounds of edits the assistant just starts piling in conditionals into existing functions until it’s a rats nest 4 layers deep and 100+ lines long. The other day it got into a loop trying to resolve a type error, where it would make a change, then revert it, then make it again
ETA: Sorry forgot about the relevancy in my rant! The one area where I’ve found the AIs helpful is enumerating and then creating test cases
- godelski 2y agoI have a simple answer for you: most people write garbage code[00]. I know this because I write garbage code but usually less garbage. A little background... I got my undergrad in physics (where I fell in love with math), spent some years working, and got really interested in coding and especially ML (especially the math). So I went to grad school. Unsurprising I had major impostor syndrome being surrounded by CS people[0], so I spent a huge amount of time trying to fill in the gap. The real problem was that many of my friends were PL people, so they're math heavy and great programmers. But after teaching a bunch of high level CS classes I realized I wasn't behind. After having to fix a lot of autograders made by my peers, I didn't feel so behind. When my lab grew and I got to meet a lot more ML people, I felt ahead, and confused. I realized the problem: I was trying to be a physicist in CS. Trying to understand things at very fundamental levels and using that to build up, not feeling like I "knew" a topic until I knew that chain. I realized people were just saying they knew at a different threshold. Back to ML: Working and researching in ML I've noticed one common flaw. People are ignoring details. I thought HuggingFace would be "my savior" where people would see that their generation outputs weren't nearly the quality you see in papers. But this didn't happen. We cherry picked results and ignored failures. It feels like writing proofs but people only look at the last line (I'd argue this is analogous to code! It's about so much more than the output! The steps are the thing that matters). So there's two camps of ML people now: the hype people and "the skeptics" (interestingly there's a large population of people with physics and math backgrounds here). I put the latter in quotes because we're not trying to stop ML progress. I'd argue we're trying to make it! The argument is we need to recognize flaws so we know what needs to be fixed. This is why Francois Chollet made the claim that GPT has delayed progress towards AGI. Because we are doing the same thing that caused the last AI winter: putting all our eggs in one basket. We've made it hard to pursue other ideas and models because to get published you need to beat benchmarks (good luck doing so out of the gate and without thousands of GPUs). Because we don't look at the limitations in benchmarks. Because we don't even check for God damn information spoilage anymore. Even HumanEval is littered with spoilage, and obviously so... There's tons of uses for LLMs and ML systems. My "rage" (as with many others) is more about over promising. Because we know if you don't fulfill those promises quickly, sentiment turns against you and funding quickly goes away. Just look at how even HN went from extremely positive on AI to a similar dichotomy (though the "skeptics" are probably more skeptical than researchers.[1]). Is playing with fire. Prometheus gave it to man to enlighten themselves but they also burned themselves quite frequently. The answer is: you evaluate in more detail than others. [00] of course it is. LLMs replicate average human code. They're optimizers. They optimize fitting data, not fitting optimal symbolic manipulation. If everyone was far better at code, LLMs would be too. That's how they work [0] boy, us physicists have big egos but CS people give us a run for the money [1] I have no doubt that AGI can be created. I have no doubt we humans can make it. But I highly doubt LLMs will get us there and we need to look in other directions. I'm not saying we shouldn't stop perusing LLMs, I'm saying don't stop the other research from happening. It's not a zero sum game. Most things in the real world are not (but for some god damn reason we always think it is)
- nzach 2y ago> the commits produced Maybe this is the problem ? I quite like using LLMs for coding, but I don't think we are in a position where a LLM is able to create a reasonable commit. For me using LLMs for coding is like a pair programming session where YOU are the co-pilot. The AI will happily fill you screen with a lot of text, but you have the responsibility to steer the session. Recently I've been using Supermaven in my editor. I like to think of it as 'LSP on steroids', it's not that smart but is pretty fast and for me this is important. Another way I use LLMs to help me is by asking open-ended questions to a more capable but slower LLM. Something like "What happens when I read a message from a deleted offset in a Kafka topic?" to o1. Most of the time it doesn't give great answers, but it generally gives good keywords to start a more focused Google search.
- fovc 2y agoThe “pair programming” approach with good models is just slow enough that I lose focus on each step. The faster models I’ve tried are not good enough except for straightforward things where it’s faster to just use emacs/LSP refactoring and editing tools. Maybe supermaven manages to beat the “good enough, fast enough” bar; I’ll have to try it!
- nzach 2y agoOne thing I've realized after using a really fast model is that the time it takes the model to generate a suggestion is proportional to the size of suggestion I'm willing to accept. And in my experience the quality of suggestions decreases when the suggestion size increases. If the model takes a couple seconds to generate a suggestion I get inclined to accept several lines of code. But if the suggestion takes just 300ms to generate I don't feel the "need" to accept the suggested code. I'm not really sure why that happens, maybe that's just the sunk cost fallacy happening right in my editor? If I wait 5 seconds for a suggestion and don't use the suggestion did I effectively just wasted 5 seconds of my life for no good reason?
- jitl 2y agoI think having supermaven or cursor-style IDE integration is really key to making LSP worth it, otherwise friction around the workflow overwhelms the gains for many tasks. Like, I would never use AI to generate code or tests if it involved copy-pasting code in and out of webpage text boxes. But with IDE integration, I often write a function signature and doc comment with no plans to ask AI for anything, but the AI happens to offer a correct tab completion of the whole function. That's great, much less typing for me.
- lumost 2y agoI use it for the boiler plate, uber google, automated reviewer, rubber duck design reviewer, and junior engineer given extremely precise instructions. The latest models (o1, Claude sonnet new) are decent at generating code for up to 1k lines. More than that, they start to struggle. On large code bases they lose the plot quickly and generate gibberish. I only use them as code summarizers and as a Google replacement in that context.
- TacticalCoder 2y ago> I still don’t understand how people are getting value out of AI coders. I’ve tried really hard and the commits produced are just a step up from garbage. These aren't mutually exclusive. I pay for ChatGPT. It sucks fat balls at coding but it's okay to do things like "Bash: ensure exactly two params are passed, 1st one is a dir, 2nd is a file". This is slightly faster than writing it myself so it's worth $20 a month but that's about it. "from now on no explanation, code only" also helps. Does is still stuck? Definitely. But it's one more tool. I can understand why one wouldn't even bother though.
- Lerc 2y ago>"from now on no explanation, code only" also helps. Does it? Without using a model with a internal monologue interface, the explanation is the only way for the model to do any long form thinking. I would have thought that requesting an explanation of the code it is about to write would be better. An explanation after it has written the code would be counterproductive, because it would be flavouring the explanation to what it actually wrote instead of what it was wanting to achieve.
- godelski 2y agoAt that point, why not learn bash? I often hear how great GPT is at bash but imo it is terrible. Granted, most bash code is terrible, though I'm not sure why. It's like people don't even know how to use functions. It's petty quick to get to an okay level too! (I suspect few people sit down to learn it and instead learn it a line at a time over a very sparse timeframe) The other part is compounding returns. This is extra obvious with bash. Getting good at shell scripting also helps you be really good at using the shell and vise versa. The returns aren't always obvious though, but you'll quickly find yourself piping into sed or xargs or writing little for loops or feeling like find actually makes sense. Pretty soon you'll be living in the terminal and questioning why others don't. Bash scripting is an insanely underrated skill. In general this is something I find problematic with AI code generation. The struggle is part of the learning process. Training wheels are great and at face value AI should help you learn. But it's like having a solution manual to your math homework. We both know 9/10 people go for the answer and not use it to get unstuck. It's also a bit hard with LLMs because they aren't great doing one line at a time and not spoiling the next steps. But I'm sure you can prompt engineer this to a decent degree of success.
- weitendorf 2y agoWielding GenAI effectively is genuinely a skill. I’m working on an AI developer tool product (a more agentic Cursor, but not fully agentic because the underlying LLM tech/hardware is simply not there yet) and have seen a lot of different techniques, good and bad, used by developers including myself. I could go on at length but relevant to what you commented: 1. Generally you want to give the LLM one well-specified task at a time. If you weren’t specific enough initially try clarifying a bit maybe, and if it makes a mistake maybe try one round of fixing it in the same conversation. Otherwise I always recommend putting followups and separate microtasks in a separate, new conversation (with some context carried over and some no-longer-relevant context pruned). Every time you call an LLM it takes the entire conversation history as a parameter and generates the most likely response, which at least last I checked was O(n^2) for leading models. Long conversations force it to sift through tons of junk, bias responses towards what the model did previously, and often confuse it regarding objectives and instructions. 2. Don’t let the model make you forget that you know how to write software, and don’t believe everything it says. Make sure you actually read and understand the code it spits out, and try to think at least a little about any errors it causes. If it got most of the way there it’s usually easier to just do the last bit of code yourself IME, and you can still Google your errors as engineers have done for decades now. 3. Treat the model like the “doer” and don’t let it do the thinking. It’s great at converting instructions and code to more code, and for knowing lots of stuff about most things on the Internet, and to use as a sounding board. Anything more intellectually challenging than that you probably want to scope down into simpler stuff. TLDR is you need to build a Theory of Mind for how to interact with LLM coding tools and know when to take over.
- Art9681 2y agoThe best developers of the future arent the ones who mastered Python or Rust. It will be people who can describe complex things using a new permutation of English. That new English is evolving right now. So using Ai isnt just about learning some stack or tinkering with a model. Its learning how to communicate with the AI under the current constraints. Perhaps this goes against the ultimate goal. Speak plainly and AI understands. That's AGI. Today, we need to speak AI English, and that's a new knowledge domain all its own. Hopefully it will be shortlived.