8 ms·
My experience creating software with LLM coding agents – Part 2 (Tips)
- efitz 1y agoI spent much of the last several months using LLM agents to create software. I've written two blog posts about my experience; this is the second post that includes all the things I've learned along the way to get better results, or at least waste less money.
- afeezaziz 1y agoyou should write more about your experience using LLM. Is this solely using LLM?
- xwowsersx 1y agoThis lines up with my own experience of learning how to succeed with LLMs. What really makes them work isn't so different from what leads to success in any setting: being careful up front, measuring twice and cutting once.
- CuriouslyC 1y agoIf I paid for my API usage directly instead of the plan it'd be like a second mortgage.
- 3abiton 1y agoTo be fair, allocating some token for planning (recursively) helps a lot. It requires more hands on work, but produce much better results. Clarifying the tasks and breaking them down is very helpful too. Just you end up spending lots of time on it. On the bright side, Qwen3 30B is quite decent, and best of all "free".
- manmal 1y agoOne weird trick is to tell the LLM to ask you questions about anything that’s unclear at this point. I tell it eg to ask up to 10 questions. Often I do multiple rounds of these Q&A and I‘m always surprised at the quality of the questions (w/ Opus). Getting better results that way, just because it reduces the degrees of freedom in which the agent can go off in a totally wrong direction.
- deadbabe 1y agoThis is a little anthropomorphic. The faster option is to tell it to give you the full content of an ideal context for what you’re doing and adjust or expand as necessary. Less back and forth.
- manmal 1y agoCan you give me the full content of the ideal context of what you mean here?
- rzzzt 1y agoCertainly!
- 7thpower 1y agoIt’s not though, one of the key gaps right now is that people do not provide enough direction on the tradeoffs they want to make. Generally LLMs will not ask you about them, they will just go off and build. But if you have them ask, they will often come back with important questions about things you did not specify.
- deadbabe 1y agoThey don’t know what to ask. They only assemble questions according to training data.
- fuzzzerd 1y agoWhile true, the questions are all points where the LLM would have "assumed" an answer and by asking you get to point in the right direction instead.
- 7thpower 1y agoIt seems like you are trying to steer toward a different point or topic. In the course of my work, I have found they ask valuable clarifying questions. I don’t care how they do it.
- rvz 1y ago[flagged]
- indigodaddy 1y agoWhy?
- compootr 1y agoI guess you need an active developer license to write blog posts
- rvz 1y agoOr maybe this industry still trusts experienced software engineers to write well maintained and robust software used by millions that make money.
- rvz 1y agoIt's quite simple. I perfer building and using software that is robust, heavily tested and thoroughly reviewed by highly experienced software engineers who understand the code, can detect bugs and can explain what each line of code they write does. Today, we are now in the phase where embracing mediocre LLM generated code over heavily tested / scrutinized code is now encoraged in this industry - because of the hype of 'vibe coding'. If you can't even begin to explain the code or point out any bugs generated by LLMs or even off-load architectural decisions to them, you're going to have a big problem in explaining that in code review situations or even in a professional pair-programming scenario.
- exe34 1y ago> I perfer building and using software that is robust, heavily tested and thoroughly reviewed by highly experienced software engineers who understand the code, can detect bugs and can explain what each line of code they write does. that's amazing. by that logic you probably use like one or two pieces of software max. no windows, macos or gnome for you.
- pmxi 1y ago> If you are a heavy user, you should use pay-as-you go pricing if you’re a heavy user you should pay for a monthly subscription for Claude Code which is significantly cheaper than API costs.
- ramesh31 1y agoAm I alone in spending $1k+/month on tokens? It feels like the most useful dollars i've ever spent in my life. The software I've been able to build on a whim over the last 6 months is beyond my wildest dreams from a a year or two ago.
- zppln 1y agoCare to show what you've built?
- fainpul 1y ago> The software I've been able to build on a whim over the last 6 months is beyond my wildest dreams from a a year or two ago. If you don't mind sharing, I'm really curious - what kind of things do you build and what is your skillset?
- stillsut 1y agoNot OP but I know it can be difficult to really difficult to measure or communicate this to people who aren't familiar with the codebase or the problem being solved. Other than just dumping 10M tokens of chats into a gist and say read through everything I said back and forth with claude for a week. But, I think I've got the start of a useful summary format: it that takes every prompt and points to the corresponding code commit produced by ai + adds a line diff amount and summary of the task. Check it out below. https://github.com/sutt/agro/blob/master/docs/dev-summary-v1.md#v017 https://github.com/sutt/agro/blob/master/docs/dev-summary-v1... (In this case it's an python cli ai-coding framework that I'm using to build the package itself)
- tovej 1y agoI would personally never. Do I want to spend all my time reviewing AI code instead of writing? Not really. I also don't like having a worse mental model of the software. What kind of software are you building that you couldn't before?
- athrowaway3z 1y ago> One of the weird things I found out about agents is that they actually give up on fixing test failures and just disable tests. They’ll try once or twice and then give up. Its important to not think in terms of generalities like this. How they approach this depends on your tests framework, and even on the language you use. If disabling tests is easy and common in that language / framework, its more likely to do it. For testing a cli, i currently use run_tests.sh and never once has it tried to disable a test. Though that can be its own problem when it hits 1 it can't debug. # run_tests.sh # Handle multiple script arguments or default to all .sh files scripts=("${@/#/./examples/}") [ $# -eq 0 ] && scripts=(./examples/*.sh) for script in "${scripts[@]}"; do [ -n "$LOUD" ] && echo $script output=$(bash -x "$script" 2>&1) || { echo "" echo "Error in $script:" echo "$output" exit 1 } done echo " OK" ---- Another tip. For a specific tasks don't bother with "please read file x.md", Claude Code (and others) accept the @file syntax which puts that into context right away.
- blarg-and-co 1y ago[dead]
- Lucasoato 1y agoI’ve seen going very successfully using both codex with gpt5 and claude code with opus. You develop a solution with one, then validate it with the other. I’ve fixed many bugs by passing the context between them saying something like: “my other colleague suggested that…”. Bonus thing: I’ve started using symlinks on CLAUDE.md files pointing at AGENTS.md, now I don’t even have to maintain two different context files.
- dizhn 1y agoOther symlinks one can do: QWEN.md, GEMINI.md, CONVENTIONS.md (for aider).
- alex-moon 1y agoAs a human dev, can I humbly ask you to separate out your LLM "readme" from your human README.md? If I see a README.md in a directory I assume that means the directory is a separate module that can be split out into a separate repo or indeed storage elsewhere. If you're putting copy in your codebase that's instructions for a bot, that isn't a README.md. By all means come up with a new convention e.g. BOTS.md for this. As a human dev I know I can safely ignore such a file unless I am working with a bot.
- kergonath 1y agoI think things are moving towards using AGENTS.md files: https://agents.md/ https://agents.md/ . I’d like something like this to become the consensus for most commonly used tools at some point. There was a discussion here 3 days ago: https://news.ycombinator.com/item?id=44957443 https://news.ycombinator.com/item?id=44957443 .
- mattmanser 1y agoWhile I agree keep Readme a for humans, Readme literally means read me. Not 'this is a separate project'. Not 'project documentation file'. You can have read mes dotted all over a project if that's necessary. It's simply a file that a previous developer is asking you to read before you start making around in that directory.
- Terr_ 1y ago> If I see a README.md in a directory I assume that means the directory is a separate module that can be split out into a separate repo or indeed storage elsewhere. While I can understand why someone might develop that first-impression, it's never been safe to assume, especially as one starts working with larger projects or at larger organizations. It's not that unusual for essential sections of the same big project to have their own make-files, specialized utility scripts, tweaks to auto-formatter, etc. In other cases things are together in a repo for reasons of coordination: Consider frontend/backend code which runs with different languages on different computers, with separate READMEs etc. They may share very little in terms of their build instructions, but you want corresponding changes on each end of their API to remain in lockstep. Another example: One of my employer's projects has special GNU gettext files for translation and internationalization. These exist in a subdirectory with its own documentation and support scripts, but it absolutely needs to stay within the application that is using it for string-conversions.
- snissn 1y agohttps://gist.github.com/snissn/4f06cae8fb4f4ac43ffdb104db1923b9 https://gist.github.com/snissn/4f06cae8fb4f4ac43ffdb104db192...
- sothatsit 1y ago> If you are a heavy user, you should use pay-as-you go pricing; TANSTAAFL. This is very very wrong. Anthropic's Max plan is like 10% of the cost of paying for tokens directly if you are a heavy user. And if you still hit the rate-limits, Claude Code can roll-over into you paying for tokens through API credits. Although, I have never hit the rate limits since I upgraded to the $200/month plan.
- yifanl 1y agoThe blogpost is transparently an advertisement, which is ironic considering the author's last blogpost was https://blog.efitz.net/blog/modern-advertising-is-litter/ https://blog.efitz.net/blog/modern-advertising-is-litter/
- theshrike79 1y agoIn this case the lunch is being paid by VC money. I acknowledge that and get like $400 worth of tokens from my $20 Claude Code Pro subscription every month. I'm building tools I can use when the VC money runs out or a clear winner gets on top and the prices shoot up to realistic levels. At that point I've hopefully got enough local compute to run a local model though.
- efitz 1y agoOP here. I was not solicited by anyone nor did I solicit or accept compensation from anyone for this or any other post. It’s not an advertisement; I apologize if I come off as a Claude Code fanboy. If you read part 1 of my post (linked in my OP) you will see that I disclosed exactly how much I paid for my usage, and also the reasons that I ended up choosing Claude Code over other agents.
- efitz 1y agoThere are many people who quickly hit the limits of the $200/month plan. I hit the limits of the $20/month plan in less than a day. So I never tried the $200/month plan but I suspect you are wrong. Also, if you sign up for Anthropic’s feedback program you get a 30% reduction on API usage.
- bgwalter 1y agoHis profile says: "I'm a technology geek and do information security for a living." The blog post starts with: "I’m not a professional developer, just a hobbyist with aspirations." Is this a vibe blog promoting Misanthropic Claude Vibe? It is hard to tell, since all "AI" promotion blogs are unstructured and chaotic.
- chrisweekly 1y agoHmm, to my eye those descriptors (profile and blog post intro) aren't contradictory.
- deleted 1y ago[deleted]
- JeremyNT 1y agoSome of these sample prompts in this blog post are extremely verbose: If you are considering leveraging any of the documentation or examples, you need to validate that the documentation or example actually matches what is currently in the code. I have better luck being more concise and avoiding anthropomorphizing. Something like: "validate documentation against existing code before implementation" Should accomplish the same thing!
- Wowfunhappy 1y agoI've had both experiences. On some projects concise instructions seem to work better. Other times, the LLM seems to benefit from verbosity. This is definitely a way in which working with LLMs is frustrating. I find them helpful, but I don't know that I'm getting "better" at using them. Every time I feel like I've discovered something, it seems to be situation specific.
- brookst 1y agoI have the best luck with RFC speak. “You MUST validate that the documentation validates existing code before implementation. You MAY update documentation to correct any mismatches.” But I also use more casual style when investigating. “See what you think about the existing inheritance model, propose any improvements that will make it easier to maintain. I was thinking that creating a new base class for tree and flower to inherit from might make sense, but maybe that’s over complicating things” (Expressing uncertainty seems to help avoid the model latching on to every idea with “you’re absolutely right!”)
- JeremyNT 1y agoRFC speak is a good way to put it. Also, there's a big difference between giving general "always on" context (as in agents.md) for vibe coding - like "validate against existing code" etc - versus bouncing ideas in a chat session like your example, where you don't necessarily have a specific approach in mind and burning a few extra tokens for a one off query is no big deal. Context isn't free (either literally or in terms of processing time) and there's definitely a balance to be found for a given task.
- 1y ago
- kvnhn 1y agoIMO, a key passage that's buried: "You can ask the agent for advice on ways to improve your application, but be really careful; it loves to “improve” things, and is quick to suggest adding abstraction layers, etc. Every single idea it gives you will seem valid, and most of them will seem like things that you should really consider doing. RESIST THE URGE..." A thousand times this. LLMs love to over-engineer things. I often wonder how much of this is attributable to the training data...
- brookst 1y agoThey’re not dissimilar to human devs, who also often feel the need to replat, refactor, over-generalize, etc. The key thing in both cases, human and AI, is to be super clear about goals. Don’t say “how can this be improved”, say “what can we do to improve maintainability without major architectural changes” or “what changes would be required to scale to 100x volume” or whatever. Open-ended, poorly-defined asks are bad news in any planning/execution based project.
- exitb 1y agoThere are however human developers that have built enough general and project-specific expertise to be able to answer these open-ended, poorly-defined requests. In fact, given how often that happens, maybe that’s at the core of what we’re being paid for.
- awesome_dude 1y agoI have to be honest, I've heard of these famed "10x" developers, but when I come close to one I only ever find "hacks" with a brittle understanding of a single architecture.
- brookst 1y agoBut if the business doesn’t know the goals, is it really adding any value to go fulfill poorly defined requests like “make it better”? AI tools can also take a swing at that kind of thing. But without a product/business intent it’s just shooting in the dark, whether human or AI.
- theptip 1y ago> Finally it occurred to me to put context where it was needed - directly in the test files. Probably CLAUDE.md is a better place? > Too much context Claude’s Sub-agents[1] seems to be a promising way of getting around this, though I haven’t had time to play with the feature too much. Eg when you need to take a context-busting action like debugging dependencies, instead spin up a new agent to read the output and summarize. Then your top-level context doesn’t get polluted. [1]: https://docs.anthropic.com/en/docs/claude-code/sub-agents https://docs.anthropic.com/en/docs/claude-code/sub-agents
- swader999 1y agoI wrote three sub agents this week, one to run unit tests, another to run playwright and a third to write playwright. These are pretty basic boundaries that aren't hard to share context between agent and the orchestrating main agent. It seemed to help a lot. I also have complex ways to run tests (docker, data seeding, auth) and previously it was getting lost. Only compacted a couple times. Was a big improvement.
- nicwolff 1y agoI've added a subagent to read the "memory_bank" files for project context after being told the task at hand, and summarize only the pertinent parts for the main agent. This is working well to keep the context focused. https://gist.github.com/nicwolff/273d67eb1362a2b1af42e822f6c8e6c6 https://gist.github.com/nicwolff/273d67eb1362a2b1af42e822f6c...
- andai 1y agoPart 1: https://efitz-thoughts.blogspot.com/2025/08/my-experience-creating-software-with.html https://efitz-thoughts.blogspot.com/2025/08/my-experience-cr...
- jona777than 1y ago> If the original function isn’t working as expected, I suspect that the agent-created test will test the functionality as it exists, not as was intended, regardless of what it calls the test case. I have experienced this on many occasions. It ultimately adds up to a sneakily false sense of code stability.