5 ms·
In my experience, the delta in agent performance is substantial if the codebase is littered with dead code, redundant code, unreachable fallbacks, leaking abstr
by i_have_an_idea 3mo ago
In my experience, the delta in agent performance is substantial if the codebase is littered with dead code, redundant code, unreachable fallbacks, leaking abstractions and half-baked design patterns vs if the code is well-organized, with clear data flow, with good encapsulation and clean architecture. Like, I've seen all the frontier models have to do several rounds of code review / QA and fix when the code is bad vs just getting it right at the 1st/2nd attempt.
- ramraj07 3mo agoI was reading your comment, agreeing with it but still feeling why this is a bad comment. It just occurred to me that an anecdotal statement like this is the antithesis of scientific discourse. We have a paper here, trying to answer a question, and anecdotal testimonials can only harm the discussion by biasing readers without adding anything of value to let anyone objectively conclude anything on the problem. The most useful discussion would be if we all read the paper and critique its methodology or results.
- vetronauta 3mo agoI was reading your comment, disagreeing with it but still feeling why this is a good comment. It just occurred to me that this is not science: science must be reproducible and this is just an historical report on artifact that will be unavailable soon.
- dnautics 3mo agoi mean this is feeling too but im too paranoid and frequently do refactoring and code organization passes and never don't do it, so i cant say i know for sure there is a delta. though people who complain that llms aren't that great strike me as the type to have messy code bases
- BobbyTables2 3mo agoFeel the same way myself when working in messy codebases… At some point, the horrible patterns start to rub off…
- NitpickLawyer 3mo agoEvery time this subject comes up, there are a bunch of takes along the lines of "would you work on a codebase maintained by agents? they'll mess up the code". And I'm asking myself where these people work, because in 20+ years I've yet to see that pristine state of a project that keeps being pristine after the honeymoon greenfield phase, and 50+ people start working on it. Every project devolves in time, old stuff gets patched in a hurry, someone tries to make it better, learns why certain things were done a certain way, hits some undocumented client needs handled by some arcane combination of code + external systems, and so on. If anything, keeping track of what does what in a project is a task where agents can shine, if only in "ask" mode so you can figure out things quicker. Not to mention onboarding and stuff for new team members.
- Muromec 3mo agoEntropy is real and with offshore contractors you don't need agents to blame for it
- hannofcart 3mo agoSome of the issues mentioned above like dead code removal, code duplication, unreachable code are already solved using deterministic linters for quite a while now for most language ecosystems. You can get the LLM to run a script which checks for all of these and also enforce them by running the same script as a pre-commit hook. Setting this up religiously in every code base I work on has been what's given me the most mileage with agentic coding. I wrote down a more detailed post of the various linters I use here: https://www.balajeerc.info/Use-Deterministic-Guardrails-for-your-LLM-Agents/ https://www.balajeerc.info/Use-Deterministic-Guardrails-for-...
- rafaelmn 3mo ago> Some of the issues mentioned above like dead code removal, code duplication, unreachable code are already solved using deterministic linters for quite a while now for most language ecosystems. I have legacy endpoints that are no longer used in practice, there for historical reasons, intertwined with existing code etc. They might be marked obsolete, services implementing it are not - agent greps those, builds off of them - produces half legacy garbage. Linters only handle trivial cases most of us already solved.
- yoyohello13 3mo agoYeah, we have a big struggle with this. We have lots of legacy code that doesn't follow our latest design patterns intermixed with new code. The LLM picks up bad habits depending on what it pulls in to context first. We have AGENTS.md configured with the right way, but old style still slips in. We obviously need to update the old code but on the other hand if it ain't broke why touch it.
- kstenerud 3mo agoI have the agent inject comments that mention that this particular code is legacy and must not be used as a reference, should not be cleaned up, etc. If you have a document that lists all of the reasons not to use or touch some code, the comments can simply be references to it. // LEGACY CODE, per docs/legacy_rules.md §14, §19
- yoyohello13 3mo agoI’ve been working with these things for quite some time now and every time I simply “treat it like I would a human” it seems to perform better. I can’t imagine agents wouldn’t perform better in a clean codebase than a giant mess of one. Just like it performs better when it has well formed specs and access to documentation.
- jaggederest 3mo agoIt actually goes even further than humans, humans can pretty rapidly get inured to things being awkward or messy and stop noticing, but the context for agents is taking up the same space and "attention" every time they're run, and they're creations entirely of context, so the quality and examples matter massively.
- maccard 3mo ago> actually goes even further than humans, humans can pretty rapidly get inured to things being awkward or messy and stop noticing, You’ve never had an agent completely lose the plot and forget/confuse its instructions due to the context filling up?
- Muromec 3mo agoThat's the moment it should take a nap to compress the context. If you are expose the context size to the agent through some diagnostic input it will even do it itself
- embedding-shape 3mo ago> I can’t imagine agents wouldn’t perform better in a clean codebase than a giant mess of one. I guess it depends on what you mean with "better" but almost all the agent-built projects I do with zero regards to code quality, design and architecture ends up with every single agent needing 10+ minutes to do even the easy changes, while the ones where I focused on those things together with the agent, large changes can take 10+ minutes but everything else is solved faster. I don't have empirical evidence of this yet, I guess I should put together some sort of test to confirm/disconfirm this.