13 ms·
Over-editing refers to a model modifying code beyond what is necessary
- jeremie_strand 6mo ago[dead]
- itopaloglu83 6mo agoI always described it as over-complicating the code, but doing too much is a better diagnosis.
- whinvik 6mo agoYeah I have always felt GPT 5.4 does too much. It is amazing at following instructions precisely but it convinces itself to do a bit too much. I am surprised Gemini 3.1 Pro is so high up there. I have never managed to make it work reliably so maybe there's some metric not being covered here.
- eterm 6mo agoIt's funny, because the wisdom that was often taught ( but essentially never practiced ) was "Refactor as you go". The idea being that if you're working in an area, you should refactor and tidy it up and clean up "tech debt" while there. In practice, it was seldom done, and here we have LLMs actually doing it, and we're realising the drawbacks.
- hyperpape 6mo agoThat's a real question, maybe the changes are useful, though I think I'd like to see some examples. I do not trust cognitive complexity metrics, but it is a little interesting that the changes seem to reliably increase cognitive complexity.
- ramesh31 6mo ago>The idea being that if you're working in an area, you should refactor and tidy it up and clean up "tech debt" while there. This is horrible practice, and very typical junior behavior that needs to be corrected against. Unless you wrote it, Chesterton's Fence applies; you need to think deeply for a long time about why that code exists as it does, and that's not part of your current task. Nothing worse than dealing with a 1000 line PR opened for a small UI fix because the code needed to be "cleaned up".
- cassianoleal 6mo agoThat is the flip side of what you're arguing against, and is also very typical junior behaviour that needs to be corrected against. Tech debt needs to be dealt with when it makes sense. Many times it will be right there and then as you're approaching the code to do something else. Other times it should be tackled later with more thought. The latter case is frequently a symptom of the absence of the former. In Extreme Programming, that's called the Boy Scouting Rule. https://furqanramzan.github.io/clean-code-guidelines/principle/bsr.html https://furqanramzan.github.io/clean-code-guidelines/princip...
- traderj0e 6mo agoThe Boy Scout "leave it better than you found it" is a good rule to follow. All code has its breaking points, so when you're adding a new feature and find that the existing code doesn't support it without hacks, it probably needs a refactor.
- ramesh31 6mo agoIndeed there's a distinction that needs to be made here between "not refactoring this code means I'll need to add hacks" and "Oh I'll just clean that up while I'm in here." The former can be necessary, but the latter is something you learn with experience to avoid.
- cassianoleal 6mo ago> the latter is something you learn with experience to avoid. The latter is something you learn to judge the right time to tackle. Sometimes a small improvement that's not required will mean you're not pressed to make the refactor to avoid hacks. The earliest you can tackle problems, the cheaper they are to solve.
- yamajun93 6mo agoThis is a real pain, I notice over-editing a lot when using AI coding tools locally and also re-inventing the wheel (e.g not using external libraries or reimplementing something over and over again). The Boy Scouting Rule is a good framing, but the tricky part is deciding when we apply it and what's actually a 'minimal' change. Maybe smt like 10% of the PR should be refactor max, and if is more open a new PR?
- localhoster 6mo agoSo I think theres some more nuance than that. A lot of the times, the abstraction is solid enough for you to work with that code area, ie tracking down some bug or extending a functionality. But sometimes you find yourself at a crossroad - which is either hacking around the existing implementation, or rethink it. With LLMs, how do you even rethink it? Does it even matter to rethink it? And on any who, those decisions are hidden away from you.
- traderj0e 6mo agoIt's only hidden if you don't read the code. Even if you don't, at some point you'll notice the LLM starting to struggle.
- bluefirebrand 6mo agoThere is a pretty substantial difference between "making changes" and "refactoring" If LLMs are doing sensible and necessary refactors as they go then great I have basically zero confidence that is actually the case though
- hirako2000 6mo agoWhen the model write new code doing the same thing as existing logic that's not a refactor. At times even when a function is right there doing exactly what's needed. Worse, when it modifies a function that exists, supposedly maintaining its behavior, but breaks for other use cases. Good try I guess. Worst. Changing state across classes not realising the side effect. Deadlock, or plain bugs.
- aerhardt 6mo agoWhen they decide to touch something as they go, they often don't improve it. Not what I would call "refactoring" but rather a yank of the slot machine's arm.
- raincole 6mo agoReally? I've never heard it's considered wise to put refactoring and new features (or bugfixes) in the same commit. Everyone I know from every place I've seen consider it bad. From harmful to a straight rejection in code review. "Refactor-as-you-go" means to refactor right after you add features / fix bugs, not like what the agent does in this article.
- xboxnolifes 6mo agoNotice how they didn't say to put it in the same commit. The real issue, and why refactor as you go isn't done as much, is the overhead of splitting changes that touch the same code into different commits without disrupting your workflow. It's not as easy as it should be to support this strategy. Instead you to do it later, and then never do it.
- raincole 6mo agoI think you're talking about a different topic unrelated to the linked article. In the linked article the LLM doesn't split it into several commits. If LLM had a button to split the bug fix and the overall refactoring, the author wouldn't complain and we wouldn't see this article.
- whimblepop 6mo ago> In practice, it was seldom done, and here we have LLMs actually doing it, and we're realising the drawbacks. I spent some time dealing with this today. The real issue for me, though, was that the refactors the agent did were bad. I only wanted it to stop making those changes so I could give it more explicit changes on what to fix and how.
- jstanley 6mo agoConversely, I often find coding agents privileging the existing code when they could do a much better job if they changed it to suit the new requirement. I guess it comes down to how ossified you want your existing code to be. If it's a big production application that's been running for decades then you probably want the minimum possible change. If you're just experimenting with stuff and the project didn't exist at all 3 days ago then you want the agent to make it better rather than leave it alone. Probably they just need to learn to calibrate themselves better to the project context.
- _pastel 6mo agoThe tradeoff is highly contextual; it's not a tradeoff an agent can always make by inspecting the project themselves. Even within the same project, for a given PR, there are some parts of the codebase I want to modify freely and some that I want fixed to reduce the diff and testing scope. I try to explain up-front to the agent how aggressively they can modify the existing code and which parts, but I've had mixed success; usually they bias towards a minimal diff even if that means duplication or abusing some abstractions. If anyone has had better success, I'd love to hear your approach.
- mncharity 6mo agoJust brainstorming, but perhaps a more tangible gradient, with social backpressure? Imagine three identical patch tools: "patch", "submit patch", and "send patch to chief architect and wait". With the "where each can be used" explained or even enforced. Having the contrast of less-aggressive options, might make it easier to encourage more aggressive refactoring elsewhere. Or pushing the impact further up the CoT, "patch'ing X requires an analysis field describing less invasive alternatives and their un/suitability; for Y, just do it, refactor aggressively".
- ZihangZ 6mo ago[dead]
- jauntywundrkind 6mo agoTo get the agent to think for itself sometimes it feels like I have to delete a bunch of code and markdown first. Instruction to refactor/reconsider broadly has such mild success, I find. I'll literally run an agent & tell it to clean up a markdown file that has too much design in it, delete the technical material, and/or delete key implementations/interfaces in the source, then tell a new session to do the work, come up with the design. (Then undelete and reconcile with less naive sessions.) Path dependence is so strong. Right now I do this flow manually but I would very much like to codify this, make a skill for this pattern that serves so well.
- anonu 6mo agoHere, the author means the agent over-edits code. But agents also do "too much": as in they touch multiple files, run tests, do deployments, run smoke tests, etc... And all of this gets abstracted away. On one hand, its incredible. But on the other hand I have deep anxiety over this: 1. I have no real understanding of what is actually happening under the hood. The ease of just accepting a prompt to run some script the agent has assembled is too enticing. But, I've already wiped a DB or two just because the agent thought it was the right thing to do. I've also caught it sending my AWS credentials to deployment targets when it should never do that. 2. I've learned nothing. So the cognitive load of doing it myself, even assembling a simple docker command, is just too high. Thus, I repeatedly fallback to the "crutch" of using AI.
- Barbing 6mo agoMust consider ourselves lucky for having the intuition to notice skill stagnation and atrophy. Only helps if we listen to it :) which is fun b/c it means staying sharp which is inherently rewarding
- ok_dad 6mo agoWhy are you letting the LLM drive? Don't turn on auto-approve, approve every command the agent runs. Don't let it make design or architecture decisions, you choose how it is built and you TELL that clanker what's what! No joke, if you treat the AI like a tool then you'll get more mileage out of it. You won't get 10x gains, but you will still understand the code.
- giraffe_lady 6mo agoI agree with this too. I decided on constraints for myself around these tools and I give my complete focus & attention to every prompt, often stopping for minutes to figure things through and make decisions myself. Reviewing every line they produce. I'm a senior dev with a lot of experience with pair programming and code review, and I treat its output just as I would those tasks. It has about doubled my development pace. An absolutely incredible gain in a vacuum, though tiny compared to what people seem to manage without these self-constraints. But in exchange, my understanding of the code is as comprehensive as if I had paired on it, or merged a direct report's branch into a project I was responsible for. A reasonable enough tradeoff, for me.
- slopinthebag 6mo agoI think the industry has leaned waaay too far into completely autonomous agents. Of course there are reasons why corporations would want to completely replace their engineers with fully autonomous coding agents, but for those of us who actually work developing software, why would we want less and less autonomy? Especially since it alienates us from our codebases, requiring more effort in the future to gain an understanding of what is happening. I think we should move to semi-autonomous steerable agents, with manual and powerful context management. Our tools should graduate from simple chat threads to something more akin to the way we approach our work naturally. And a big benefit of this is that we won't need expensive locked down SOTA models to do this, the open models are more than powerful enough for pennies on the dollar.
- NitpickLawyer 6mo agoI'm hearing this more and more, we need new UX that is better suited for the LLM meta. But none that I've seen so far have really got it, yet.
- grttww 6mo agoWhen you steer a car, there isn’t this degree of probability about the output. How do you emulate that with llm’s? I suppose the objective is to get variance down to the point it’s barely noticeable. But not sure it’ll get to that place based on accumulating more data and re-training models.
- slopinthebag 6mo agoWell, the point is by steering it you can get both more expected/reproducible output, and you can correct bad assumptions before they become solidified in your codebase. You can get pretty close to reproducible output by narrowing the scope and using certain prompts/harnesses. As in, you get roughly the same output each time with identical prompts, assuming you're using a model which doesn't change every few hours to deal with load, and you aren't using a coding harness that changes how it works every update. It's not deterministic, but if you ask it for a scoped implementation you essentially get the same implementation every time, with some minor and usually irrelevant differences. So you can imagine with a stable model and harness, with steering you can basically get what you ask it for each time. Tooling that exploits this fact can be much more akin to using an autocomplete, but instead of a line of code it's blocks of code, functions, etc. A harness that makes it easy to steer means you can basically write the same code you would have otherwise written, just faster. Which I think is a genuine win, not only from a productivity standpoint but also you maintain control over the codebase and you aren't alienated or disenfranchised from the output, and it's much easier to make corrections or write your own implementations where you feel it's necessary. It becomes more of an augmentation and less of a replacement.
- lo1tuma 6mo agoI’m not sure if I share the authors opinion. When I was hand-writing code I also followed the boy-scout rule and did smaller refactorings along the line.
- exitb 6mo agoAs mentioned in the article, prompting for minimal changes does help. I find GPT models to be very steerable, but it doesn’t mean much when you take your hands of the wheel. These type of issues should be solved at planning stage.
- deleted 6mo ago[deleted]
- Almured 6mo agoI feel ambivalent about it. In most cases, I fully agree with the overdoing assessment and then having to spend 30min correcting and fixing. But I also agree with the fact sometimes the system is missing out on more comprehensive changes (context limitations I suppose)! I am starting to be very strict when coding with these tool but still not quite getting the level of control I would like to see
- lopsotronic 6mo agoWhen asked to show their development-test path in the form of a design document or test document, I've also noticed variance between the document generated and what the chain-of-thought thingy shows during the process. The version it puts down into documents is not the thing it was actually doing. It's a little anxiety-inducing. I go back to review the code with big microscopes. "Reproducibility" is still pretty important for those trapped in the basements of aerospace and defense companies. No one wants the Lying Machine to jump into the cockpit quite yet. Soon, though. We have managed to convince the Overlords that some teensy non-agentic local models - sourced in good old America and running local - aren't going to All Your Base their Internets. So, baby steps.
- aerhardt 6mo agoI'm building a website in Astro and today I've been scaffolding localization. I asked Codex 5.4 x-high to follow the official guidelines for localization and from that perspective the implementation was good. But then it decides to re-write the copy and layout of all pages. They were placeholders, but still? Codex also has a tendency to apply unwanted styles everywhere. I see similar tendencies in backend and data work, but I somehow find it easier to control there. I'm pretty much all in on AI coding, but I still don't know how to give these things large units of work, and I still feel like I have to read everything but throwaway code.
- jasonjmcghee 6mo agoI never use xhigh due to overthinking. I find high nearly always works better. Purely anecdotal.
- sabas123 6mo agoThis is even the official OpenAI guideline too.
- magicalhippo 6mo agoYou can steer it though. When I see it going off the reservation I steer it back. I also commit often, just about after every prompt cycle, so I can easily revert and pick up the ball in a fresh context. But yeah, I saw a suggestion about adding a long-lived agent that would keep track of salient points (so kinda memory) but also monitor current progress by main agent in relation to the "memory" and give the main agent commands when it detects that the current code clashes with previous instructions or commands. Would be interesting to see if it would help.
- agdexai 6mo ago[dead]
- traderj0e 6mo agoThey also don't understand how exceptions work. They'll try-catch everything, print the error, and continue. If I see a big diff, I know it just added 10 try-catches in random parts of my codebase.
- pilgrim0 6mo agoLike others mentioned, letting the agent touch the code makes learning difficult and induces anxiety. By introducing doubt it actually increases the burden of revision, negating the fast apparent progress. The way I found around this is to use LLMs for designing and auditing, not programming per se. Even more so because it’s terrible at keeping the coding style. Call it skill issue, but I’m happier treating it as a lousy assistant rather than as a dependable peer.
- pyrolistical 6mo agoI attempt to solve most agent problems by treating them as a dumb human. In this case I would ask for smaller changes and justify every change. Have it look back upon these changes and have it ask itself are they truly justified or can it be simplified.
- graybeardhacker 6mo agoI use Claude Code every day and have for as long as it has been available. I use git add -p to ensure I'm only adding what is needed. I review all code changes and make sure I understand every change. I prompt Claude to never change only whitespace. I ask it to be sure to make the minimal changes to fix a bug. Too many people are treating the tools as a complete replacement for a developer. When you are typing a text to someone and Google changes a word you misspelled to a completely different word and changes the whole meaning of the text message do you shrug and send it anyway? If so, maybe LLMs aren't for you.
- sebringj 6mo ago[dead]
- deleted 6mo ago[deleted]
- dbvn 6mo agoDon't forget the non-stop unnecessary comments
- tim-projects 6mo ago> The model fixes the bug but half the function has been rewritten. The solution to this is to use quality gates that loop back and check the work. I'm currently building a tool with gates and a diff regression check. I haven't seen these problems for a while now. https://github.com/tim-projects/hammer https://github.com/tim-projects/hammer
- Isolated_Routes 6mo agoI think building something really well with AI takes a lot of work. You can certainly ask it to do things and it will comply, and produce something pretty good. But you don't know what you don't know, especially when it speaks to you authoritatively. So checking its work from many different angles and making sure it's precise can be a challenge. Will be interesting to see how all of this iterates over time.
- deepfriedbits 6mo agoI agree 100%. At the same time, I feel like this piece, and our comments on it are snapshots in time because of the rate of advancement in the industry. These coding models are already significantly better than they were even nine months ago. I can't help but read complaints about the capabilities of AI – and I'm certainly not accusing you of complaining about AI, just a general thought – and think "Yet" to myself every time.
- Isolated_Routes 6mo agoExactly! I completely agree. I think figuring out how to use this new tool well develop into a bit of an art form, which we will race to keep up with.
- ValentineC 6mo ago> But you don't know what you don't know, especially when it speaks to you authoritatively. So checking its work from many different angles and making sure it's precise can be a challenge. I've spent far more time pitting one AI context against another (reviewing each other's work) than I have using AI to build stuff these days. The benefit is that since it mostly happens asynchronously, I'm free to do other stuff.
- mleo 6mo agoIf I don’t know what I don’t know, how am I going to build something any better than a coding agent? An approach on a couple of projects has been to prototype with the agent, learn, write a design and then start over. I then know the areas to look into more detail.
- simonw 6mo agoI've not seen over-editing in Claude Code or Codex in quite a while, so I was interested to see the prompts being used for this study. I think they're in here, last edited 8 months ago: https://github.com/nreHieW/fyp/blob/5a4023e4d1f287ac73a616b5b944a14f28422c7e/partial_edits/utils/prompts_utils.py https://github.com/nreHieW/fyp/blob/5a4023e4d1f287ac73a616b5...
- qweiopqweiop 6mo agoLikewise, this felt like an early agent problem to me.
- 59nadir 6mo agoJust had one today where GPT-5.4, instead of adding the 10 lines I asked for (an addition that could be done pretty mechanically by just looking at some previous code and adding a similar thing with different/new variable names) proceeded to rewrite 50 lines instead, because it was "cleaner". It was not. It also didn't originally add the thing I asked for either, which was perplexing. Over-editing is definitely not some long gone problem. This was on xhigh thinking, because I forgot to set it to lower.
- jollyllama 6mo agoIt's called code churn. Generally, LLMs make code churn.
- ricardorivaldo 6mo agoduplicated ? https://news.ycombinator.com/item?id=47866913 https://news.ycombinator.com/item?id=47866913
- tantalor 6mo ago> Code review is already a bottleneck Counterpoint: no it isn't > makes this job dramatically harder No it doesn't
- esafak 6mo agoHow many LOC do you generate and read a day? Only your own code or others' too?
- LetsGetTechnicl 6mo agoWell seeing as they don't KNOW anything this isn't surprising at all
- vibe42 6mo agoWith the pi-mono coding agent (running local, open models) this works very well: "Do not modify any code; only describe potential changes." I often add it to the end when prompting to e.g. review code for potential optimizations or refactor changes.
- foo12bar 6mo agoI've noticed AI's often try and hide failure by catching exceptions and returning some dummy value maybe with some log message buried in tons of extraneous other log messages. And the logs themselves are often over abbreviated and missing key data to successfully debug what is happening. I suspect AI's learned to do this in order to game the system. Bailing out with an exception is an obvious failure and will be penalized, but hiding a potential issue can sometimes be regarded as a success. I wonder how this extrapolates to general Q&A. Do models find ways to sound convincing enough to make the user feels satisfied and the go away? I've noticed models often use "it's not X, it's Y", which is a binary choice designed to keep the user away from thinking about other possibilities. Also they often come up with a plan of action at the end of their answer, a sales technique known as the "assumptive close", which tries to get the user to think about the result after agreeing with the AI, rather than the answer itself.
- hexaga 6mo agoAI behavior is pretty easy to understand and predict if you view it from the lens of: they will shamelessly do any/everything possible to game whatever metric they are trained on. Because... that's how hill-climbing a metric looks. It's A/B enshittification taken to inscrutable heights. They are trained on human feedback, so there is no other way this goes. Every bit of every response is pointed toward subversion of the assumed evaluator.
- samusiam 6mo agoIn my experience this "gaming" behavior is easily caught by just asking another agent (could just be another session of Claude Code) to review the code changes.
- Bengalilol 6mo agoTangent and admittedly off-topic but I've come to see LLM-assisted coding as a kind of teleportation. With LLMs, you glimpse a distant mountain. In the next instant, you're standing on its summit. Blink, and you are halfway down a ridge you never climbed. A moment later, you're flung onto another peak with no trail behind you, no sense of direction, no memory of the ascent. The landscape keeps shifting beneath your feet, but you never quite see the panorama. Before you know it, you're back near the base, disoriented, as if the journey never happened. But confident, you say you were on the top of the mountain. Manual coding feels entirely different. You spot the mountain, you study its slopes, trace a route, pack your gear. You begin the climb. Each step is earned steadily and deliberately. You feel the strain, adjust your path, learn the terrain. And when you finally reach the summit, the view unfolds with meaning. You know exactly where you are, because you've crossed every meter to get there. The satisfaction isn't just in arriving, nor in saying you were there: it is in having truly climbed.
- jdkoeck 6mo agoThe thing is, with manual coding, you spot a view in the distance, you trek your way for a few hours, and you realize when you get there that the view isn’t as great as you thought it was. With LLM-assisted coding, you skip the trek and you instantly know that’s not it.
- btbuildem 6mo agoI wish there was a reliable way to choke the agents back and prevent them from doing this. Every line of code added is a potential bug, and they overzealously spew pages and pages of code. I've routinely gone through my (hobby) projects and (yes, still with the aid of an LLM) trimmed some 80% of the generated code with barely any loss of functionality. The cynic in me thinks it's done on purpose to burn more tokens. The pragmatist however just wants full control over the harness and system prompts. I'm sure this could be done away with if we had access to all the knobs and levers.
- qurren 6mo ago> if we had access to all the knobs and levers. We do, just tell it what you want in your AGENTS.md file. Agents also often respond well to user frustration signs, like threatening to not continue your subscription.
- gobdovan 6mo ago> Agents also often respond well to user frustration signs, like threatening to not continue your subscription. From the phrasing, I can't but imagine you as a very calm, completely unemotional person that only emulates user frustration signs, strategically threatening AI that you'll close your subscription when it nukes your code.
- brianwmunz 6mo agoI feel like a core of this is that agents aren't exactly a replacement for a junior developer like some people say. A junior dev has its own biases, predispositions, history and understanding of the internal and external aspects of a product and company. An AI agent wants to do what you ask in the best way possible which is...not always what a dev wants :) The fix the article talks about is simple but shows that these models have no inherent sense of project scope or proportionality. You have to give context (as much context as possible) explicitly to fill in the gaps so it infers less and makes smaller decisions.
- spullara 6mo agothis is one of the best things about using claude over gpt. claude understands the bigger assignment and does all the work and sometimes more than necessary but for me it beats the alternative.
- deleted 6mo ago[deleted]
- standardly 6mo agoI've had a bad experience using AI for front-end stuff, where I replace or deprecate a feature only to notice later all the artifacts it left behind, some which were never even used in the first place. I re-did an entire UI recently, and when one of the elements failed to render I noticed the old UI peeking out from underneath. It had tried just covering up old elements instead of adjusting or replacing them. Like telling your son to clean their room, so they push all the clothes under the bed and hope you don't notice LOL It saves 2 hours of manual syntax wrangling but introduces 1 .5 hours of clean up and sanity checking. Still a net productivity increase, but not sure if its worth how lazy it seems to be making me (this is an easy error to correct, im sure, but meh Claude can fix it in 2 seconds so...)
- recursivecaveat 6mo agoMy experience is usually the opposite. The code they write is verbose yes, but the diffs are over-minimal. Whenever I see a comment like "Tool X doesn't support Y or has a bug with Z [insert terrible kludge]" and actually fixing the problem in the other file would be very easy, I know it is AI-generated. I suspect there is a bias towards local fixes to reduce token usage.
- m463 6mo agoYou know, this made me think of over-engineering. ...and that led me to believe that AI might be very capable to develop over-engineered audio equipment. Think of all the bells and whistles that could be added, that could be expressed in ridiculous ways with ridiculous price tags.
- janalsncm 6mo agoThis is a really solid writeup. LLMs are way too verbose in prose and code, and my suspicion is this is driven mainly by the training mechanism. Cross entropy loss steers towards garden path sentences. Using a paragraph to say something any person could say with a sentence, or even a few precise words. Long sentences are the low perplexity (low statistical “surprise”) path.
- kgeist 6mo agoInteresting, my assumption used to be that models over-edit when they're run with optimizations in attention blocks (quantization, Gated DeltaNet, sliding window etc.). I.e. they can't always reconstruct the original code precisely and may end up re-inventing some bits. Can't it be one of the reasons too?
- scotty79 6mo agoThis seems like something that should be easy to prevent in pi harness. Just tell it to make an extension that before calling file edit tool asks the model to make sure that no lines unconnected with the current topic are going to be unnecessarily changed by this edit.
- hathawsh 6mo agoI'm either in a minority or a silent majority. Claude Code surpasses all my expectations. When it makes a mistake like over-editing, I explain the mistake, it fixes it, and I ask it to record what it learned in the relevant project-specific skills. It rarely makes that mistake again. When the skill file gets big, I ask Claude to clean and compact it. It does a great job. It doesn't really make sense economically for me to write software for work anymore. I'm a teacher, architect, and infrastructure maintainer now. I hand over most development to my experienced team of Claude sessions. I review everything, but so does Claude (because Claude writes thorough tests also.) It has no problem handling a large project these days. I don't mean for this post to be an ad for Claude. (Who knows what Anthropic will do to Claude tomorrow?) I intend for this post to be a question: what am I doing that makes Claude profoundly effective? Also, I'm never running out of tokens anymore. I really only use the Opus model and I find it very efficient with tokens. Just last week I landed over 150 non-trivial commits, all with Claude's help, and used only 1/3 of the tokens allotted for the week. The most commits I could do before Claude was 25-30 per week. (Gosh, it's hard to write that without coming across as an ad for Anthropic. Sorry.)
- swader999 6mo agoI feel the same way. Doesn't make sense economically or even in good faith for me to use company paid time writing code for line of business apps at anymore and I'm 28 years into this kind of work.
- Powdering7082 6mo agoIs your claude.md, skills or other settings that you have honed public?
- hathawsh 6mo agoSorry, no, and they're highly project specific anyway. I just started with the "/init" skill a few weeks ago and gradually improved it from there.
- p1necone 6mo agoWhich subscription tier are you using?
- collimarco 6mo agoOver-editing and over-adding... I can find solutions that are just a few lines of code in a single file where AI would change 10 files and add 100s of lines of code. Writing less code is more important than ever. Too much code means more technical debt, a maintainability nightmare and more places where bugs can hide.
- Gigachad 6mo agoI've seen this at work where people submit PRs that implement whole internal libraries to do something that could have been done with an existing tool or just done simpler in general. It's impossible to properly review this in a reasonable time and they always introduce tons of subtle bugs.
- maxbeech 6mo ago[dead]
- ArielTM 6mo ago[dead]
- aroido-bigcat 6mo ago[dead]
- rcvassallo83 6mo agoThis resonates I've had success with greenfield code followed by frustration when asking for changes to that code due to over editing And prompting for "minimal changes" does keep the edits down. In addition to this instruction, adding specifics about how to make the change and what not to do tends to get results I'm looking for. "add one function that does X, add one property to the data structure, otherwise leave it as is, don't add any new validation"
- Meterman 6mo agoI really felt this. total pain point for me.
- EverMemory 6mo ago[dead]
- ozozozd 6mo agoThere is no need for a new name. It’s called a high-impact change. As opposed to a low-impact change, where one changes or adds the least number of lines necessary to achieve the goal. Not surprised to see this, since once again, because some of us didn’t like history as a subject, lines of code is a performance measure, like a pissing contest.
- jacek-123 6mo agoFeels like a training-data artifact. SFT and preference data are full of "here's a cleaner version of your file", not "here's the minimum 3-line diff". The model learned bigger, more polished outputs win. Prompting around it helps a bit but you're fighting the prior.
- panavm 6mo ago[flagged]
- bluequbit 6mo agoI call this overcooking. Adding unnecesary features.
- rosscorinne96 6mo ago[dead]
- EthanFrostHI 6mo ago[dead]
- DrokAI 6mo ago[dead]
- devdevai 6mo ago[dead]
- BoredomIsFun 6mo agoIt feels like a pointless conversation, if no sampler settings (min_p, temperature etc.) mentioned.
- jimmypk 6mo ago[flagged]
- figassis 6mo agoOver editing is one of the biggest tells of junior engineers. Often, a very big task is reduced to a 1 line change if you spend the time to understand the core problem. I feel like this was one of the most valuable skills an engineer could learn, as it protects the integrity of the system by making minimum viable changes. If you need to refactor something, it should be clear that that is the task. But we're all tokenmaxing now.
- ryanshrott 6mo ago[dead]
- lw1981 6mo agoI’ve run into this in debugging too. I had one case where the model fixed a bug in concurrent execution code by basically making it sequential. So yes, the bug disappeared, but only because it removed the concurrency property instead of actually solving the underlying issue. That kind of thing has made me much more cautious about judging these tools purely by whether the immediate error went away.
- Abby_101 6mo agoThe over-editing cost is asymmetric when you're solo. If an agent rewrites 50 lines when you asked it to touch 5, there's no second reviewer behind you. You're the writer and the reviewer, and reviewer (you) is usually the one whose attention is already depleted. What helped me most wasn't prompting tricks but committing after every turn so I can revert cheaply and re-prompt with tighter constraints. Cheaper than auditing a 400-line diff at 11pm.
- SummSolutions 6mo agoGreat article covering an issue that I can see growing as more general users try their hand at coding. Prompting, reviewing, and testing are crucial parts of my updating/ editing process. Any suggestions?