19 ms·
Getting AI to work in complex codebases
- dhorthy 1y ago[dead]
- jb2403 1y agoIt’s refreshing to read a full article this was written by a human. Content +++
- malfist 1y agoThis article bases its argument on the predicate that AI _at worst_ will increase developer productivity be 0-10%. But several studies have found that not to be true at all. AI can, and does, make some people less effective
- dhorthy 1y agodefinitely - the standford video has a slide about how many cases caused people to be even slower than without AI
- mcny 1y agoQuestion for discussion - what steps can I take as a human to set myself up for success where success is defined by AI made me faster, more efficient etc?
- 0xblacklight 1y agoIn many cases (though not all) it's the same thing that makes for great engineering managers: smart generalists with a lot of depth in maybe a couple of things (so they have an appreciation for depth and complexity) but a lot of breadth so they can effectively manage other specialists, and having great technical communication skills - be able to communicate what you want done and how without over-specifying every detail, or under-specifying tasks in important ways.
- Planktonne 1y ago>where success is defined by AI made me faster, more efficient etc? I think this attitude is part of the problem to me; you're not aiming to be faster or more efficient (and using AI to get there), you're aiming to use AI (to be faster and more efficient). A sincere approach to improvement wouldn't insist on a tool first.
- keeda 1y agoAccording to the Stanford video the only cases (statistically speaking) where that happened was high-complexity tasks for legacy / low popularity languages, no? I would imagine that is a small minority of projects. Indeed, the video cites the overall productivity boost at 15 - 20% IIRC.
- CharlesW 1y agoBoth (1) "AI can, and does, make some people less effective" and (2) "the average productivity boost (~20%) is significant" (per Stanford's analysis) can be true. The article at the link is about how to use AI effectively in complex codebases. It emphasizes that the techniques described are "not magic", and makes very reasonable claims.
- dingnuts 1y agothe techniques described sound like just as much work, if not more, than just writing the code. the claimed output isn't even that great, it's comparable to the speed you would expect a skilled engineer to move at in a startup environment
- CharlesW 1y ago> the techniques described sound like just as much work, if not more, than just writing the code. That's very fair, and I believe that's true for you and for many experienced software developers who are more productive than the average developer. For me, AI-assisted coding is a significant net win.
- dhorthy 1y agoI tend to think about it like vim - you will feel slow and annoyed for the first few weeks, but investing in these skills are massive +EV long term
- criemen 1y agoYet a lot of people never bother to learn vim, and are still outstanding and productive engineers. We're surely not seeing any memos "Reflexive vim usage is now a baseline expectation at [our company]" (context: https://x.com/tobi/status/1909251946235437514 https://x.com/tobi/status/1909251946235437514) The as-of-yet unanswered question is: Is this the same? Or will non-LLM-using engineers be left behind?
- f59b3743 1y ago[flagged]
- simonw 1y ago"AI can, and does, make some people less effective" So those people should either stop using it or learn to use it productively. We're not doomed to live in a world where programmers start using AI, lose productivity because of it and then stay in that less productive state.
- bgwalter 1y agoIf managers are convinced by stakeholders who relentlessly put out pro-"AI" blog posts, then a subset of programmers can be forced to at least pretend to use "AI". They can be forced to write in their performance evaluation how much (not if, because they would be fired) "AI" has improved their productivity.
- telliott1984 1y agoThere's also the more insidious gap between perceived productivity and actual productivity. Doesn't help that nobody can agree on how to measure productivity even without AI.
- tschellenbach 1y agoI wrote this blogpost on the same topic: https://getstream.io/blog/cursor-ai-large-projects/ https://getstream.io/blog/cursor-ai-large-projects/ It's super effective with the right guardrails and docs. It also works better on languages like Go instead of Python.
- dhorthy 1y agowhy do you think go is better than python (i have some thoughts but curious your take)
- polishdude20 1y agoprobably because it's typed?
- 0xblacklight 1y agoAmong other things; coding agents that can get feedback by running a compile step on top of the linter will tend to produce better output. Also, strongly-typed languages tend to catch more issues through the language server which the agent can touch through LSP.
- heavyset_go 1y agoPython is strongly typed
- mholm 1y agoimo: 1. Go's spec and standard practices are more stable, in my experience. This means the training data is tighter and more likely to work. 2. Go's types give the llm more information on how to use something, versus the python model. 3. Python has been an entry-level accessible language for a long time. This means a lot of the code in the training set is by amateurs. Go, ime, is never someone's first language. So you effectively only get code from someone who has already has other programming experience. 4. Go doesn't do much 'weird' stuff. It's not hard to wrap your head around.
- ath3nd 1y agoWhy though. Why should we do that? If AI is so groundbreaking, why do we have to have guides and jump through 3000 hoops just so we can make it work?
- ej88 1y agowhy do we have guides and lessons on how to use a chainsaw when we can hack the tree with an axe?
- leptons 1y agoThe chainsaw doesn't sometimes chop off your arm when you are using it correctly.
- crent 1y agoIf you swing an axe with a lack of hand eye coordination you don't think it's possible to seriously injure yourself?
- leptons 1y agoWas the axe or the chainsaw designed in such a way that guarantees that it will definitely miss the log and hit your hand fair amount of the times you use it? If it were, would you still use it? Yes, these hand tools are dangerous, but they were not designed so that it would probably cut off your hand even 1% of the time. "Accidents happen" and "AI slop" are not even remotely the same. So then with "AI" we're taking a tool that is known to "hallucinate", and not infrequently. So let's put this thing in charge of whatever-the-fuck we can? I have no doubt "AI" will someday be embedded inside a "smart chainsaw", because we as humans are far more stupid than we think we are.
- logicchains 1y agoEven if we had perfectly human-level AI it'd still need management, just like human workers do, and turns out effective management is actually nontrivial.
- shafyy 1y ago> Heck even Amjad was on a lenny's podcast 9 months ago talking about how PMs use Replit agent to prototype new stuff and then they hand it off to engineers to implement for production. Please kill me now
- ath3nd 1y ago[dead]
- cube00 1y agoI got lectured this week that I wasn't working fast enough because the client had already vibe coded (a broken, non-functional prototype) in under an hour. They saw the the first screen assembled by Replit and figured everything they could see would work with some "small tweaks" which is where I was allegedly to come into the picture. They continued to lecture me about how the app would need Web Workers for maximum client side performance (explanations full of em-dashes so I knew they were pasting in AI slop at me) and it must all be browser based with no servers because "my prototype doesn't need a server" Meanwhile their "prototype" had a broken Node.js backend running alongside the frontend listening on a TCP port. When I asked about this backend they knew nothing about it be assured me their prototype was all browser based with no "servers". Needless to say I'm never taking on any work from that client again, one of the small joys of being a contractor.
- shafyy 1y agoSounds like hell
- cheschire 1y agoAs an aside, this single markdown file as an entire GitHub repo is a unique approach to blog posts.
- dhorthy 1y agos/unique/lazy
- merlincorey 1y agoIt seems we're still collectively trying to figure out the boundaries of "delegation" versus "abstraction" which I personally don't think are the same thing, though they are certainly related and if you squint a bit you can easily argue for one or the other in many situations. > We've gotten claude code to handle 300k LOC Rust codebases, ship a week's worth of work in a day, and maintain code quality that passes expert review. This seems more like delegation just like if one delegated a coding task to another engineer and reviewed it. > That in two years, you'll be opening python files in your IDE with about the same frequency that, today, you might open up a hex editor to read assembly (which, for most of us, is never). This seems more like abstraction just like if one considers Python a sort of higher level layer above C and C a higher level layer above Assembly, except now the language is English. Can it really be both?
- dhorthy 1y agoI would say its much more about abstraction and the leverage abstractions give you. You'll also note that while I talk about "spec driven development", most of the tactical stuff we've proven out is downstream of having a good spec. But in the end a good spec is probably "the right abstraction" and most of these techniques fall out as implementation details. But to paraphrase sandy metz - better to stay in the details than to accidentally build against the wrong abstraction (https://sandimetz.com/blog/2016/1/20/the-wrong-abstraction https://sandimetz.com/blog/2016/1/20/the-wrong-abstraction) I don't think delegation is right - when me and vaibhav shipped a week's worth of work in a day, we were DEEPLY engaged with the work, we didn't step away from the desk, we were constantly resteering and probably sent 50+ user messages that day, in addition to some point-edits to markdown files along the way.
- sarchertech 1y agoIt’s definitely not abstraction. You don’t watch a compiler output machine code and constantly “resteer” it.
- sothatsit 1y agoI continue to write codebases in programming languages, not English. LLM agents just help me manipulate that code. They are tools that do work for me. That is delegation, not abstraction. To write and review a good spec, you also need to understand your codebase. How are you going to do that without reading the code? We are not getting abstracted away from our codebases. For it to be an abstraction, we would need our coding agents to not only write all of our code, they would also need to explain it all to us. I am very skeptical that this is how developers will work in the near future. Software development would become increasingly unreliable as we won't even understand what our codebases actually do. We would just interact with a squishy lossy English layer.
- techlatest_net 1y ago[dead]
- fusslo 1y agoMaybe I am just misunderstanding. I probably am; seems like it happens more and more often these days But.. I hate this. I hate the idea of learning to manage the machine's context to do work. This reads like a lecture in an MBA class about managing certain types of engineers, not like an engineering doc. Never have I wanted to manage people. And never have I even considered my job would be to find the optimum path to the machine writing my code. Maybe firmware is special (I write firmware)... I doubt it. We have a cursor subscription and are expected to use it on production codebases. Business leaders are pushing it HARD. To be a leader in my job, I don't need to know algorithms, design patterns, C, make, how to debug, how to work with memory mapped io, what wear leveling is, etc.. I need to know 'compaction' and 'context engineering' I feel like a ship corker inspecting a riveted hull
- jnwatson 1y agoI've started to use agents on some very low-level code, and have middling results. For pure algorithmic stuff, it works great. But I asked it to write me some arm64 assembly and it failed miserably. It couldn't keep track of which registers were which.
- jmkni 1y agoI imagine the LLM's have been trained on a lot less firmware code than say, HTML
- dolebirchwood 1y agoGuess it boils down to personality, but I personally love it. I got into coding later in life, and coming from a career that involved reading and writing voluminous amounts of text in English. I got into programming because I wanted to build web applications, not out of any love for the process of programming in and of itself. The less I have to think and write in code, the better. Much happier to be reading it and reviewing it than writing it myself.
- skydhash 1y agoNo ones like programming that much. That's like saying someone love speaking English. You have an idea and you express it. Sometimes there's additional complexity that got in the way (initializing the library, memory cleanup,...), but I put those at the same level as proper greetings in a formal letter. It also helps starting small, get something useful done and iterate by adding more features overtime (or keeping it small).
- daxfohl 1y ago> And yeah sure, let's try to spend as many tokens as possible It'd be nice if the article included the cost for each project. A 35k LOC change in a 350k codebase with a bunch of back and forth and context rewriting over 7 hours, would that be a regular subscription, max subscription, or would that not even cover it?
- CharlesW 1y agoFrom a cost perspective, you would definitely want a Claude Max subscription for this.
- dhorthy 1y agoyes - correct. For the record, if spending raw tokens, the 2 prs to baml cost about $650. but yes we switched off per-token this week because we ran out of anthropic credits, we're on max plan now
- daxfohl 1y agoHaha, when I asked claude the question, it estimated $20-45. https://claude.ai/share/5c3b0592-7bc9-4c40-9049-459058b16920 https://claude.ai/share/5c3b0592-7bc9-4c40-9049-459058b16920 Horrible, right? When I asked gemini, it guessed 37 cents! https://g.co/gemini/share/ff3ed97634ba https://g.co/gemini/share/ff3ed97634ba
- daxfohl 1y agoOh, oops it says further down > oh, and yeah, our team of three is averaging about $12k on opus per month I'll have to admit, I was intrigued with the workflow at first. But emm, okay, yeah, I'll keep handwriting my open source contributions for a while.
- vanillax 1y agoDoesnt githubs new speckit solve this? https://github.com/github/spec-kit https://github.com/github/spec-kit
- 0xblacklight 1y agohow does this solve it?
- ooopakaj 1y agoI’m not an expert in either language, but seeing a 20k LoC PR go up (linked in the article) would be an instant “lgtm, asshole” kind of review. > I had to learn to let go of reading every line of PR code Ah. And I’m over here struggling to get my teammates to read lines that aren’t in the PR. Ah well, if this stuff works out it’ll be commoditized like the author said and I’ll catch up later. Hard to evaluate the article given the authors financial interest in this succeeding and my lack of domain expertise.
- ActionHank 1y agoI dunno man, I usually close the PR when someone does that and tell them to make more atomic changes. Would you trust an colleague who is over confident, lies all the time, and then pushes a huge PR? I wouldn't.
- Our_Benefactors 1y agoClosing someone else’s PR is an actively hostile move. Opening a 20k LOC isn’t great either, but going ahead and closing it is rude as hell.
- ActionHank 1y agoDumping a huge PR across a shared codebase wherein everyone else also has to deal with the risk of you monumental changes is pretty rude as well, I would even go so far as to say that it is likely selfishly risky.
- bak3y 1y agoOpening a 20k LOC PR is an actively hostile move worthy of an appropriate response. Closed > will not review > make more atomic changes.
- wwweston 1y agoA 20k LOC PR isn’t reviewable in any normal workflow/process. The only moves are refusing to review it, taking it up the chain of authority, or rubber stamping it with a note to the effect that it’s effectively unreviewable so rubber stamping must be the desired outcome.
- koakuma-chan 1y agoContext has never been the bottleneck for me. AI just stops working when I reach certain things that AI doesn't know how to do.
- jmkni 1y agoMy problem is it keeps working, even when it reaches certain things it doesn't know how to do. I've been experimenting with Github agents recently, they use GPT-5 to write loads of code, and even make sure it compiles and "runs" before ending the task. Then you go and run it and it's just garbage, yeah it's technically building and running "something", but often it's not anything like what you asked for, and it's splurged out so much code you can't even fix it. Then I go and write it myself like the old days.
- koakuma-chan 1y agoI have same experience with CC. It loves to comment out code, add a "fallback" implementation that returns mock data, and act like the thing works.
- 0xblacklight 1y ago> Context has never been the bottleneck for me. AI just stops working when I reach certain things that AI doesn't know how to do. It's context all the way down. That just means you need to find and give it the context to enable it to figure out how to do the thing. Docs, manuals, whatever. Same stuff that you would use to enable a human that doesn't know how to do it to figure out how.
- koakuma-chan 1y agoAt that point it's easier to implement the thing yourself, and then let AI work with that.
- bluefirebrand 1y agoOr just forget the AI entirely, if you can build it yourself then do it yourself I treat "uses AI tools" as a signal that a person doesn't know what they are doing
- philipp-gayret 1y agoCan't agree with the formula for performance, on the "/ size" part. You can have a huge codebase, but if the complexity goes up with size then you are screwed. Wouldn't a huge but simple codebase be practical and fine for AI to deal with? The hierarchy of leverage concept is great! Love it. (Can't say I like the 1 bad line of CLAUDE.md is 100K lines of bad code; I've had some bad lines in my CLAUDE.md from time to time - I almost always let Claude write it's own CLAUDE.md.).
- dhorthy 1y agoi mean there's also the fact that claude code injects this system message into your claude.md which means that even if your claude.md sucks you will probably be okay: <system-reminder> IMPORTANT: this context may or may not be relevant to your tasks. You should not respond to this context or otherwise consider it in your response unless it is highly relevant to your task. Most of the time, it is not relevant. </system-reminder> lots of others have written about this so i won't go deep but its a clear product decision, but if you don't know what's in your context window, you can't respond/architect your balance between claude.md and /commands well.
- jascha_eng 1y agoExcept for ofc pushing their own product (humanlayer) and some very complex prompt template+agent setups that are probably overkill for most, the basics in this post about compaction and doing human review at the correct level are pretty good pointers. And giving a bit of a framework to think within is also neat
- hellovai 1y agoif you haven't tried the research -> plan -> implementation approach here, you are missing out on how good LLMs are. it completely changed my perspective. the key part was really just explicitly thinking about different levels of abstraction at different levels of vibecoding. I was doing it before, but not explicitly in discrete steps and that was where i got into messes. The prior approach made check pointing / reverting very difficult. When i think of everything in phases, i do similar stuff w/ my git commits at "phase" levels, which makes design decision easier to make. I also do spend ~4-5 hours cleaning up the code at the very very end once everything works. But its still way faster than writing hard features myself.
- 0xblacklight 1y agotbh I think the thing that's making this new approach so hard to adopt for many people is the word "vibecoding" Like yes vibecoding in the lovable-esque "give me an app that does XYZ" manner is obviously ridiculous and wrong, and will result in slop. Building any serious app based on "vibes" is stupid. But if you're doing this right, you are not "coding" in any traditional sense of the word, and you are *definitely* not relying on vibes Maybe we need a new word
- dhorthy 1y agoalex reibman proposed hyperengineering i've also heard "aura coding", "spec-driven development" and a bunch of others I don't love. but we def need a new word cause vibe coding aint it
- Xss3 1y agoVibe coding is accepting ai output based on vibes. Simple as that. You can vibe code using specs or just by having a conversation.
- simonw 1y agoI'm sticking to the original definition of "vibe coding", which is AI-generated code that you don't review. If you're properly reviewing the code, you're programming. The challenge is finding a good term for code that's responsibly written with AI assistance. I've been calling it "AI-assisted programming" but that's WAY too long.
- r2ob 1y agoI refactored CPython using GPT-5, turning the compiler bilingual for english and portuguese keywords. https://github.com/ricardoborges/cpython https://github.com/ricardoborges/cpython what web programming task GPT-5 can't handle?
- iambateman 1y agoI built a package which I use for large codebase work[0]. It starts with /feature, and takes a description. Then it analyzes the codebase and asks questions. Once I’ve answered questions, it writes a plan in markdown. There will be 8-10 markdowns files with descriptions of what it wants to do and full code samples. Then it does a “code critic” step where it looks for errors. Importantly, this code critic is wrong about 60% of the time. I review its critique and erase a bunch of dumb issues it’s invented. By that point, I have a concise folder of changes along with my original description, and it’s been checked over. Then all I do is say “go” to Claude Code and it’s off to the races doing each specific task. This helps it keep from going off the rails, and I’m usually confident that the changes it made were the changes I wanted. I use this workflow a few times per day for all the bigger tasks and then use regular Claude code when I can be pretty specific about what I want done. It’s proven to be a pretty efficient workflow. [0] GitHub.com/iambateman/speedrun
- j45 1y agoThis looks very cool. I see it has a pseudo code step, was it helpful at all to try to define a workflow, process or procedure beforehand? I've also heard that keeping each file down to 100 lines is critical before connecting them. Noticed the same but haven't tried it in depth.
- dhorthy 1y agoFile size matters if you don’t have strategically placed “read the entire file” instructions for certain parts of the workflow (we do)
- eddywebs 1y agoThe biggest challenge i found with LLMs on large codebase is making the same mistakes again and again How do keep track of the architecture decisions in context of every tasks on the large codebase ?
- tom_m 1y agoVery very clear, unambiguous, prompts and agent rules. Use strong language like "must" and "critical" and "never" etc. I would also try working on smaller sections of a large codebase at a time too if things are too inaccurate. The AI coding tools are going to be looking at other files in the project to help with context. Ambiguity is the death of AI effectiveness. You have to keep things clear and so that may require addressing smaller sections at a time. Unless you can really configure the tools in ways to isolate things. This is why I like tools that have a lot of control and are transparent. If you ask a tool what the full system and user prompt is and it doesn't tell you? Run away from that tool as fast as you can. You need to have introspections here. You have to be able to see what causes a behavior you don't want and be able to correct it. Any tool that takes that away from you is one that won't work.
- iagooar 1y agoI am working on a project with ~200k LoC, entirely written with AI codegen. These days I use Codex, with GPT-5-Codex + $200 Pro subscription. I code all day every day and haven't yet seen a single rate limiting issue. We've come a long way. Just 3-4 months ago, LLMs would start doing a huge mess when faced with a large codebase. They would have massive problems with files with +1k LoC (I know, files should never grow this big). Until recently, I had to religiously provide the right context to the model to get good results. Codex does not need it anymore. Heck, even UI seems to be a solved problem now with shadcn/ui + MCP. My personal workflow when building bigger new features: 1. Describe problem with lots of details (often recording 20-60 mins of voice, transcribe) 2. Prompt the model to create a PRD 3. CHECK the PRD, improve and enrich it - this can take hours 4. Actually have the AI agent generate the code and lots of tests 5. Use AI code review tools like CodeRabbit, or recently the /review function of Codex, iterate a few times 6. Check and verify manually - often times, there are a few minor bugs still in the implementation, but can be fixed quickly - sometimes I just create a list of what I found and pass it for improving With this workflow, I am getting extraordinary results. AMA.
- mentos 1y agoWhat platform are you developing for, web? Did you start with Cursor and move to Codex or only ever Codex?
- drewnick 1y agoNot OP, but I use Codex for back-end, scripting, and SQL. Claude Code for most front-end. I have found that when one faces a challenge, the other often can punch through and solve the problem. I even have them work together (moving thoughts and markdown plans back and fourth) and that works wonders. My progression: Cursor in '24, Roo code mid '25, Claude Code in Q2 '25, Codex CLI in Q3 `25.
- iagooar 1y agoCursor for me until 3-4 weeks ago, now Codex CLI most of the time. These tools change all the time, very quickly. Important to stay open to change though.
- potamic 1y agoThere are a lot of people declaring this, proclaiming that about working with AI, but nobody presents the details. Talk is cheap, show me the prompts. What will be useful is to check in all the prompts along with code. Every commit generated by AI should include a prompt log recording all the prompts that led to the change. One should be able to walkthrough the prompt log just as they may go through the commit log and observe firsthand how the code was developed.
- troupo 1y agoThis blog post of mine will be evergreen: https://dmitriid.com/everything-around-llms-is-still-magical-and-wishful-thinking https://dmitriid.com/everything-around-llms-is-still-magical...
- rhetocj23 1y agoMoreover, show me the money!!
- an0malous 1y agoI agree, the rare times when someone has shared prompts and AI generated code I have not been impressed at all. It very quickly accrues technical debt and lacks organization. I suspect the people who say it’s amazing are like data engineers who are used to putting everything in one script file, React devs where the patterns and organization are well defined and constrained, or people who don’t code and don’t even understand the issues in their generated code yet.
- spariev 1y agoThanks for sharing, I wonder how do you keep the stylistic and mental alignment of the codebase - is this happens during the code review or there are specific instructions during at the plan/implement stages?
- jgilias 1y agoI used to do these things manually in Cursor. Then I had to take a few months off programming, and when I came back and updated Cursor I found out that it now automatically does ToDos, as well as keeps track of the context size and compresses it automatically by summarising the history when it reaches some threshold. With this I find that most of the shenanigans of manual context window managing with putting things in markdown files is kind of unnecessary. You still need to make it plan things, as well as guide the research it does to make sure it gets enough useful info into the context window, but in general it now seems to me like it does a really good job with preserving the information. This is with Sonnet 4 YMMV
- gloosx 1y agoIt's strange that author is bragging that this 35K LOC was researched and implemented in 7 hours, but there are 40 commits spanning across 7 days. Was it 1 hour per day or what? Also quite funny that one of the latest commits is "ignore some tests" :D
- dhorthy 1y agoif you read further down, I acknowledge this > While the cancelation PR required a little more love to take things over the line, we got incredible progress in just a day.
- daxfohl 1y agoFWIW I think your style is better and more honest than most advocates. But I'd really love to see some examples of things that completely failed. Because there have to be some, right? But you hardly ever see an article from an AI advocate about something that failed, nor from an AI skeptic about something that succeeded. Yet I think these would be the types of things that people would truly learn from. But maybe it's not in anyone's financial interest to cross borders like that, for those who are heavily vested in the ecosystem.
- daxfohl 1y agoBut, yeah, looking again, that was a pretty big omission. And even moreso, a missed opportunity! I think if this had been called out more explicitly, then rather than arguing whether this is a realistic workflow or not, we'd be seeing more thoughtful conversation about how to fix the remaining problems. I don't mean to sound discouraging. Keep up the good work!
- dhorthy 1y agothere is a portion in the article where I talk about how our hadoop refactor completely failed
- svieira 1y ago
- wrs 1y ago1. Research -> Plan -> Implement 2. Write down the principles and assumptions behind the design and keep them current In other words, the same thing successful human teams on complex projects do! Have we become so addicted to “attention-deficit agile” that this seems like a new technique? Imagine, detailed specs, design documents, and RFC reviews are becoming the new hotness. Who would have thought??
- dhorthy 1y agoyeah its kinda funny how some bigger more sophisticated eng orgs that would be called "slow and ineffective" by smaller teams are actually pretty dang well set-up to leverage AI. All because they have been forced to master technical communication at scale. but the reason I wrote this (and maybe a side effect of the SF bubble) is MOST of the people I have talked to, from 3-person startups to 1000+ employee public companies, are in a state where this feels novel and valuable, not a foregone conclusion or something happening automatically
- rybosworld 1y agoTLDR: We're taking a profession that attracts people who enjoy a particular type of mental stimulation, and transforming it into something that most members of the profession just fundamentally do not enjoy. If you're a business leader wondering why AI hasn't super charged your company's productivity, it's at least partly because you're asking people to change the way they work so drastically, that they no longer derive intrinsic motivation from it. Doesn't apply to every developer. But it's a lot.
- GoatInGrey 1y agoHello, I noticed your privacy policy is a black page with text seemingly set to 1% or so opacity. Can you get the slopless AI to fix that when time permits? - Mr. Snarky
- dhorthy 1y agothank you for the feedback! themes are hard. Update going out now
- procaryote 1y agoThey wanted a transparent privacy policy
- ghm2199 1y agoThis article is like a bookmark in time of where I exactly gave up (in July) managing context in Claude code. I made specs for every part of the code in a separate folder and that had in it logs on every feature I worked on. It was an API server in python with many services like accounts, notifications, subscriptions etc. It got to the point where managing context became extremely challenging. Claude would not be able to determine business logic properly and it can get complex. e.g. if you want to do a simple RBAC system with an account and profile with a junction table for roles joining an account with profile. In the end what kind of worked was I had to give it UML diagrams of the relationship with examples to make it understand and behave better.
- dhorthy 1y agoi think that was one of the key reasons we built research_codebase.md first - the number one concern is "what happens if we end up owning this codebase but don't know how it works / don't know how to steer a model on how to make progress" There are two common problems w/ primarily-AI-written code 1. Unfamiliar codebase -> research lets you get up to speed quickly on flows and functionality 2. Giant PR Reviews Suck -> plans give you ordered context on what's changing and why Mitchell has praised ampcode for the thread sharing, another good solution to #2 - https://x.com/mitchellh/status/1963277478795026484 https://x.com/mitchellh/status/1963277478795026484
- svieira 1y ago> the number one concern "what happens if we end up owning this codebase but ... don't know how to steer a model on how to make progress" > Research lets you get up to speed quickly on flows and functionality This is the _je ne sais quoi_ that people who are comfortable with AI have made peace with and those who are not have not. If you don't know what the code base does or how to make progress you are effectively trusting the system that built the thing you don't understand to understand the thing and teach you. And then from that understanding you're going to direct the teacher to make changes to the system it taught you to understand. Which suggests a certain _je ne sais quoi_ about human intelligence that isn't present in the system, but which would be necessary to create an understanding of the thing under consideration. Which leads to your understanding being questionable because it was sourced from something that _lacks_ that _je ne sais quoi_. But the order time of failure here is "lifetimes". Of features, of codebases, of persons.
- varjag 1y agoA few weeks later, @hellovai and I paired on shipping 35k LOC to BAML, adding cancellation support and WASM compilation - features the team estimated would take a senior engineer 3-5 days each. Sorry, had they effectively estimated that an engineer should produce 4-6KLOC per day (that's before genAI)?
- henry2023 1y agoThe missing detail here is that the senior engineer would probably have shipped it in 2k lines of code
- gigel82 1y agoOr 1k lines of functional, readable, testable, commented code... but who cares, we'll abstract it all away soon enough.
- rsynnott 1y agoAnd note that, as admitted elsewhere, it _actually_ took a week: https://news.ycombinator.com/item?id=45351546 https://news.ycombinator.com/item?id=45351546
- afiodorov 1y ago> It was uncomfortable at first. I had to learn to let go of reading every line of PR code. I still read the tests pretty carefully, but the specs became our source of truth for what was being built and why. This is exactly right. Our role is shifting from writing implementation details to defining and verifying behavior. I recently needed to add recursive uploads to a complex S3-to-SFTP Python operator that had a dozen path manipulation flags. My process was: * Extract the existing behavior into a clear spec (i.e., get the unit tests passing). * Expand that spec to cover the new recursive functionality. * Hand the problem and the tests to a coding agent. I quickly realized I didn't need to understand the old code at all. My entire focus was on whether the new code was faithful to the spec. This is the future: our value will be in demonstrating correctness through verification, while the code itself becomes an implementation detail handled by an agent.
- cm2012 1y agoClaude Plays Pokemon showed that too. AI is bad at deciding when something is "working" - it will go in circles forever. But an AI combined with a human to occasionally course correct is a powerful combo.
- nine_k 1y ago> My entire focus was on whether the new code was faithful to the spec This may be true, but see Postel's Law, that says that the observed behavior of a heavily-used system becomes its public interface and specification, with all its quirks and implementation errors. It may be important to keep testing that the clients using the code are also faithful to the spec, and detect and handle discrepancies.
- patrickmay 1y agoI believe that's Hyrum's Law.
- lunarcave 1y ago> Our role is shifting from writing implementation details to defining and verifying behavior. I could argue that our main job was always that - defining and verifying behavior. As in, it was a large part of the job. Time spent on writing implementation details have always been on a downward trend via higher level languages, compilers and other abstractions.
- iLoveOncall 1y ago> Within an hour or so, I had a PR fixing a bug which was approved by the maintainer the next morning An hour for 14 lines of code. Not sure how this shows any productivity gain from AI. It's clear that it's not the code writing that is the bottleneck in a task like this. Looking at the "30K lines" features, the majority of the 30K lines are either auto-generated code (not by AI), or documentation. One of them is also a PoC and not merged...
- mwigdahl 1y agoThe author said he was not a Rust expert and had no prior familiarity with the codebase. An hour for a 14 line fix that works and is acceptable quality to merge is pretty good given those conditions.
- asdev 1y agoThe problem is the research phase will fail because you can't glean tribal product knowledge from just looking at the code
- faxmeyourcode 1y agoI've used this pattern on two separate codebases. One was ~500k LOC apache airflow monolith repo (I am a data engineer). The other was a greenfield flutter side project (I don't know dart, flutter, or really much of anything regarding mobile development). All I know is that it works. On the greenfield project the code is simple enough to mostly just run `/create_plan` and skip research altogether. You still get the benefit of the agents and everything. The key is really truly reviewing the documents that the AI spits out. Ask yourself if it covered the edge cases that you're worried about or if it truly picked the right tech for the job. For instance, did it break out of your sqlite pattern and suggest using postgres or something like that. These are very simple checks that you can spot in an instant. Usually chatting with the agent after the plan is created is enough to REPL-edit the plan directly with claude code while it's got it all in context. At my day job I've got to use github copilot, so I had to tweak the prompts a bit, but the intentional compaction between steps still happens, just not quite as efficiently because copilot doesn't support sub-agents in the same way as claude code. However, I am still able to keep productivity up. ------- A personal aside. Immediately before AI assisted coding really took off, I started to feel really depressed that my job was turning into a really boring thing for me. Everything just felt like such a chore. The death by a million paper cuts is real in a large codebase with the interplay and idiosyncrasies of multiple repos, teams, personalities, etc. The main benefit of AI assisted coding for me personally seems to be smoothing over those paper cuts. I derive pleasure from building things that work. Every little thing that held up that ultimate goal was sucking the pleasure out of the activity that I spent most of my day trying to do. I am much happier now having impressed myself with what I can build if I stick to it.
- dhorthy 1y agoI appreciate the share. Yes as I said it was a pretty dang uncomfortable to transition to this new way of working but now that it’s settled we’re never going back
- qazxcvbnmlp 1y agothe ratio of {solve interesting problem} to {remember how to compile foo bar with update 23.12.21938 on a full moon} is so much better now
- rationalfaith 1y ago[dead]
- wobblyasp 1y agoVerifying behavior is great and all if you can actually exhaustively test the behaviors of your system. If you can't, then not knowing what your code is actually doing is going to set you back when things do go belly up.
- ipnon 1y agoI love this comment because it makes perfect sense today, it made perfect sense 10 years ago, it would have made perfect sense in 1970. The principles of software engineering are not changed by the introduction of commodified machine intelligence.
- dhorthy 1y agoi 100% agree - the folks who are best at ai-first engineering, they spend 3 days designing the test harness and then kick off an agent unsupervised for 2+ days and come back to working software. not exactly valuable as guidance since programming languages are very easy to verify, but the https://ghuntley.com/ralph https://ghuntley.com/ralph post is an example of whats possible on the very extreme end of the spectrum
- robertoallende 1y agoThanks to write such detailed article... lot of very well supported information. I've been working on something what I call Micromanaged Driven Development https://mmdd.dev https://mmdd.dev and wrote about it at https://builder.aws.com/content/2y6nQgj1FVuaJIn9rFLThIslwaJ/code-with-ai-micromanagement-is-all-you-need https://builder.aws.com/content/2y6nQgj1FVuaJIn9rFLThIslwaJ/... I'm in a similar search and I'm stoked to see that many people riding the wave of coding with AI is moving in this direction. Lots of learning ahead.
- marcuschong 1y agoI'm using GPT Pro and a VS extension that makes it easy to copy code from multiple files at once. I'm architecting the new version of our SaaS and using it to generate everything for me on the backend. It’s a huge help with modeling and coding, though it takes a lot of steering and correction. I think I’ll end up with a better result than if I did it alone, since it knows many patterns and details I’m not aware of (even simple things like RRULE). I’m designing this new project with a simpler, more vertical architecture in the hopes that Codex will be able to create new tables and services easily once the initial structure is ready and well documented. Edit: typo.
- dhorthy 1y agoyeah flat, simple code is good to start, but I find I'm still developing instincts around right balance between "when to let duplicate code sprawl" vs. "when to be the DRY police".
- CuriouslyC 1y agoAgents get really confused by duplicate code, so I advise DRYing out early and often.
- jwpapi 1y agoHonestly I was reading that article and I smelled sales pitch and then in the end of course. This is not my experience at all. I also don’t get the line obsession. Good code has less lines not more
- skydhash 1y agoIt seems like a different universe from the openbsd and the suckless guys.
- deleted 1y ago[deleted]
- onscreencomb 1y agoI created an account to say this: RepoPrompt's 'Context Builder' feature helps a ton with scoping context before you touch any code. It's kind of like if you could chat with Repomix or Gitingest so they only pull the most relevant parts of your codebase into a prompt for planning, etc I'm a paying RepoPrompt user but not associated in any other way. I've used it in conjunction with Codex, Claude Code, and any other code gen tool I have tried so far. It saves a lot of tokens and time (and headaches)
- SafeDusk 1y agoOpenAI Codex has an `update_plan` function[0]. I'm wondering if switching the implementation to this would improve the coding agent's capabilities or is the default for simplicity better. [0]: https://blog.toolkami.com/openai-codex-tools/ https://blog.toolkami.com/openai-codex-tools/
- jongjong 1y agoWhen I read about people dumping 2000 lines of code every few days, I'm extremely skeptical about the quality of this code. All the people I've met who worked at this rate were always going for naive solutions and their code was full of hard-to-see bugs which only reared their ugly heads once in a while and were impossible to debug.
- pcannons 1y agoNice, my experience writing production code in large codebase here, granted it's evolved a lot since: https://philippcannons.com/100x-ai-coding-for-real-work-summon-the-ghost-army/ https://philippcannons.com/100x-ai-coding-for-real-work-summ... Not surprisingly, building really fast is not the silver bullet you'd think it is. It's all about what to build and how to distribute it. Otherwise bigcos/billionaires would have armies of engineers growing their net worth to epic scales.
- c54 1y agoRegarding billionaires having armies of engineers growing their wealth to massive scale: is that not what they have?
- pcannons 1y agoMy current world view: for monster multiples you need someone who knows how to go 0 to 1, repeatedly. That's almost always only the founder. People after are incremental. If they weren't, they'd just be a founder. Hence why everything is done through acquisitions post-founder. So there's armies of engineers incrementally scaling and maintaining dollars. But not creating that wealth or growing it in a significant % way.
- kkpattern 1y agoI used a similar pattern. When ask AI to do a large implementation. I ask gemini-2.5-pro to write a very detailed overview implementation plan. Then review it. Then ask gemini-2.5-pro to split the plan into multiple stages and write detail implementation plan for each stage. Then I ask claude sonnat to read the overview plan and implement the stage n. I found that this is the only way to complete a major implementation with a relatively high success rate.
- cadamsdotcom 1y agoRe the meta of running multiple phases of "document expansion": Research helps with complex implementations and for brownfield. But it isn't always needed - simple bugfixes can be one-shot! So all AI workflows could be expressed with some number "N" of "document expansion phases": N(0): vibe coding. N(1): "write a spec then implement it while I watch". N(2): "research then specify". At this point you start to get serious steerability. What's N(3) and beyond? Strategy docs, industry research, monetization planning? Can AI do these too, all of it ending up in git? Interesting to muse on.
- pvncher 1y ago[dead]
- mercurialsolo 1y agoGood pointers on decompositing and looking at implementation or fixing in chunks. 1. Break down the feature or bug report into a technical implementation spec. Add in COT for the splits. 2. Verify the implementation spec. Feed reviews back to your original agent that has created the spec. Edit, merge, integrate feedback. 3. Transform implementation spec into an implementation plan - logically split into modules look at dependency chain. 4. Build, test and integrate continuously with coding agents 5. Squash the commits if needed into a single one for the whole feature. Generally has worked well as a process when working on a complex feature. You can add in HITL at each stage if you need more verification. For larger codebases always maintain an ARCHITECTURE.md and for larger modules a DESIGN.md
- spike021 1y agoI admittedly haven't tried this approach at work yet but at home while working on a side project, I'll make a new feature branch and give CLAUDE a prompt about what the feature is with as much detail as possible. i then have it generate a CLAUDE-feature.md and place an implementation plan along with any supporting information (things we have access to in the codebase, etc.). i'll then prompt it for more based on if my interpretation of the file is missing anything or has confusing instructions or details. usually in-between larger prompts I'll do a full /reset rather than /compact, have it reference the doc, and then iterate some more. once it's time to try implementing I do one more /reset, then go phase by phase of the plan in increments /reset-ing between each and having it update the doc with its progress. generally works well enough but not sure i'd trust it at work.
- dhorthy 1y agoMy advice - never use compact, always stash your context to Md or a wordy git commit message and then clear context You want control over and visibility into what’s being compacted, and /compact doesn’t do great on either
- Amaury-El 1y agoUsing AI to help with code felt like working with a smart but slightly unreliable teammate. If I wasn’t clear, it just couldn’t follow. But once I learned to explain what I wanted clearly and specifically, it actually saved me time and helped me think more clearly too.
- ath_ray 1y agoI enjoyed the emphasis on optimising the context window itself. I think that's the most important bit. An abstraction for this that seems promising to me for its completeness and size is a User Story paired with a research plan(?). This works well for many kinds of applications and emphasizes shipping concrete business value for every unit of work. I wrote about some of it here: https://blog.nilenso.com/blog/2025/09/15/ai-unit-of-work/ https://blog.nilenso.com/blog/2025/09/15/ai-unit-of-work/ I also think a lot of coding benchmarks and perhaps even RL environments are not accounting for the messy back and forth of real world software development, which is why there's always a gap between the promise and reality.
- dhorthy 1y agoI have had a user story and a research plan and only realized deep in the implementation that a fundamental detail about how the code works was missing (specifically, that types and sdks are generated from OpenAPI spec) - this missing meant the plan was wrong (didn’t read carefully enough) and the implementation was a mess
- ath_ray 1y agoYeah I agree. There's a lot more needed than just the User Story, one way I'm thinking about it is that the "core" is deliverable business value, and the "shells" are context required for fine-grained details. There will likely need to be a step to verify against the acceptance criteria. I hope to back up this hypothesis with actual data and experiments!
- HellDunkel 1y agoI am still sceptical of the roi and the time i am supposed to sink into trying and learning these AI tools which seem to be replacing each other every week.
- maltalex 1y agoInteresting read, and some interesting ideas, but there's a problem with statements like these: > Sean proposes that in the AI future, the specs will become the real code. That in two years, you'll be opening python files in your IDE with about the same frequency that, today, you might open up a hex editor to read assembly. > It was uncomfortable at first. I had to learn to let go of reading every line of PR code. I still read the tests pretty carefully, but the specs became our source of truth for what was being built and why. This doesn't make sense as long as LLMs are non-deterministic. The prompt could be perfect, but there's no way to guarantee that the LLM will turn it into a reasonable implementation. With compilers, I don't need to crack open a hex editor on every build to check the assembly. The compiler is deterministic and well-understood, not to mention well-tested. Even if there's a bug in it, the bug will be deterministic and debuggable. LLMs are neither.
- beefnugs 1y agosounds like a good nudge to make tests better
- oblio 1y agoIf the tests are written by the AI, who watches the watchers? :-)
- ozim 1y agoThe fun part is that specs already are non-deterministic. If you spend time to write out requirements in English in a way that cannot be misinterpreted in any way you end up with programming language.
- diavolodeejay 1y agoSo… COBOL?
- jlawson 1y agoSpecs are ambiguous but not necessarily non-deterministic. The same entity interpreting the spec in exactly the same way will resolve the ambiguities the same way each time. Human and current AI interpretation of specs is non-deterministic process. But, if we wanted to build a deterministic AI we could.
- Michael_Keller 1y ago[dead]
- Emma_Schmidt 1y ago[dead]
- jrecyclebin 1y agoLots of gold in this article. It's like discovering a basket of cheat codes. This will age well. Great links, BAML is a crazy rabbithole and just found myself nodding along to frequent /compact. These tips are hard-earned and very generously given. Anyone here can take it or leave it. I have theft on my mind, personally. (ʃƪ¬‿¬)
- lexoj 1y agoTo minimise context bloat and provide more holistic context, I extract on first step the important elements from a codebase via AST which then the LLM uses to determine which files to get in full for given task. https://github.com/piqoni/vogte https://github.com/piqoni/vogte
- madcocomo 1y agoFor me the biggest difficulty is I find it hard to read unverifiable documentation. It's like dyslexia - if I can't connect the text content with runnable code, I feel lost in 5 minutes. So with this approach of spending 3 hours on planning without verification in code, that's too hard for me. I agree the context compaction sounds good. But I'm not sure if an md file is good enough to carry the info from research to plan and implementation. Personally I often find the context is too complex or the problem is too big. I just open a new session to resolve a smaller, more specific problem in source code, then test and review the source code.
- jillesvangurp 1y agoWe're currently in a transition phase where we're using agentic coding on systems developed with tools and languages designed for humans. Ironically, this makes things unnecessarily hard as things that are easy for us aren't necessary easy to deal with; or that optimal for agentic coding systems. People like languages that are expressive and concise. That means they do things like omit types, use type inference, macros, syntactic sugar, allow for ambiguities and all the other stuff that gives us shorter, easier to type code that requires more effort to figure out. A good intuition here might be that the harder the compiler/interpreter has to work to convert it into running/executable code, the harder an LLM will have to work to figure out what that code does. LLMs don't mind verbosity and spelling things out. Things that are long winded and boring to us are helpful for an LLM. The optimal language for an LLM is going to be different than one that is optimal for a human. And we're not good at actually producing detailed specifications. Programming actually is the job of coming up with detailed specifications. Easy to forget when you are doing that but that's literally what programming is. You write some kind of specification that is then "compiled" into something that actually works as specified. The solution to agentic coding isn't writing specifications for our specifications. That just moves the problem. We've had a few decades of practice where we just happen to stuff code into files and use very primitive tools to manipulate those files. Agentic coding uses a few party tricks involving command line tools to manipulate those files and reading them by one into the precious context window. We're probably shoveling too much data around. But since that's the way we store code, there are no better tools to do that. From having used things like Codex, 99% of what it does is interrogating what's there via tediously slow prodding and poking around the code base using simple command line commands and build tool invocations. It's like watching paint dry. I usually just go off doing something else while it boils the oceans and does god knows what before finally doing the (usually) relatively straightforward thing that I asked it to do. It's easy to see that this doesn't scale that well. The whole point of a large code base is that it probably won't all fit in the context window. We can try to brute force the problem; or we can try to be more selective. The name of the game here is being able to be able to quickly select just the right stuff to put in there and discard all the rest. We can either do that manually (tedious and a lot of work, sort of as the article proposes), or make it easier for the LLM to use tools that do that. Possibly a bunch of poorly structured files in some nested directory hierarchy isn't the optimal thing here. Most non AI based automated refactorings require something that more closely resembles the internal data structures of what a compiler would use (e.g. symbol tables, definitions, etc.). A lot of what an agentic coding system has to do is reconstruct something similar enough to that just so it can build a context in which it can do constructive things. The less ambiguous and more structured that is, the easier the job. The easier we make it to do that, the more it can focus on solving interesting problems rather than getting ready to do that. I don't have all the answers here but if agentic coding is going to be most of the coding, it makes sense to optimize the tools, languages, etc. for that rather than for us.
- hufdr 1y ago[dead]
- djgrant 1y agoHow granular are the specs? Is it at the level of "this is the code you must write, and here is how to do it", or are you letting AI work some of that out?
- alanfranz 1y ago> our team of three is averaging about $12k on opus per month That’s usd 150k per year. Probably low for SF, but may be a lot in other areas.
- habinero 1y agoYou could almost hire a real engineer for that money.
- aitchnyu 1y ago> context management, and keeping utilization in the 40%-60% range (depends on complexity of the problem). Is this a rule of thumb? Will the cheaper (fewer params) models dumb down at 25%?
- rs186 1y ago> Sean proposes that in the AI future, the specs will become the real code. That in two years, you'll be opening python files in your IDE with about the same frequency that, today, you might open up a hex editor to read assembly (which, for most of us, is never). Only if AI code generation is correct 99.9% of the time and almost never hallucinates. We trust compilers and don't read assembly code because we know it's deterministic and the output can never be wrong (barring bugs and certain optimization issues, which are rare/one-time fixes). As long as generated code is not doing what the original "code" (in this case, specs) doing, humans need to go back to fix things themselves.
- afro88 1y agoI use a similar pattern but without the subagents. I get good results with it. I review and hand edit "research" and plans. I follow up and hand edit code changes. It makes me faster, especially in unfamiliar codebases. But the write up troubles me. If I'm reading correctly, he did 1 bugfix (approved and merged) and then 2 larger PRs (1 merged, 1 still in draft over a month later). That's an insanely small sample size to draw conclusions from. How can you talk like you've just proven the workflow works "for brownfield codebases"? You proved it worked for 2/3 tasks in 2 codebases, one failure (we can't say it works until the code is shipped IMO).
- klysm 1y agoTasted like a sales pitch the whole way and what do ya know at the very end there it is
- suninsight 1y agoSo I can attest to the fact that all of the things proposed in this article actually works. And you can try it out yourself on any arbitrary code base within few minutes. This is how: I work for a company called NonBioS.ai - we already implement most of what is mentioned in this article. Actually we implemented this about 6 months back and what we have now is an advanced version of the same flow. Every user in NonBioS gets a full linux VM with root access. You can ask nonbios to pull in your source code and ask it to implement any feature. The context is all managed automatically through a process we call "Strategic Forgetting" which is in someways an advanced version of the logic in this article. Strategic Forgetting handles the context automatically - think of it like automatic compaction. It evaluates information retention based on several key factors: 1. Relevance Scoring: We assess how directly information contributes to the current objective vs. being tangential noise 2. Temporal Decay: Information gets weighted by recency and frequency of use - rarely accessed context naturally fades 3. Retrievability: If data can be easily reconstructed from system state or documentation, it's a candidate for pruning 4. Source Priority: User-provided context gets higher retention weight than inferred or generated content The algorithm runs continuously during coding sessions, creating a dynamic "working memory" that stays lean and focused. Think of it like how you naturally filter out background conversations to focus on what matters. And we have tried it out in very complex code bases and it works pretty well. Once you know how well it works, you will not have a hard time believing that the days of using IDE's to edit code is probably numbered. Also - you can try it out for yourself very quickly at NonBioS.ai. We have a very generous free tier that will be enough for the biggest code base you can throw at nonbios. However, big feature implementations or larger refactorings might take time longer than what is afforded in the free tier.
- grbsh 1y agoThe fundamental frustration most engineers have with AI coding is that they are used to the act of _writing_ code being expensive, and the accumulation of _understanding_ happening for free during the former. AI makes the code free, but the understanding part is just as expensive as it always was (although, maybe the 'research' technique can help here). But let's assume you're much better than average at understanding code by reviewing it -- you have another frustrating experience to get through with AI. Pre-AI, let's say 4 days of the week are spend writing new code, while 1 day is spent fixing unforseen issues (perhaps incorrect assumption) that came up after production integration or showing things to real users. Post-AI, someone might be able to write those 4 days worth of code in 1 day, but making decisions about unexpected issues after integration doesn't get compressed -- that still takes 1 day. So post-AI, your time switches almost entirely from the fun, creative act of writing code to the more frustrating experience of figuring out what's wrong with a lot of code that is almost correct. But you're way ahead -- you've tested your assumptions much faster, but unfortunately that means nearly all of your time will now be spent in a state of feeling dumb and trying to figure out why your assumptions are wrong. If your assumptions were right, you'd just move forward without noticing.
- jarek83 1y agoNice read and I'm trying to follow and use your tools. But it's just seems hard to make Claude Code follow the instructions - it always diverges into researching and proposing fixes as well as searching for root causes - which is strictly said it the research_codebase not to do. From many tries I had a success only twice. Could you give some hints on how to use it?