4 ms·
You write a generic architecture document on how you want your code base to be organized, when to use pattern x vs pattern y, examples of what that looks like i
by nickstinemates 8mo ago
You write a generic architecture document on how you want your code base to be organized, when to use pattern x vs pattern y, examples of what that looks like in your code base, and you encode this as a skill.
Then, in your prompt you tell it the task you want, then you say, supervise the implementation with a sub agent that follows the architecture skill. Evaluate any proposed changes.
There are people who maximize this, and this is how you get things like teams. You make agents for planning, design, qa, product, engineering, review, release management, etc. and you get them to operate and coordinate to produce an outcome.
That's what this is supposed to be, encoded as a feature instead of a best practice.
- satellite2 8mo agoAren't you just moving the problem a little bit further? If you can't trust it will implement carefully specified features, why would you believe it would properly review those?
- frde_me 8mo agoIt's hard to explain, but I've found LLMs to be significantly better in the "review" stage than the implementation stage. So the LLM will do something and not catch at all that it did it badly. But the same LLM asked to review against the same starting requirement will catch the problem almost always The missing thing in these tools is that automatic feedback loop between the two LLMs: one in review mode, one in implementation mode.
- resonious 8mo agoI've noticed this too and am wondering why this hasn't been baked into the popular agents yet. Or maybe it has and it just hasn't panned out?
- bashtoni 8mo agoAnecdotaly I think this is in Claude Code. It's pretty frequent to see it implement something, then declare it "forgot" a requirement and go back and alter or add to the implementation.
- bethekidyouwant 8mo agoYou have to dump the context window for the review to work good.
- cbovis 8mo agoAFAICT this is already baked into the GitHub Copilot agent. I read its sessions pretty often and reviewing/testing after writing code is a standard part of its workflow almost every time. It's kind of wild seeing how diligent it is even with the most trivial of changes.
- tclancy 8mo agoHow does this not use up tokens incredibly fast though? I have a Pro subscription and bang up against the limits pretty regularly.
- doctoboggan 8mo agoIt _does_ use up tokens incredibly fast, which is probably why Anthropic is developing this feature. This is mostly for corporations using the API, not individuals on a plan.
- digdugdirk 8mo agoI'd love to see a breakdown of the token consumption of inaccurate/errored/unused task branches for claude code and codex. It seems like a great revenue source for the model providers.
- shafyy 8mo agoYeah, that's what I was thinking. They do have an incentive to not get everything right on the first try, as long as they don't over do it... I also feel like that they try to get more token usage by asking unnecesary follow up questions that the user may say yes to etc.
- andyferris 8mo agoIt does use tokens faster, yes.
- nickstinemates 8mo agoI don't know, all I can say is with API-based billing, doing multi-thousand like refactors that would take days to do costs like $4. In terms of value : effort, it's incredible.
- indemnity 8mo agoI had to go to Max, Pro is more like a taster. At work tho we use Claude Code thru a proxy that uses the model hosted on AWS bedrock. It’s slower than consumer direct-to-Anthropic and you have to wait a bit for the latest models (Opus 4.5 took a while to get), but if our stats are to be believed it’s much much cheaper.