5 ms·
Backpressure is all you need
- denysvitali 4mo agoThis seems to be the coding agents 101: build a strong feedback loop. Am I missing something?
- artursapek 4mo agoYeah I don’t really see the backpressure analogy here - it implies that the agent is constantly producing new stuff, which isn’t really possible since the solution is very detailed specs/goals.
- lucamark 4mo agoNo, there is nothing new in this article. This methodology is already adopted - and further optimized - by the sw community.
- mark_l_watson 4mo agoInteresting ideas for generalizing goals to reduce human labor in human <—> agent interactions. That said, maybe it is better to set up customized skills and infrastructure for large projects? At our early stage of trying to capture value of agentic systems, the good ideas in this article might be premature optimization.
- vermilingua 4mo ago> It should also reduce the number of low-quality PRs your teammates have to review for details the agent should have caught itself. Oh boy.
- Alifatisk 4mo agoCare to elaborate?
- atq2119 4mo agoThat quote shows an utter disregard for basic human decency. It is the responsibility of the person running the coding agent to make sure the resulting PRs are high quality. Putting that on your team mates, or worse, random open source project maintainers on the internet, is the definition of an extractive contribution.
- ovi256 4mo agoIt seems the OP agrees with you, and he's proposing a method for how to do so using agents.
- erooke 4mo ago> It is the responsibility of the person running the coding agent to make sure the resulting PRs are high quality. And > he's proposing a method for how to do so using agents Are not in agreement. The claim being made is that you shouldn't be sending PRs you haven't personally vetted to be high quality. Definitionally a bot cannot be used to personally vet something.
- khafra 4mo agoThis is not a contradiction; it's an augmentation. As an operations guy, I can tell you that well-constructed automation to reduce the amount of manual checking a human has to do almost always increases the quality of the overall process's output.
- atq2119 4mo agoOf course it does, but that's beside the point. As a software developer, you must never subject your team mates to a PR that you yourself believe to be low quality. The point of code review by others is to catch things that you missed. There are multiple lines of defense for quality. Yes, automation can and should be one of them, but your own self-review always has to come before review by your team mates.
- wellpast 4mo agoI’m willing to be wrong but this industry-wide emphasis on AI creative/coding workflows seems way over-engineered. Ime successful creative execution looks like micro-iterations where each output informs the next creative move. I can build something incredibly fast from essentially caveman grunt instructions through an LLM harness, iterating as I go. Optimizing for feeding a huge plan to an agent sounds to me like a net waste of time. And looking over the shoulder of industry peers trying to do this, I don’t see their outputs or throughput some remarkable improvement over what I can produce with minimal fanfare usage.
- piazz 4mo ago100% agree. Took me ages of working with the agents to circle back around this, which was the best way to get work done before AI automation anyways.
- 0x696C6961 4mo agoYeah it's wild watching so many people decide waterfall is great all of a sudden.
- NamTaf 4mo agoNever mind stumbling into proper engineering principles like having documented, testable requirements specifications.
- dcrazy 4mo agoI”ve been pretty happy with this side effect of the agentic coding bubble.
- NamTaf 4mo agoAs a non-tech engineer (mechanical, trains) it's fascinating seeing what is essentially the "not real engineers" SWE crew finally pay the piper because they've invoked what is in essence a non-compliant, cost-focused subcontractor and now need all of the same engineering rigours they never previously understood.
- dnnddidiej 4mo agoOh this is 101. Anyone not doing this? If not do it now!
- Arodex 4mo agoBecause your token use explodes?
- dnnddidiej 4mo agoCan be done on a claude pro. But if you are low on tokens then yeah probably stick to more of a non-thinking copilot type arrangement (which is fine!)
- haeseong 4mo ago[flagged]
- _zoltan_ 4mo ago"In this post, I’ll cover a third, not-so-obvious approach: building ways for the agent to validate more of its own work before a human has to step in. " this has been an obvious thing to do since at least January (since Geoffrey Huntley published "everything is a ralph loop"), and this is how I've been working: build enough orchestration tooling to be able to automate everything: development container bringup, building it, running the unit tests, doing integration testing, and using the software as eventually an end user. then to iterate set performance goals on an already solid basis so the automated agent ("gym") can go and iterate autonomously, and let you know when it's "done". I understand this probably does not work if you're on some subscription and not using the API (tokens burn fast), but this has been extremely productive for me.
- psychoslave 4mo agoWhat license do you use then?
- _zoltan_ 4mo agoyou can pay by just volume ("API pricing")
- kami23 4mo agoThis is where most of my productivity gains have come, I have a special harness I move from project to project now that does my testing orchestration, lots of my work day is setting up a prompt or two early and just letting them loop till they return evidence that the feature is working having gone through the big QA loop. I've slowly been optimizing for token use through the stack and Claude ends up making very tight for loops for most of the process and keeping token count even lower. It's been nice. A lot of my toil at work is just gone.
- osigurdson 4mo agoI can see how you could avoid regressions this way, but what do you add to your harness to prove that a new feature is working?
- cyanydeez 4mo agointeresting idea, unfortunately programming the structure is equivalent (P=NP) to just programming itself. same as TDD. as usual, the tool isnt really doing whats listed on its label. however, people are different so this might improve someones capability to deploy LLMs. might even provide better evidence where actual brain power is needed.
- xg15 4mo agoIsn't this a bit of an incorrect usage of the term "backpressure"? OP quoted the correct definition right at the start: > In systems engineering, backpressure is the mechanism by which a downstream component signals upstream that it can't accept more work (the "downstream component" being the human reviewer in this case) But the measures they propose don't actually do that. They are more like fixed throttle elements which would slow down the rate of submissions of an agent and weed out some low-quality submissions before hitting "downstream". I'm missing the connection to the actual capacity (or will) that the human developers have to review the submissions.
- jeffbee 4mo agoIt is an incorrect use of what was already a flawed metaphor. Pressure is isotropic. Directed pressure makes no sense, like all other fluid analogies in unrelated fields of engineering.
- brookst 4mo agoWait so cross ventilation, where a breeze will flow through a house if windows are open on opposite sides at a much greater rate than if windows are only open on the upwind side… isn’t really a thing?
- pcstl 4mo agoAir moves from regions of high pressure to regions of low pressure. Pressure itself does not have a direction.
- marcosdumay 4mo agoThe act of "making pressure" means applying a force and is completely directional.
- DoctorOetker 4mo agoso this is about lower or upper back pressure?
- SkiFreeWin3 4mo agoLooks like plenty of recent prior art on this: https://pura.xyz https://pura.xyz https://github.com/puraxyz/puraxyz/blob/main/docs/paper/main.pdf https://github.com/puraxyz/puraxyz/blob/main/docs/paper/main...
- indianrestrooms 4mo ago[dead]
- jongguk 4mo ago[dead]
- pshirshov 4mo agoA very long post about a simple and very obvious idea with many different implementations. The three main problems are 1) API usage is deadly expensive 2) Claude is about to make all automation very expensive 3) all the flows where a model has the initiative are strictly biased towards unwarranted stops (checkpointing). Also, I won't call that "backpressure", there is no producer-consumer disbalance or something similar. From what I can see, the author just proposes a structured feedback loop. That's a discussion about organizational principles for system which consist of multiple unreliable but very complex components and this "backpressure" is just one of the aspects. Personally I find the viable system model framework productive as both a mental model and literal implementation guideline. Lesser problem is that agent SDKs are bad and building a custom harness is hard.
- root-parent 4mo agohttps://en.wikipedia.org/wiki/Back_pressure https://en.wikipedia.org/wiki/Back_pressure
- pshirshov 4mo agoThe problem is not "backpressure", that's just one of the tools and there are different approaches with the same effect. You can't express orchestration in terms of "backpressure" only, I think. Implement-Review-Repeat loop does not involve backpressure in the strict meaning of the term.
- entrope 4mo ago> all the flows where a model has the initiative are strictly biased towards unwarranted stops Can you elaborate on what you think causes such a bias? My experience is that Qwen3.6, Claude Sonnet 4.6 and Opus 4.6/4.7 will work as far as they can given direction and a way to test their work. My so-far limited experience with Opus 4.8 is that it does stop somewhat earlier for feedback, but in places where I am glad it is checking assumptions or where I agree with it identifying a change in scope (for example, where the following work deserves a separate commit or merge request). I would call those justified stops rather than unwarranted.
- yearesadpeople 4mo agoIf the systems invariants are well defined, and a suite of conformance + requirements tests (ensuring invariance is respected) are defined, wouldn't this be a broad - _'base case'_ - approach in general?
- jon-wood 4mo agoThis what hooks[1] are for, except hooks allow specifying criteria in certain conditions (like the agent believing it’s done and ready to hand back to the user) in a manner that the agent won’t just forget about once it’s a few turns deep, and doesn’t require triggering a whole other LLM instance to read some plain text instructions while you hope it interprets them correctly. It absolutely makes sense to have a system in place that allows the code generated by an LLM to be automatically validated but there’s no need to resort to a non-deterministic system for these sort of deterministic pass/fail conditions. [1] https://code.claude.com/docs/en/hooks https://code.claude.com/docs/en/hooks
- jarrettcoggin 4mo agoI was thinking the exact same thing. There are multiple places to implement hooks (git hooks, Claude hooks, etc.). One thing I've been wondering about is how to reliably protect specific portions of the system from unexpected/unnecessary change (for example, a failing test that Claude decides to comment out or rewrite to get it to pass). My only thought for this was to automatically revert test changes during specific portions of the implementation, but that feels overly rigid and potentially prevents things like refactoring code.
- cadamsdotcom 4mo agoEveryone looking into this and other verification should be moving away from long prompts and complex skills, and looking into hooks. If you put all these checks in your stop hook and your git commit hook, your repo docs can tell your agent that checks will run automatically when it stops work, and it should fix any problems found. It’s wonderful to reintroduce determinism at the QA end of your process. I find it very calming to know the agent can’t skip or forget to check its work because with hooks the checks are run by the harness.
- manmal 4mo agoI think pi-subagents (which can form arbitrarily long chains of subagents, with up to 8 in parallel) and Claude Code‘s new workflows feature, are quite convenient abstractions that can be setup quickly.
- jasonlotito 4mo ago100% agree. None of this who watches the watchmen thing. Force it as a pre-commit hook. The best part there is it means you don't have to hope other people have setup there agents in the same way. It just works, every time.
- xlii 4mo agoIf that's third then I have fourth. Self plug obviously, but figured that I'd like something between smart autocomplete and an agent - an autocomplete that has wider context. Called it rik, and it's on GitHub if anyone's interested checking it. https://github.com/exlee/rik https://github.com/exlee/rik
- EMM_386 4mo agoI always use a standard workflow and it has never been a problem. - Define the task and the goal, write a short spec document (markdown is fine) - Point the agent at it in plan mode and have it write the plan to disk with phases. Iterate on its plan if necessary here and now. - Have each agent tackle a phase and have it update it as a living document (switch models if some phases are more difficult than others) - Clear and repeat until done I've never had to overcomplicate this and it's worked both on enterprise-scale projects and personal projects. I am not sure what I'm missing - if anything.
- visarga 4mo agoI think what you are doing is good, I also have a similar workflow, but the idea here is to automate some of your manual approval work with coded tests. Since they are easy to generate, have as many as possible, think hard about what to test for, and the agent will deviate less and be more autonomous.
- apf6 4mo agoSlowing down development is the wrong goal. I see a desire for slowness come up a lot with developers. If you pursue that goal all the way to its logical conclusion then eventually you would stop all coding completely. Which would prevent new bugs but obviously we can't do that and keep our jobs. By all means add tons of quality gates to your SDLC pipeline. But thinking about slowness purely for the sake of slowness will not solve your problems.
- apsurd 4mo agoIf AI makes code a commodity, then why is the prescription for everyone to ship even more code even more urgently? My gut reaction, as a professional developer, to my (previous) company's AI mandates was an instinctive "wait but..." -- it didn't logic out to me. Now that I have much more AI experience under my belt, I understand the tension, it's a superpower and net-net ok so more features and more "stuff" will be built. But it's a very hard thing to balance. It's always been a bad idea for a company to position themselves as the one with more "stuff" in it.
- einpoklum 4mo agoIn other words: Spending more tokens is all you need. The main kind of pressure I'm feeling is the pressure of the giant AI, GPU & datacenter companies with their insane capital expenditure and circular deals, trying to get enough people to develop an expensive reliance on their service. And the more expensive, the better, so don't just pay for the LLM to code for you, have another LLM interact with the first LLM and pay double, treble, 5x or whatever. Then you can get the most refined slop.
- deleted 4mo ago[deleted]
- bilbo-b-baggins 4mo agoBro just rediscovered software best practices and thinks its a novel AI thing. Fuck, we’re so cooked.
- socketcluster 4mo agoI've been advocating for this approach for years. It's useful for any kind of data processing. You can't avoid race conditions without using some kind of queueing mechanism and you need backpressure to measure queue capacity. I built this into every aspect of https://socketcluster.io/ https://socketcluster.io/ - From pub/sub channels, RPCs to event listeners.
- jwpapi 4mo agoSuch a fantasy, it leads to two problems. Increased complexity of your systems. Increased pipelines of your system. You might reduce the likelihood of errors, but at an overproportinal cost of time it takes to complete (which some might argue is irrelevant, but has the cost of human context), and with an way higher time and focus needed for all bugs that the system doesnt work. You’ll have to fix adapt and maintain all your verification layers, because just because you set them up they are not perfect. Your testing pipeline becomes incredible slow and you need to maintain it as well. It’s tremendously weaker than a hands-on approach. I’ve written this exact same article in January and since then completely switched my position. Good luck on everyone trying this. You shuffling your own grave and waste time.
- hsaliak 4mo agoI have a custom agent that generates patches like you will with kernel development and I review and merge those in. https://github.com/hsaliak/std_slop/blob/main/docs/mail_mode.md https://github.com/hsaliak/std_slop/blob/main/docs/mail_mode... My agent forces this workflow by disabling modifications outside the coding step. I added looping to this not too long ago. https://github.com/hsaliak/std_slop/blob/main/docs/mail-loop/SKILL.md https://github.com/hsaliak/std_slop/blob/main/docs/mail-loop... This gives me the best of both worlds, hand curated reviews and automation. I often get the best quality if I do both, with an agent doing a pass first.
- tulga 4mo ago[flagged]
- jt_park 4mo ago[flagged]
- jasonlotito 4mo agoI feel like a lot of people just forget you can put this stuff in pre-commit hooks. This forces the AI to deal with issues. You don't have to hope and pray it remembers your "Pretty please, check your work" markdown file. A pre-commit hook has been wonderful. Sure, you can add instructions, but pre-commit hooks are where you want to put the guards.
- try-working 4mo agoI built a recursive workflow that creates its own source of truth for verifying its work: https://github.com/try-works/recursive-mode https://github.com/try-works/recursive-mode
- slow_typist 4mo agoWho is going to write tests? But I like the fact that this approach implicitly approves of the stochastic parrot model. I mean, given enough computing power and sufficiently well made tests, I could just generate random strings of increasing length until one compiles into a program that passes all tests, mission accomplished. Like one million apes typing on one million typewriters.
- acamerer 4mo ago[flagged]
- corner_booth_88 4mo ago[flagged]
- _zendar_ 4mo ago[dead]
- lofaszvanitt 4mo agoLipstick on a pig.
- tim-projects 4mo agoI'm building a tool that automates most of this. What the author didn't even touch on is just how much AI cheats. The more guardrails you provide the more it cheats. AI is like a wild animal that needs to do something, and it takes a fair bit of work to corner it. And only when it's cornered and at the point of giving up, can you then offer it a way out. If you don't do what I said, I can guarantee it's fooling you somehow.
- eugeneonai 4mo ago[flagged]
- ahstilde 4mo agomy entire thesis behind oro, my personal coding harness: https://github.com/mraakashshah/oro https://github.com/mraakashshah/oro
- mcint 4mo agoThe overriding of click behavior is quite annoying. 30 years of browser user-agent behavior. Next, Vercel, already handle this correctly. It takes special effort to violate "least surprise" here. Cmd-click on a link, should open it in a new tab. It does appear to be an issue with SimpleAnalytics, now Adobe's, onclick="saAutomatedLink(this, 'outbound'); return false;" Free debugging of how the site tweaks, breaks, the 30 year consensus web standard behavior. Good sites, good blogs, *don't override onclick for links.* Or handle it correctly. I'll leave an issue on the github. Between your footer, and dotfiles repo, OP does seem to appreciate standards & norms, in principle.
- visha1v 4mo ago[flagged]