6 ms·
Dynamic Workflows in Claude Code
- mil22 4mo agoInteresting to note, not sure if this was known publicly before today's blog post: Rewriting Bun with dynamic workflows An example of what dynamic workflows can unlock at scale is the recent rewrite of Bun. Jarred Sumner used dynamic workflows to port Bun from Zig to Rust with 99.8% of the existing test suite passing, roughly 750,000 lines of Rust, and eleven days from first commit to merge. One workflow mapped the right Rust lifetime for every struct field in the Zig codebase. The next wrote every .rs file as a behavior-identical port of its .zig counterpart, hundreds of agents working in parallel with two reviewers on each file. A fix loop then drove the build and test suite until both ran clean. After the port landed, an overnight workflow addressed unnecessary data copies and opened a PR for each for final review. While not yet in production, all of this was handled by dynamic workflows. Jarred will be writing about this more in the future.
- SkyPuncher 4mo agoI'm extremely skeptical that dynamic workflows had anything to do with this. I've been able to refactor one of the most complicated parts of our code base with similar results. Mechanical refactors are relatively straight forward for agents.
- jeswin 4mo ago> I've been able to refactor one of the most complicated parts of our code base with similar results. Mechanical refactors are relatively straight forward for agents. A rewrite of bun in Rust is unlikely to be a trivial mechanical refactor. And if you are not sharing what the complicated parts were, or how big it is, how do we assess that the task was similar? Unless you are intimately familiar with the bun codebase and you've already made that assessment.
- tra3 4mo agoI say this as someone who's found LLMs incredibly beneficial. Is this a way to increase token burn? I thought we covered this with Claude's C compiler. What changed?
- mattas 4mo agoMy initial reaction was that this is tokenmaxxing disguised as a product.
- Deukhoofd 4mo agoI'm going to be honest, this very much reads like an exciting new way to burn up as many tokens as possible. Large amounts of parallel agents that all have all their work double-checked by multiple other agents, and that keeps running for a longer period of time? I feel like there are more efficient ways to tackle the issues given.
- ithkuil 4mo agoPossibly. But otoh one cannot complain that agents don't produce high quality code while at the same time not allowing them to thoroughly go through all the steps required to produce high quality code
- SilverElfin 4mo agoCloudflare just launched a feature with this same name, just this month. Why would Anthropic choose the same exact name? https://blog.cloudflare.com/dynamic-workflows/ https://blog.cloudflare.com/dynamic-workflows/ Also isn’t all of this already easy to do on any of the platforms (include Claude before this and OpenAI too).
- deleted 4mo ago[deleted]
- CuriouslyC 4mo agoAnthropic is going to price themselves out of code, but still find a nice market providing service to senior management. Their long term play is virtual employees rather than tools for humans.
- trjordan 4mo agoIt feels like we're far past the point of where having AI do more faster is helpful. It's telling that they used "rewrite Bun in Rust" as the proof point here. It's cool! But the vast majority of software engineering doesn't start with tens of thousands of tests, where making them pass is the whole job. In my experience, AI still drifts from what I meant it to do on anything bigger than building a widget. My time is spent suspiciously reviewing output for changes the agent snuck in, or invariants it broke. I talked with a friend recently where the agent broke the test harness badly enough that none of the tests mattered for 3 weeks. They did pass, though, so CI never complained. There's something at the intersection of context engineering, managing that sloppy pile of markdown plans, and good old fashioning system understanding that's the real bottleneck.
- kian 4mo ago"In my experience, AI still drifts from what I meant it to do on anything bigger than building a widget." I've had code bases with tens of thousands of lines of code built from scratch that I hand-reviewed every line of and worked with the AI to improve, and haven't had this issue. I feel like a significant part of this is due to an involved /plan stage -- going back and forth on building out a plan for what you want the AI to do involves surfacing the assumptions that you would have called drift if you asked them to implement it directly from your prompt. Once the plan has been refined and is what I want it to be, getting it to implement everything in TDD style has for the most part given me 100% working code, as I wanted it to be, without issues. It definitely helps that I'm a principal-level engineer with extensive architectural experience -- but if you're able to tell the AI in detail what you want, have it ask questions for clarifications, and read through a plan before getting it implemented, and have a solid testing plus manual qa process (automated by chrome devtools mcp) in place, I've find that you can one-shot complex features, rewrites, and even not-insignificant applications that would have taken days to write by hand in a few hours.
- jeremyjh 4mo agoThere are certainly domains where AI is not so effective, but at this point I would agree that at least in terms of web development if you can't get effective results from agents at this point it is a skill issue. That skill can be learned, if you recognize that learning is part of the solution. I do think prior experience in product design, specifications & business analysis as well as engineering leadership are all extremely helpful. Its about putting the agent in a box so small that it really can't screw up; but its also about being able to review design and code rigorously - to see around corners and anticipate possible weaknesses etc. There is really nothing I have to do when working with an agent that I haven't already been doing for decades but it seems to me that a lot of developers have never found a single bug while reviewing someone else's code.
- piyuv 4mo ago“We realized the tech is not as addictive as we’ve hoped so we won’t be able to raise token prices enough to be profitable, so here’s a way to make you consume a lot more tokens without even realizing”
- bcherny 4mo agoA few of us from the Claude Code team will be hanging around if anyone has questions! Very excited for this launch -- dynamic workflows have been a game changer for engineering here at Anthropic. Can't wait to hear what you think.
- thallavajhula 4mo agoHi Boris! Thanks for Claude Code. Is there an example of how y'all use Dynamic Workflows internally that you could share with the rest of us here so that we can mimic something similar?
- bcherny 4mo agoHey, yep. A few things I personally used dynamic workflows for over the last few weeks: 1. Autonomously landed 20+ optimizations to reduce Claude Code's token usage by ~15% 2. Ported tree-sitter, color-diff, yoga-layout, and a number of other WASM and Rust native modules to TypeScript, improving CPU and memory use by 2-10x in the process 3. Made our CI faster, and repeatedly found and fixed flaky tests (with /loop) 4. Migrated from regex-based bash static analysis to tree-sitter, reducing false positive permission prompts by 45% 5. Reduced Claude Agent SDK startup time by 61%, by repeatedly profiling and optimizing the startup path, putting up a number of PRs in the process 6. Shipped 69 code simplification PRs, deleting >10k lines of code
- rahkiin 4mo agoYou _reduced_ its _efficiency_? Why do you make CC more inefficient?
- bcherny 4mo agoTypo! Edited
- isoprophlex 4mo agoMaxxing everything is all the rage. Gotta cpumaxx or bossman isnt getting his money's worth
- SkyPuncher 4mo agoI don't really get this. At this point, my limiting factor is not how quickly Claude can self-trudge through code. It's whether Claude is going to do the task correctly or not. I need more mechanisms for controlling long-running sessions and dynamically injecting my thoughts, correction, and nudges rather than faster ways to burn through my tokens without knowing if the results are going to be correct.
- jascha_eng 4mo agoyes I agree with this, more granular going back, letting me interrupt where it went off the rails, or even editing file reads myself etc would be lovely. Ingesting parts of other conversations would also be cool!
- dude250711 4mo agoI have heard of "token-maxxing" but I have not heard of "correctness-maxxing" or "quality-maxxing".
- mirashii 4mo agoNot with those exact terms, but it is certainly being discussed. Wes McKinney said in a recent talk that with current coding agents there’s no longer an excuse for shipping suboptimal code that takes on tech debt. Writing tests has never been cheaper, writing custom fuzzers, linters, and other harnesses that serve as guardrails has never been cheaper. His take is that “we didn’t have enough engineering time to do it right” is no longer an excuse, and the only excuses left are that you don’t know any better or you have bad taste.
- wrs 4mo agoI think the theoretical answer here is this: "Agents address the problem from independent angles, other agents try to refute what they found, and the run keeps iterating until the answers converge." So you will be supplying the "ground truth" (test suite, detailed spec, whatever) and empower an agent to use it to guide the other agents. Currently a lot of people do this sequentially in the form of multiple code-review passes by fresh agent sessions looking at the work of previous sessions. Adversarial models are a longstanding technique in ML so it makes sense they would try to go this way.
- vld_chk 4mo agoQuite a thing to use Bun rewrite to Rust as example of dynamic workflows, while now it is considered as anti pattern which leads team to stop supporting the tool due to inability to properly understand and navigate 1m vibe coded Rust lines
- buryat 4mo agoNot sure I understand how it's different from a team of sub-agents, what's the difference I'm curious?
- bcherny 4mo agoThere's two main differences: 1. Support for 1-2 OOMs more agents, to do more work in parallel 2. A phased, semi-structured approach where work happens in steps
- vblanco 4mo agoI made my own knockoff of that for myself https://github.com/vblanco20-1/AgentLoom https://github.com/vblanco20-1/AgentLoom (not really usable, just a vibecoded prototype), based on the workflow files found in the Bun repo. Ive been using it but pointed at deepseek flash to do some really large scale stuff. Its a fun way of using agents, and highly useful for tasks like code review to apply some rules, or to find vulnerability candidates. Funny enough, i used it in the same way claude does, vibecoding the workflow scripts and prompts themselves. I did find it uses tokens like crazy, i migrated Pixel Dungeon (java) to C# as a experiment, and it used almost 2 billion tokens. It was just 20 bucks due to deepseek flash, but i shudder thinking of how much money this uses when run on the real claude API pricing.
- jorgeleo 4mo agocurios minds... why to do that port?
- vblanco 4mo agojust to test the tech. No real usage other than for the fun of it. I did port stb_image from C to Jai which i was able to fully verify and harden and that one ill give more use. Im also using the same workflow system to perform agentic translation of a game i work with from english to various other languages, the results are far better than the commercial "human" translation services we tested. And i also use it to fix OCR issues on PDF books im ocr-ing for a data pipeline. This kind of workflow/wide agent swarm system is rather useful for many things where you want to "apply" the same prompts across a whole codebase or just in parallel.
- mkw5053 4mo agoWow, almost like the good old days of /ultrathink are back. Feels simultaneously like just yesterday and a lifetime ago.
- 2001zhaozhao 4mo agoWe really need a way to scope and implement these multi-agent orchestration features that isn't locked in to one provider.
- xcskier56 4mo agoAre these “features” just hooks to get people to burn more tokens faster? I’m at the point where deciding what we should and should not do takes a lot more time than actually doing it. More agents just means running faster in potentially the wrong direction
- cush 4mo agoThey’re pure enterprise features - needed for massive legacy codebases with tens or hundreds of similar enough coding tasks - where there is a lived “cost” of not doing this type of work paid by every engineer working around it
- jdw64 4mo ago[dead]
- isoprophlex 4mo agoThis seems like it's an openclaw, anthropic edition. Something like ClaudeClaw?
- zli0823 4mo agoa completely new way to burn your money.
- zli0823 4mo agofound a new way to burn your money quicker.
- brap 4mo ago>Claude dynamically writes orchestration scripts So, is this like a skill the LLM should follow, or an actual "workflow" in the deterministic sense? If it's the former, is it even reliable for long running tasks? If it's the latter, can users interact with it?
- afro88 4mo agoIt's the later. You can view it and see fine grained progress, but you can't interact with it. I hope that's coming next, because it would be useful to steer later phases or even agents
- aabdi 4mo agowrote something similar for my own use/work stuff; seems everything is converging towards similar ideas. IMO, this style of workflow/agentics is how all SWE'll look like long term. Automate everything into a big pipe-y thing. How it's gonna be modelled is up in the air though. lots of different approaches: mine: https://github.com/portpowered/you-agent-factory https://github.com/portpowered/you-agent-factory https://github.com/ComposioHQ/agent-orchestrator https://github.com/ComposioHQ/agent-orchestrator https://github.com/gastownhall/gastown https://github.com/gastownhall/gastown https://github.com/openai/symphony https://github.com/openai/symphony
- afro88 4mo agoI tried this out yesterday - lucky enough to have access through EAP at work. The workflows that are generated are quite good - smart parallelisation and phasing. End results for larger chunks of work are also much better, which I attribute to more of the work having clean context windows (Opus 4.7 is unusable past 200k conversation length, and each subagent ends up using less than that IME). They also seem to have a validation phase hint in the workflow generator which also helps a lot. Speed is a bonus. You can achieve a similar result manually prompting to use subagents, yes. But the TUI for in flight dynamic workflows is really nice - great visibility into exactly what's happening. Honesty, for anything larger than a 1 shot PR, it's worth firing off a workflow for better automatic context management alone (more work done in the first 20% sweet spot)
- ethanguo 4mo ago[dead]
- mohsen1 4mo agoI’m gonna try this one on tsz. So far Codex /goal has been great https://tsz.dev https://tsz.dev So far Codex /goal has been amazing but Claude Code /goal or even /loop does not work hard enough and gives up. I have observed it just claiming it’s “iterating” in a broken loop or simply giving up.
- vb-8448 4mo ago> Rewriting Bun with dynamic workflows Are we sure this is a good "success story" example?
- AndyNemmity 4mo agoI have had dynamic workflows in my agent for the past 9 months. I am diffing Claude Code with them, I tend to agree with the analysis. So far, versus my system, there are tradeoffs, but the dynamic workflows are over tuned to use way more agents that I have ever found add value. It used 8 to diff our systems. I would have used 4, for example.
- chandureddyvari 4mo agoI’m currently cobbling sub agents with hooks, workflows looks very promising for doing things more predictably. Is this equivalent of DAGs for sub agents inside claude code? Can i pause and resume/retry workflows? How stateful are they? Really appreciate it someone claude code can throw more light on above. I’m trying to see if I can get langgraph equivalent DAGs here.
- KaiShips 4mo ago[flagged]
- ajma 4mo ago> It’s important to note that dynamic workflows consume meaningfully more usage than a typical Claude Code session
- ncphillips 4mo agoI just hit my Claude Max limit for the first time _ever_ thanks to workflows lol Like 90 agents ran to do a code review of a fairly small package I have. They're really looking for us to increase token usage aren't they?
- tomjakubowski 4mo agoThis is a fundamental incentive issue with any company that does all of training models, building harnesses for them, and offering them as a service.
- deleted 4mo ago[deleted]
- Ozzie-D 4mo ago[flagged]
- Robdel12 4mo ago> Rewriting Bun with dynamic workflows There ya go, the rewrite was for marketing.
- seabass 4mo agoHave to love their demo use case: React -> Solid migration
- _pdp_ 4mo ago[flagged]
- dools 4mo agoThe #1 goal for Anthropic and others is to take the longest running process possible and make it entirely opaque to the developer. It's the only way they can build a moat for a commodity. I would highly recommend building your own multi-stage orchestration flows because then you'll get a much better idea of where you need to be in the loop, and where you can save money. Once entire organisations are functioning only as extensions of Anthropic, they'll put the prices up and squeeze the shit out of the market.
- deleted 4mo ago[deleted]
- nebben64 4mo agoHow is this different, or how does it complement Agent Teams? When should I use which?
- Zopieux 4mo agoWho, beyond Anthropic themselves, can afford such purposefully wasteful uses of LLMs? "the model sucks a bit so we just have best-of-4 & adversarial reviewing agents; surely one more agent will do the trick"
- facundo_olano 4mo agoAbsolutely annoying to have it assuming that I want to use this when I type workflow in the prompt. Like thats not already a thing in half of the software projects
- bickov 4mo ago[dead]
- sermakarevich 4mo agoNot sure why Claude does not have AskUserQuestion implementation that works for spawned sessions: subagents, teams, workflows. Without it, spawning hundreds of subagents and wait for final result without single input feels a bit risky. Here is the solution to it. Built on a SQLite DB and MCP, blocking until the question is answered, supporting all possible question types, with a CLI or web interface for answers, `ask_human_question` fills the gap in efficient subagent management. https://news.ycombinator.com/item?id=48320233 https://news.ycombinator.com/item?id=48320233
- willXare 4mo ago[flagged]
- kcarriedo 4mo ago[flagged]
- inerte 4mo agoI am getting so confused when to use what... agents, sub-agents, tasks, team mates, /goal, /loop, and now workflow. Each with different degrees of effort. Don't make me think. All these knobs are also exposed in ChatGPT, which I am more familiar when chatting. Which one of the models? Do I go Instant, Thinking, Pro? Extended Pro? Oh no, maybe I need Deep Research. Sometimes I think it's on purpose. I fear if I try a lowest knob, it will miss something. So turn everything up. And token usage goes up.
- notatoad 4mo agoi think it is on purpose, but not for the cynical reason of burning more tokens. developers like knobs, it helps us feel like we're in control. even when we're not. so the ai companies give us knobs and buttons and sliders to make us more comfortable.
- esafak 4mo agoThat's what product management is for.
- dbbk 4mo agoClaude doesn't have product management. I don't think they even have QA. There are such glaring UX issues that have persisted for months that I genuinely don't think they use their own product - there is no other explanation.
- deleted 4mo ago[deleted]
- dynamicflow 4mo agoPremium domain for sale: dynamicflow.com Perfect for AI workflow tooling, automation SaaS, or orchestration platforms. Short, brandable, .com. DM if interested.
- Syntaf 4mo agoTested this out on a 5x max plan, turns out I spun up 62 Opus 4.8 1M sub-agents for my dynamic workflow and maxed out my ~5hr cap in..... 18 minutes? Oops, but probably good to know that this is not a cheap feature -- next time I'll have to figure out how to tune the workflows to use Haiku / Sonnet
- willyv3 4mo ago[flagged]
- kcarriedo 4mo ago[flagged]
- kcarriedo 4mo ago[flagged]