13 ms·
Orchestrate teams of Claude Code sessions
- Sol- 8mo agoWith stuff like this, might be that all the infra build-out is insufficient. Inference demand will go up like crazy.
- kylehotchkiss 8mo agoIt'd be nice if CC could figure out all the required permissions upfront and then let you queue the job to run overnight
- LtWorf 8mo agoExcept it cannot really do anything unattended
- MantraHQ 8mo ago[dead]
- intellegix 8mo agoIt actually can with the right wrapper. I built an open source loop driver that runs Claude Code CLI autonomously with --dangerously-skip-permissions. It handles session continuity (--resume), budget enforcement, stagnation detection (two-strike system if turns stay low), and auto model fallback (Opus -> Sonnet on consecutive timeouts). The key is streaming NDJSON output to track cost per iteration and detect completion markers. The human stays in control by editing CLAUDE.md between runs to steer the project. https://github.com/intellegix/intellegix-code-agent-toolkit https://github.com/intellegix/intellegix-code-agent-toolkit
- Der_Einzige 8mo agoAnyone paying attention has known that demand for all type of compute than can run LLMs (i.e. GPUs, TPUs, hell even CPUs) was about to blow up, and will remain extremely large for years to come. It's just HN that's full of "I hate AI" or wrong contrarian types who refuse to acknowledge this. They will fail to reap what they didn't sow and will starve in this brave new world.
- emp17344 8mo agoThis reads like a weird cult-ish revenge fantasy.
- RGamma 8mo agoAnd what about you? Show your "I used AI today" badge, right now!
- mrkeen 8mo agoOh yeah I mean if you're a webdev and you haven't built several data centres already you're basically asking to be homeless.
- ffffuuuuuccck 8mo ago[flagged]
- aaaalone 8mo agoIf ai progresses slow enough, we will end in a society were high unemployment numbers are the norm and we are stuck in capitalism. And if I think about one 'senior' in my team I would pref an expensive ai subscription over that one person already.
- Der_Einzige 8mo ago[flagged]
- emp17344 8mo agoWhat the fuck is wrong with you? This guy is either a troll or legitimately mentally ill.
- sciencejerk 8mo agoBlue collar work won't be safe for long. Just longer.
- anthem2025 8mo ago
- RGamma 8mo agoUnlocking the next order of magnitude of software inefficiency! Though I do hope the generated code will end up being better than what we have right now. It mustn't get much worse. Can't afford all that RAM.
- Sol- 8mo agoDunno, it's probably less energy efficient than a human brain, but being able to turn electricity into intelligence is pretty amazing. RAM and power generation are engineering problems to be solved for civilization to benefit from this.
- bhasi 8mo agoSeems similar to Gas Town
- nickorlow 8mo agoyeah, seems like a much simpler design though (i.e. only seems like one 'special/leader' agent, and the rest are all workers vs gastown having something like 8 different roles mayor, polecat, witnesses, etc). Wonder how they compare?
- greenfish6 8mo agoi would have to imagine the gastown design isn't optimal though? why 8, and why does there need to multiple hops of agent communications before two arbitrary agents communicate with each other as opposed to single shared filespace?
- Ethee 8mo agoI've been using Gas Town a decent bit since it was released. I'd agree with you that it's design is sub-optimal, but I believe that's more due to the way the actual agents/harnesses have been designed as opposed to optimal software design. The problem you often run into is that agents will sometimes hang thinking they need human input for a problem they are on, or they think they're at a natural stopping point. If you're trying to do fully orchestrated agentic coding where you don't look at the code at all (putting aside whether that's good or not for a second) then this is sub-optimal behavior, and so these extra roles have been designed to 'keep the machine going' as it were. Often times if I'm only working on a single project or focus, then I'm not using most of those roles at all and it's as you describe, one agent divvying out tasks to other agents and compiling reports about them. But due to the fact that my velocity with this type of coding is now based on how fast I can tell that agent what I want, I'm often working on 3 or 4 projects simultaneously, and Gas Town provides the perfect orchestration framework for doing this.
- cstejerean 8mo ago
- taikahessu 8mo agoClean up the team
- Retr0id 8mo agoClaude Town
- IhateAI 8mo ago[flagged]
- theappsecguy 8mo agoThe crash and burn can't come soon enough.
- tjr 8mo agoPeople often compare working with AI agents to being something like a project manager. I've been a project manager for years. I still work on some code myself, but most of it is done by the rest of the team. On one hand, I have more bandwidth to think about how the overall application is serving the users, how the various pieces of the application fit together, overall consistency, etc. I think this is a useful role. On the other hand, I definitely have felt mental atrophy from not working in the code. I still think; I still do things and write things and make decisions. But I feel mentally out of shape; I lack a certain sharpness that I perceived when I was more directly in tune with the code. And I'm talking, all orthogonal to AI. This is just me as a project manager with other humans on the project. I think there is truth to, well, operate at a higher level! Be more systems-minded, architecture-minded, etc. I think that's true. And there are surely interesting new problems to solve if we can work not on the level of writing programs, but wielding tools that write programs for us. But I think there's also truth to the risk of losing something by giving up coding. Whether if that which might be lost is important to you or not, is your own decision, but I think the risk is real.
- IhateAI 8mo agoI definitely think what you're losing is extremely important, and can't be compensated with LLMs once its gone. Back when automatic piano players came out, if all the world's best piano players stopped playing and mostly just composing/writing music instead, would the quality of the music have increased or decreased. I think the latter.
- sathish316 8mo agoI do think there’s a real risk of Brain Atrophy when you rely on AI coding tools for everything and while learning something new. About a year ago, I dealt with this problem by using Neovim and having shortcuts like below to easily toggle GitHub Copilot on/off. Now that AI is baked into almost every part of the toolchain in VSCode, Cursor, ClaudeCode, Intellij, I don't know how the newer engineers will learn without AI assistance.
- greenfish6 8mo agoExcited to try this out. I've seen a lot of working systems on my own computer that share files to talk between different Claude Code agents and I think this could work similarly to that. (i thought gas town was satire? people in comments here seem to be saying that gas town also had multi-agent file sharing for work tracking)
- nkmnz 8mo agoI’m looking for something like this, with opus in the driver seat, but the subagents should be using different LLMs, such as Gemini or Codex. Anyone know if such a tool? just-every/code almost does this, but the lead/orchestrator is always codex, which feels too slow compared to opus or Gemini.
- fosterfriends 8mo agoI think this is where future cursor features will be great - to coordinate across many different model providers depending on the sub-jobs to be done
- nkmnz 8mo agoWhat I want is something else: I want them to work in parallel on the same problem, and the orchestrator to then evaluate and consolidate their responses. I’m currently doing this manually, but it’s tedious.
- sathish316 8mo agoYou can run an ensemble of LLMs (Opus, Gemini, Codex) in Claude Code Router via OpenRouter or any Agent CLI that supports Subagents and not tied to a single LLM like Opencode. I have an example of this in Pied-Piper, a subagent orchestrator that runs in Claude Code or ClaudeCodeRouter and uses distinct model/roles for each Subagent: 1. GPT-5.2 Codex Max for planning 2. Opus 4.5 for implementation 3. Gemini for reviews It’s easy to swap models or change responsibilities. Doc and steps here: https://github.com/sathish316/pied-piper/blob/main/docs/playbook/PLAYBOOK_DREAM_TEAM_ENSEMBLE_MODELS.md https://github.com/sathish316/pied-piper/blob/main/docs/play...
- knes 8mo agoAt Augment' we've been working on this. Multi agents orchestration, spec driven, different models for different tasks, etc. https://www.augmentcode.com/product/intent https://www.augmentcode.com/product/intent can use the code AUGGIE to skip the queue. Bring your own agent (powered by codex, CC, etc) coming to it next week.
- 8mo ago
- morleytj 8mo agoGas Town decimated by Claude bomb from orbit
- greenfish6 8mo agosomething i really like from tryin git out over the last 10 minutes is that the main agent will continue talking to you while other agents are working, so you don't have to queue a message
- ottah 8mo agoI absolutely cannot trust Claude code to independently work on large tasks. Maybe other people work on software that's not significantly complex, but for me to maintain code quality I need to guide more of the design process. Teams of agents just sounds like adding a lot more review and refactoring that can just be avoided by going slower and thinking carefully about the problem.
- BonoboIO 8mo agoYou definitely have to create some sort of PLAN.md and PROGRESS.md via a command and an implement command that delegates work. That is the only way that I can get bigger things done no matter how „good“ their task feature is. You run out of context so quickly and if you don’t have some kind of persistent guidance things go south
- koakuma-chan 8mo agoI tried doing that and it didn't work. It still adds "fallbacks" that just hide errors or the fact that there is no actual implementation and "In a real app, we would do X, just return null for now"
- ottah 8mo agoIt's not sufficient, especially if I am not learning about the problem by being part of the implementation process. The models are still very weak reasoners, writing code faster doesn't accelerate my understanding of the code the model wrote. Even with clear specs I am constantly fighting with it duplicating methods, writing ineffective tests, or implementing unnecessarily complex solutions. AI just isn't a better engineer than me, and that makes it a weak development partner.
- vonneumannstan 8mo ago>AI just isn't a better engineer than me, and that makes it a weak development partner. This would also be true of Junior Engineers. Do you find them impossible to work with as well?
- 8mo ago
- ndesaulniers 8mo agoSubagents are out, put it all on agent teams!
- pronik 8mo agoTo the folks comparing this to GasTown: keep in mind that Steve Yegge explicitely pitched agent orchestrators to among others Anthropic months ago: > I went to senior folks at companies like Temporal and Anthropic, telling them they should build an agent orchestrator, that Claude Code is just a building block, and it’s going to be all about AI workflows and “Kubernetes for agents”. I went up onstage at multiple events and described my vision for the orchestrator. I went everywhere, to everyone. (from "Welcome to Gas Town" https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16dd04 https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16d...) That Anthropic releases Agent Teams now (as rumored a couple of weeks back), after they've already adopted a tiny bit of beads in form of Tasks) means that either they've been building them already back when Steve pitched orchestrators or they've decided that he's been right and it's time to scale the agents. Or they've arrived at the same conclusions independently -- it won't matter in the larger scale of things. I think Steve greately appreciates it existing; if anything, this is a validation of his vision. We'll probably be herding polecats in a couple of months officially.
- isoprophlex 8mo agoThere seems to be a lot of convergent evolution happening in the space. Days before the gas town hype hit, I made a (less baroque, less manic) "agent team" setup: a shell script to kick off a ralph wiggum loop, and CLAUDE-MESSAGE-BUS.md for inter-ralph communication (Thread safety was hacked into this with a .claude.lock file). The main claude instance is instructed to launch as many ralph loops as it wants, in screen sessions. It is told to sleep for a certain amount of time to periodically keep track of their progress. It worked reasonably well, but I don't prefer this way of working... yet. Right now I can't write spec (or meta-spec) files quick enough to saturate the agent loops, and I can't QA their output well enough... mostly a me thing, i guess?
- pronik 8mo ago> Right now I can't write spec (or meta-spec) files quick enough to saturate the agent loops, and I can't QA their output well enough... mostly a me thing, i guess? Same for me, however, the velocity of the whole field is astonishing and things change as we get used to them. We are not talking that much about hallucinating anymore, just 4-5 months ago you couldn't trust coding agents with extracting functionality to a separate file without typos, now splitting Git commits works almost without a hinch. The more we get used to agents getting certain things right 100% of the time, the more we'll trust them. There are many many things that I know I won't get right, but I'm absolutely sure my agent will. As soon as we start trusting e.g. a QA agent to do his job, our "project management" velocity will increase too. Interestingly enough, the infamous "bowling score card" text on how XP works, has demonstrated inherently agentic behaviour in more way than one (they just didn't know what "extreme" was back then). You were supposed to implement a failing test and then implement just enough functionality for this test to not fail anymore, even if the intended functionality was broader -- which is exactly what agents reliably do in a loop. Also, you were supposed to be pair-driving a single machine, which has been incomprehensible to me for almost decades -- after all, every person has their own shortcuts, hardware, IDEs, window managers and what not. Turns out, all you need is a centralized server running a "team manager agent" and multiple developers talking to him to craft software fast (see tmux requirement in Gas Town).
- GoatOfAplomb 8mo agoI wonder if my $20/mo subscription will last 10 minutes.
- simlevesque 8mo agoI've had good results with Haiku for certain tasks.
- tclancy 8mo agoAh ok, same. I keep wondering about how this would ever accomplish anything.
- mohsen1 8mo agoAt this point, if you're paying out of pocket you should use Kimi or GLM for it to make sense
- bluerooibos 8mo agoThese are super slow to run locally, though, unless you've got some great hardware - right? At least, my M1 Pro seems to struggle and take forever using them via Ollama.
- corysama 8mo agoTry this https://unsloth.ai/docs/models/qwen3-coder-next https://unsloth.ai/docs/models/qwen3-coder-next
- andai 8mo agoGLM is OK (haven't used it heavily but seems alright so far), a bit slow with ZAI's coding plan, amazingly fast on Cerebras but their coding plan is sold out. Haven't tried Kimi, hear good things.
- asdev 8mo agoI personally have no use for this type of workflow. I like parallel claude code instances in worktrees but nothing beyond that
- hpdigidrifter 8mo agoAm not a fan of dealing with worktrees Maybe for larger longer lived tasks but the time spent on merges from different agents is definitely a big headwind for parallel work. This seems handled by this new agent which is cool. I gave up on worktrees and hacked together a solution with fine-grained lockfiles for editing, running builds, etc that worked surprisingly good for what it was
- avereveard 8mo ago"finish Claude tokens quota in 3 minutes, largely over delegation and result messages instead of code writing"
- giancarlostoro 8mo agoI was working on my own alternative to Beads... then I realized I could do exactly this with something similar to Beads, I'm planning on open sourcing it soon because I like what I have so far, I also made it so I can sync my tasks directly to my GitHub projects as well. I think its more useful to have agent tasks eventually synched back up to real ticketing systems for historical reasons. Besides, its better to have alternatives that are agent agnostic.
- mcintyre1994 8mo agoI’ve been mostly holding off on learning any of the tools that do this because it seemed so obvious that it’ll be built natively. Will definitely give this a go at some point!
- khaliqgant 8mo agoBeen waiting for this to drop and excited to test it out. We've been building something in this space - https://github.com/AgentWorkforce/relay https://github.com/AgentWorkforce/relay, a real-time messaging layer that lets AI coding agents talk to each other across any CLI. Assign roles to different models and have them coordinate: Claude as the lead, Codex on backend, Gemini on frontend, etc. I wrote about my experiences with multi-agent orchestration here: https://x.com/khaliqgant/status/2019124627860050109?s=46 https://x.com/khaliqgant/status/2019124627860050109?s=46
- drbscl 8mo agoI just built a quick plugin to automatically add agents & skills then fire off a team with them, depending on your task: https://github.com/drbscl/dream-team https://github.com/drbscl/dream-team
- d4rkp4ttern 8mo agoThis sounds very promising. Using multiple CC instances (or mix of CLI-agents) across tmux panes has always been a workflow of mine, where agents can use the tmux-cli [1] skill/tool to delegate/collaborate with others, or review/debug/validate each others work. This new orchestration feature makes it much more useful since they share a common task list and the main agent coordinates across them. [1] https://github.com/pchalasani/claude-code-tools?tab=readme-ov-file#-tmux-cli--terminal-automation https://github.com/pchalasani/claude-code-tools?tab=readme-o...
- vardalab 8mo agoYeah, I've been using your tools for a while. They've been nice.
- bluerooibos 8mo agoThis is great and all but, who can actually afford to let these agents run on tasks all day long? Is anyone here actually using this or are these rollouts aimed at large companies? I'm burning through so many tokens on Cursor that I've had to upgrade to Ultra recently - and i'm convinced they're tweaking the burn rate behind the scenes - usage allowance doesn't seem proportional. Thank god the open source/local LLM world isn't far behind.
- anupamchugh 8mo ago[flagged]
- simianwords 8mo agoNeed it be actually disjoint? Interested in learning about the limitation here because apparently the agents can coordinate. Otherwise what’s the difference between what they are providing vs me creating two independent pull requests using agents and having an agent resolve merge conflicts?
- anupamchugh 8mo ago[flagged]
- Aditya_Garg 8mo agoAnthropic themselves were able to write a c compiler using teams all at the same time https://www.anthropic.com/engineering/building-c-compiler https://www.anthropic.com/engineering/building-c-compiler Here is the relevant excerpt: "To prevent two agents from trying to solve the same problem at the same time, the harness uses a simple synchronization algorithm: Claude takes a "lock" on a task by writing a text file to current_tasks/ (e.g., one agent might lock current_tasks/parse_if_statement.txt, while another locks current_tasks/codegen_function_definition.txt). If two agents try to claim the same task, git's synchronization forces the second agent to pick a different one. Claude works on the task, then pulls from upstream, merges changes from other agents, pushes its changes, and removes the lock. Merge conflicts are frequent, but Claude is smart enough to figure that out."
- dangus 8mo agoA cynical read of this is that it’s all a ploy to maximize usage. Why do agents need to speak to each other if they’re just doing the work correctly the first time? Is it an admission that a single agent is not useful and reliable enough?
- WXLCKNO 8mo agoI run a loop where I have 4 agents review in parallel after each implementation phase. It just increases the odds of finding issues. I've switched this over to a team of 4 now that talk to each other to discuss issues they find and it's amazing. They confirm between themselves and if they wrongly identified something the others correct them.
- dangus 8mo agoSo, the answer is yes, a single agent makes too many mistakes and you have to run four of them (4x usage cost) to improve the quality. I understand that it works better, but I am rightfully pointing out that it's less efficient. An analogy would be putting a V8 engine into a pickup truck to make it go as fast as a Mazda Miata.
- imiric 8mo agoI find it amusing that the innovation in this space for the past year+ has been mostly centered around engineering: MCP, "agents", "skills", etc. Now "agent" orchestration is the new hotness. Meanwhile, the same issues that have plagued these tools since their inception are largely ignored: hallucination, innacuracy, context collapse, etc. These won't be solved by engineering, but by new research and foundational improvements. On one hand, solid engineering was sorely needed, and can extract a lot of value from the current tech. But on the other, all these announcements and improvements feel like companies grasping at straws to keep the hype cycle going by any means necessary. Charts must go up and to the right, or investors get antsy. It's all adding to the mountain of signs that suggest that this isn't the path to artificial intelligence. It's interesting tech, with possibly many valuable applications, but the "AI" narrative is frankly tiring. I wish I could fast forward on this speculative phase, go past the inevitable crash, and arrive at a timeframe where we've figured out what this tech is actually good for, and where we hopefully use it more for good than evil.
- traviscline 8mo agoBeen using these types of flows across agent harnesses for a while. Check out https://github.com/tmc/it2 https://github.com/tmc/it2
- rektlessness 8mo agoAre people using Claude max 20x plan for personal pet projects? Are these expensed? Have you liquidated all other hobbies to fund this? Asking for a friend.
- jFriedensreich 8mo agoWhile i appreciate anthropic making a proof of concept like they did with claude code cli on which they can then do RL to optimise the patterns that work, I expect this to be as unusable as the cli itself. Its a big difference if a model provider internalises something like thinking mode which mainly depends on context and text or if they try to grab a part of the agent loop which has to run on the side of the systems we build and use. We cannot allow model providers to own the browsers, CLIs, memory, IDEs, extensions and other tooling. Its not just a matter of power but also they just suck at it as i experience every time i have to use claude code instead of amp. I truly hope we get the pattern of innovation that looks like: - some dude vibecodes a really cool idea - model providers build into their reference implementations - model providers optimize models to work optimally - startup and/or open source projects step in and build something that is actually usable and opens a new market segment We saw this play out beautifully with amp, kilo, roo, cline, continue Another aspect is that we do not want interfaces just made for agents to work in teams, we want software made for humans and agents, that are true platforms for these agent teams to collaborate in.
- pipejosh 8mo ago[dead]