5 ms·
Show HN: Optio – Orchestrate AI coding agents in K8s to go from ticket to PR
I think like many of you, I've been jumping between many claude code/codex sessions at a time, managing multiple lines of work and worktrees in multiple repos. I wanted a way to easily manage multiple lines of work and reduce the amount of input I need to give, allowing the agents to remove me as a bottleneck from as much of the process as I can. So I built an orchestration tool for AI coding agents:
Optio is an open-source orchestration system that turns tickets into merged pull requests using AI coding agents. You point it at your repos, and it handles the full lifecycle:
- Intake — pull tasks from GitHub Issues, Linear, or create them manually
- Execution — spin up isolated K8s pods per repo, run Claude Code or Codex in git worktrees
- PR monitoring — watch CI checks, review status, and merge readiness every 30s
- Self-healing — auto-resume the agent on CI failures, merge conflicts, or reviewer change requests
- Completion — squash-merge the PR and close the linked issue
The key idea is the feedback loop. Optio doesn't just run an agent and walk away — when CI breaks, it feeds the failure back to the agent. When a reviewer requests changes, the comments become the agent's next prompt. It keeps going until the PR merges or you tell it to stop.
Built with Fastify, Next.js, BullMQ, and Drizzle on Postgres. Ships with a Helm chart for production deployment.
- MrDarcy 7mo agoLooks cool, congrats on the launch. Is there any sandbox isolation from the k8s platform layer? Wondering if this is suitable for multiple tenants or customers.
- jawiggins 7mo agoOh good question, I haven't thought deeply about this. Right now nothing special happens, so claude/codex can access their normal tools and make web calls. I suppose that also means they could figure out they're running in a k8s pod and do service discovery and start calling things. What kind of features would you be interested in seeing around this? Maybe a toggle to disable internet connections or other connections outside of the container?
- nevon 7mo agoNetwork policies controlling egress would be one thing. I haven't seen how you make secrets available to the agent, but I would imagine you would need to proxy calls through a mitm proxy to replace tokens with real secrets, or some other way to make sure the agent cannot access the secrets themselves. Specifically for an agent that works with code, I could imagine being able to run docker-in-docker will probably be requested at some point, which means you'll need gvisor or something.
- jordanedev 7mo agoThat's exactly what i did personnaly on my oss repo https://github.com/ysa-ai/ysa https://github.com/ysa-ai/ysa I want to run my agents fully isolated with headless mode. To achieve that safely you have to run a proxy
- navilai 6mo ago[dead]
- antihero 7mo agoAnd what stops it making total garbage that wrecks your codebase?
- upupupandaway 7mo agoTicket -> PR -> Deployment -> Incident
- zvqcMMV6Zcr 7mo ago> To make error is human. To propagate error to all server in automatic way is #devops I am not sure how AI agent variation of that joke would look like. Every now and then some blog posts lands on HN asking "Where are all new apps created thanks to LLM productivity boost"?. I am more surprised there are no news about some serious fuck-ups that can be traced back to LLM usage in code.
- jawiggins 7mo agoThere are a few things: a) you can create CI/build checks that run in github and the agents will make sure pass before it merges anything b) you can configure a review agent with any prompt you'd like to make sure any specific rules you have are followed c) you can disable all the auto-merge settings and review all the agent code yourself if you'd like.
- kristjansson 7mo ago> to make sure you've really got to be careful with absolute language like this in reference to LLMs. A review agent provides no guarantees whatsoever, just shifts the distribution of acceptable responses, hopefully in a direction the user prefers.
- jawiggins 7mo agoFair, it's something like a semantic enforcement rather than a hard one. I think current AI agents are good enough that if you tell it, "Review this PR and request changes anytime a user uses a variable name that is a color", it will do a pretty good job. But for complex things I can still see them falling short.
- rafaelbcs 7mo ago[dead]
- QubridAI 7mo ago[flagged]
- knollimar 7mo agoI don't want to accuse you of being an LLM but geez this sounds like satire
- weird-eye-issue 7mo agoIt's AI.
- conception 7mo agoWhat’s the most complicated, finished project you’ve done with this?
- jawiggins 7mo agoRecently I used to to finish up my re-implementation of curl/libcurl in rust (https://news.ycombinator.com/item?id=47490735 https://news.ycombinator.com/item?id=47490735). At first I started by trying to have a single claude code session run in an iterative loop, but eventually I found it was way to slow. I started tasking subagents for each remaining chunk of work, and then found I was really just repeating the need for a normal sprint tasking cycle but where subagents completed the tasks with the unit tests as exit criteria. So optio came to my mind, where I asked an agent to run the test suite, see what was failing, and make tickets for each group of remaining failures. Then I use optio to manage instances of agents working on and closing out each ticket.
- hmokiguess 7mo agothe misaligned columns in the claude made ASCII diagrams on the README really throw me off, why not fix them? | | | |
- jawiggins 7mo agoShould be fixed now :)
- hmokiguess 7mo agothank you x)
- denysvitali 7mo agoFWIW, a "cheaper" version of this is triggering Claude via GitHub Actions and `@claude`ing your agents like that. If you run your CI on Kubernets (ARC), it sounds pretty much the same
- naultic 7mo agoI'm working on something a little similar but mines more a dev tool vs process automation but I love where yours is headed. The biggest issue I've run into is handling retries with agents. My current solution is I have them set checkpoints so they can revert easily and when they can't make an edit or they can't get a test passing, they just restart from earlier state. Problem is this uses up lots of tokens on retries how did you handle this issue in your app?
- jawiggins 7mo agoGenerally I've found agents are capable of self correcting as long as they can bash up against a guardrail and see the errors. So in optio the agent is resumed and told to fix any CI failures or fix review feedback.
- abybaddi009 7mo agoDoes this support skills and MCP?
- jawiggins 7mo agoYup. MCP can be configured on a repo level. At task execution time, enabled MCP servers are written as a .mcp.json file into the agent's worktree. Enabled skills are written as .claude/commands/{name}.md files in the worktree, making them available as slash commands to the agent
- pianopatrick 7mo agoI wonder, based on your experience, how hard would it be to improve your system to have an AI agent review the software and suggest tickets? Like, can an AI agent use a browser, attempt to use the software, find bugs and create a ticket? Can an AI agent use a browser, try to use the software and suggest new features?
- smokeyfish 7mo agoDatadog have a feature like that.
- ramon156 7mo agoI think it's more important to pin down where a human must be in order for this not to become a mess. Or have we skipped that step entirely?
- pianopatrick 7mo agoPersonally my theory is that to solve the messiness we will need some new frameworks and even languages that are designed to catch AI mistakes in large code bases. For example, AIs in the past would sometimes hallucinate methods that do not exist. But in a language with a strong type system a static type checker should be able to catch that mistake and give the AI automated feedback to fix that mistake without a human in the loop. As far as humans in the loop, the only human we ultimately cannot get rid of is the user. But I think with a combo of user feedback forms and automated metrics we can give AI a lot of feedback about how good software is just from users using the software.
- mlsu 7mo agoperhaps we can give the AI a bit of money, make it the customer, then we can all safely get off the computer and go outside :)
- stingraycharles 7mo agoAI agents can absolutely use web browsers to do these things, but the hard part is accurately defining the acceptance criteria.
- verdverm 7mo agoI love k8s, but having it as a requirement for my agent setup is a non-starter. Kubernetes is one method for running, not the center piece.
- raised_hand 7mo agoWhy K6? Is there a way I could run it without
- stingraycharles 7mo agoI’ve come to the realization that these kind of systems don’t work, and that a human in the loop is crucial for task planning; the LLM’s role being to identify issues, communicate the design / architecture, etc before it’s handed off, otherwise the LLM always ends up doing not entirely the correct thing. How is this part tackled when all that you have is GH issues? Doesn’t this work only for the most trivial issues?
- mshark 7mo agoHad the same realization which inspired eforge (shameless plug) https://github.com/eforge-build/eforge https://github.com/eforge-build/eforge - planning stays in the developer’s control with all engineering (agent orchestration) handed off to eforge. This has been working well for a solo or siloed developer (me) that is free to plan independently. Allows the developer to confidently stay in the planning plane while eforge handles the rest using a methodology that in my experience works well. Of course, garbage in garbage out - thorough human planning (AI assisted, not autonomous) is key.
- stingraycharles 7mo agoTo me that doesn't do enough yet in terms of up-front planning and visualization, but it's a step in the right direction. I prefer Traycer myself.
- mshark 7mo agoHadn’t seen Traycer, that looks really polished. An important difference is that eforge is open source (Apache 2.0). I purposefully left out planning features from eforge because I don’t want the same tool that builds my code to force me into a planning methodology. Our role as developers has shifted heavily into planning (offloading implementation), and I’m still getting comfortable with that and want to be free to explore the planning space. Maybe I’ll change my mind after my planning opinions evolve.
- berkay 7mo agoI like the separation of planning and execution. I think the right set of artifacts to pass on to the execution will evolve but may be it's different for different types of work. From the project: "The plugin enqueues the input and a daemon picks it up - planning, building, reviewing, and validating autonomously." The part that is not clear to me (and causes most problems for me) is the "validating". It makes a mistake, or decides mocking an interface is fine, etc. declares success and moves on to the next. The bigger the project the more small mistakes compound. It sounds like the agent is doing the validation. What's the approach here for validation?
- saltpath 7mo agoThe parallel execution model makes sense for independent tickets but I'm wondering what happens when agent A is halfway through a PR touching shared/utils.py and agent B gets assigned a ticket that needs the same file. Does the orchestrator do any upfront dependency analysis to detect that, or do you just let them both run and deal with the conflict at merge time?
- vidarh 7mo agoIt's generally not worth it worrying about it too much other than at a very high level vs. letting them fight it out, as long as your test suite is good enough and your orchestrator is even moderately prepared to handle retries.
- Andrew_McCarron 6mo ago[flagged]
- Acacian 7mo ago[dead]
- fhouser 7mo agoHot take: You should want to review your agents' output and progress.
- the_real_cher 7mo agoThis project should be called the Rube Goldeberg machine creator.
- fhouser 7mo agoThe Hitchhiker's Guide to issue-tracking.
- jawiggins 7mo agoYeah totally, you don't have to auto-merge anything - you can review the PRs yourself
- fhouser 7mo agoYeah, I think that's the most important part in these new types of processes. Although it is tempting to just let an agent run with it for a while.
- vidarh 7mo agoI prefer to have my agents review my agents output and progress, and have them improve the prompts for future runs.
- ferreyadinarta 7mo ago[flagged]
- maxdo 7mo agoIs the pod per repo or per task ?
- jawiggins 7mo agoOne pod is an instance of a repo, you can set the number of instances of each agent/task that can be running on a pod at a time. For >1, each agent should be using it's own worktree.
- hustleracer 7mo ago[flagged]
- MarcelinoGMX3C 7mo ago[dead]
- pistoriusp 7mo agoHey @jawiggins, would you considering using https://github.com/redwoodjs/agent-ci https://github.com/redwoodjs/agent-ci?
- bmd1905 7mo ago[dead]
- georaa 7mo ago[flagged]
- Andrew_McCarron 6mo ago[flagged]
- olegbk 6mo agoThe feedback loop is what most people miss when they build these systems. You spin up the agent, it submits a PR, CI goes red, and suddenly you're back to being the bottleneck you were trying to eliminate. One thing I ran into building something similar, agents are surprisingly good at fixing the exact error message they're given, but terrible at recognizing when they're going in circles. After the third retry on the same failing test, you're not getting a fix, you're getting increasingly creative excuses for why the test is wrong. How deep does the self-healing go? Is there a retry limit before it escalates, or does it just keep going until you manually intervene?
- psychomfa_tiger 6mo ago[flagged]