2 ms·
Nikhil here, CTO @ BuildBetter, YC founder (W19). We’re building ZeroShot, an agent session monitoring tool that builds and shares self-improving skills with yo
by chaos_emergent 2mo ago
Nikhil here, CTO @ BuildBetter, YC founder (W19). We’re building ZeroShot, an agent session monitoring tool that builds and shares self-improving skills with your team so you can work with AI agents the way your most productive engineers do.
Download the Mac Desktop App at tryzeroshot.com/download or run curl -fsSL tryzeroshot.com | sh to give it a spin!
What ZeroShot does: it captures your coding-agent sessions locally (Claude, Codex, Cursor, Pi, OMP, more on request), then runs a classification pipeline on them in the background, looking for signals that the agent could improve. When it finds a big enough cluster of signals related to the same frictions, we submit them to your local agent and get it to generate a skill. It's different from something that you can bootstrap off of your own agent sessions locally in the following ways:
1. Cross-session skill drafting: agents aren't good at telling you when they encounter the same friction again in a separate session, even when pointing an agent at session logs retroactively.
2. Team-level enrichment (paid tier): We look across your team's sessions/GitHub PR comments and contribution history, and weight skills toward the people who are experts in a given codebase or area. Our current development focus is on surfacing the strategies that your most productive teammates use to work with AI; early days for agentic engineering in general. If you have suggestions on how to track that productivity, I'm all ears.
3. Your agentic documentation becomes self-maintaining: we're trying to offload the cognitive burden of managing harness engineering so that you can focus on delivering business value.
The problem we kept hitting: Between September and November of last year, I found myself getting more upset with the vibe-slop that made its way into main from engineers who knew better (and some who didn’t). So we started harness engineering. At first we just wanted to make sure that the models wouldn't make the same dumb mistakes that we kept on repeating in PR comments and to our own agents. We realized that specific engineers were able to “hold” the models better and were more productive because of it. ZeroShot is born from that internal experimentation that led us to being 5x more productive as a team.
We were inspired to build the app by our own internal self-improving-skill system, which itself was inspired by systems that the Codex team has espoused on X.
Privacy: the app runs entirely on your machine for free. Your agent's session logs only leave the computer when you opt in to a paid team plan and otherwise stay entirely local.
A non-exhaustive list of limitations:
1. Our classifiers work best when there's a repeated friction pattern across sessions. If you're doing highly varied work, we may not be able to pick up on patterns that are worth incorporating into the harness.
2. The numbers on our website are from teams with heavy repeated work. I don't want to oversell the outcomes you'll get.
3. Works with Claude Code, Codex, Cursor, Pi, and OMP today, we're adding more on request.
4. Mac only for now, Windows in the works, Linux CLI.
Pricing: local app is free, paid plans are for team features.
Happy to answer any questions or hear any feedback!