4 ms·
I've built an orchestrator that solves some of the issues you ran into (although it doesn't do anything about cheating): https://navels.dev/blog/neal/ https://n
by navels 1mo ago
I've built an orchestrator that solves some of the issues you ran into (although it doesn't do anything about cheating): https://navels.dev/blog/neal/ https://navels.dev/blog/neal/. Features:
- lets you configure different models for planner, coder, and reviewer roles. (e.g., using Claude as an adversarial reviewer against Codex)
- breaks your plan up into reasonable-sized chunks of work with clearly defined success criteria
- runs each chunk of work through a coder / read-only reviewer loop. Once both agents are satisfied, neal moves on to the next chunk. Once everything is complete there is a final pass through the coder / reviewer loop to ensure the implementation satisfies the entire plan.
- resets the coder's context with each chunk of work to prevent context drift, leaving the reviewer's context long-running.
- chrisweekly 1mo agoWow, "neal" looks excellent. Good on you for creating and sharing it, and for the awesome blog post.
- navels 1mo agoThanks!
- avadodin 1mo agoI don't think even the frontier models recognize something was produced by the same model in order to maliciously review it positively. They may share some blind spots with the producer but generally I think they will review the other agent's output as harshly as they can if that is their task.
- gregwebs 1mo agoThat's a neat project for doing a large scale migration. I do the same for normal feature develompent but just with skills that are in this repo: https://github.com/gregwebs/skills-sdlc https://github.com/gregwebs/skills-sdlc I have accomplished code base (small size) migrations with it as well. Currently I do review each PR. For a large code base migration I think the core skills would still work but need a different way of driving it as you have come up with.
- killix 1mo ago[flagged]
- navels 1mo agoFollowing up on @killix's comments, which were helpful but he was flagged (presumably for sounding too much like AI). Thanks for the feedback. I've made a couple of updates: - Starting with 0.4.0, the reviewer gets the diff of the earlier chunk for any file the current chunk touches again. Also, if a new chunk weakens or removes a test or assertion from an earlier chunk, the reviewer will block it unless the plan says to do so. - About read-only: The reviewer's tools were already limited by the SDK (no shell or write tools). However, it was still finding MCP servers from my Claude config. I have now blocked those. The docs now explain what is enforced by the system and what is just a prompt instruction.