3 ms·
Models are already good enough. Last week I had Fable plan out a project that took approx 4 days end to end with each phase orchestrated by a supervisor agent d
by FearNotDaniel 15d ago
Models are already good enough. Last week I had Fable plan out a project that took approx 4 days end to end with each phase orchestrated by a supervisor agent delegating individual tasks to other agents, coordinating everything and checking status by simple text files in the repo. I didn’t have to tell the agent to do it that way, it just came up with it and set up the infra as part of the planning overview. Great that everyone’s posting their “my secret sauce” cookbooks just to jump on the hype train but it looks like the models have already figured it out for themselves.
- simianwords 15d agotrue haha > Great that everyone’s posting their “my secret sauce” cookbooks just to jump on the hype train but it looks like the models have already figured it out for themselves Unfortunate lesson to be learned here: there's not much leverage here other than just using AI. Previously, us devs could get a head start and build some institutional knowledge but not this time. I'm bearish on all the custom harnesses stuff that people talk about. What helps me is to understand the failure modes of LLMs - it can't be articulated in easy words but something you can learn slightly by just using it. For example I have an intuition of when to start compacting but Codex already does it for you now haha. My take: the highest leverage move for us is to write AGENTS.md and provide everything that the model can't learn on its own or might take time to learn.
- frumiousirc 15d agoWhat is the interplay between such delegation and token cache timeout? If Fable is truly active for 4 days, enough to keep the cache hot, the cost would be astronomical. OTOH, if Fable idles while subagents are active, then each awakening is a cache miss.
- FearNotDaniel 15d agoTrue, that’s why I didn’t have a single Fable thread running for the whole time. Each phase had its own orchestrator that ran about 4-5h, in Fable or Opus depending on complexity, delegating to sub agents along the way - they had tasks broken down in a granular enough way that each returned in anything from 5 to 30 mins approx, then the lead agent did some validating and status updates before delegating the next step. So the main agent’s cache never went cold. Each phase raised its own PR and had an opportunity for human validation/code review before the next phase kicked off in a new context.