3 ms·
No, this is an agent-level thing, not a feature of the model (ish, unsure for Fable). You talk to a smart, heavy model to build a plan composed of smaller step
by everforward 4mo ago
No, this is an agent-level thing, not a feature of the model (ish, unsure for Fable).
You talk to a smart, heavy model to build a plan composed of smaller steps. Then you have the heavy model spin up smaller, cheaper LLMs to actually implement the tasks.
The heavy model is basically read-only in that mode. It can read files, execute tests, etc, but it can’t write code. It just tracks what needs to be done, offloads the work to dumber LLMs, validates the task is done, and moves on to the next step.
It sort of pushes humans up the stack. Instead of having a human sitting there prompting the LLM to start the next task, you have another LLM do that loop.
It’s been on my list to try out.
- alchemism 4mo agoThe AWS Kiro (https://kiro.dev https://kiro.dev) spec-driven coding harness operates this way in Auto mode which offers the base token rate. Manually-specifying Sonnet or Opus is a multiplier on the base token rate; specifying Qwen fractions it. Left to its own, it presumably uses the heavier models to create the plan and orchestrate the work; the bite-sized task definitions are delegated to smaller models.