3 ms·
Could you say more about how caching works? One major advantage of sticking with a single model is saving money on cached input tokens. I'd imagine if you swap
by Areibman 1mo ago
Could you say more about how caching works? One major advantage of sticking with a single model is saving money on cached input tokens. I'd imagine if you swap between a bunch of models, you may improve performance but cost would would balloon out of control
- purplecats 1mo agoand caching is related to performance too ofc
- SilenN 1mo agoThe trick is to rarely switch, or switch at task boundaries. Often the conclusion of routing is actually "this one model is actually at the pareto front for this task, just use it always".
- cameronh90 1mo agoBut then it's better to just not have a gateway switch models at all. Just have the harness able to choose which model its sub-agents use, then tell it how to split up tasks and which models to use when doing so.
- SilenN 1mo agoThat is another way to do. Or we can automatically figure out which models the subagents should be using for you. And update them as new models come out and the work your subagents do changes. More than one way to skin a cat.
- aerzen 1mo agoThis does make sense. I generally only switch between models in pi when creating a new session. And it is apparent from the promt if this just a "how to see open ports on linux" or "make a concrete plan for feature X"
- try-working 1mo agoGenerally you should only have two models in the pool per domain. I wrote some of my learnings building a router here: https://try.works/first-principles-of-model-routing https://try.works/first-principles-of-model-routing