4 ms·
> NeMo Switchyard, an open source library for smart routing > When deployed, NeMo Switchyard can intelligently direct each request to the most capable and suit
by thehamkercat 2mo ago
> NeMo Switchyard, an open source library for smart routing
> When deployed, NeMo Switchyard can intelligently direct each request to the most capable and suitable model for the job
How do routers like this handle prompt caching when you send the second request?
Sticky models per session? but then the second message of that session won't be sent to a suitable model, and will only be sent to the same model as previous one.
- embedding-shape 2mo agoThe repo is probably a better entrypoint to it, bit more concise description than the press releases: https://github.com/NVIDIA-NeMo/Switchyard https://github.com/NVIDIA-NeMo/Switchyard (Notably: "Experimental software. Not for production use."). Unclear if they actually want you to deploy it or not, press release says yes, README says no, do with that what you will. Doesn't seem to mention "cache" in the README nor the docs, but the code has mentions of it (https://github.com/search?q=repo%3ANVIDIA-NeMo%2FSwitchyard+cache&type=code https://github.com/search?q=repo%3ANVIDIA-NeMo%2FSwitchyard+...), I'm not sure what their thinking is there. "Good luck" essentially? Seems to be per-provider at best, but weird position for a routing library to take.
- thehamkercat 2mo agoI personally think it's snake-oil marketing with all these smart-model-routing products/projects prompt-cache won't work with these
- try-working 2mo agoTo keep it simple, forget about routers and imagine you're in Cursor using GPT for a while, reaching a cache of says 200k. You decide to switch to DeepSeek in the same session via the model picker, and continue as usual. What happens is that the cache for DeepSeek is created with the 200k + the incremental message. After this, cache can be kept warm for both models; two instances of the cache exists, one for GPT and one for DS. You switch back to GPT. The whole session is sent to the model with the 200k original from GPT and the incremental messages you sent to DS. The 200k is read from cache and the incrementals are new, and then added to the cache. Let's say every second message you switch between GPT and DS; cache was 200k and each incremental message is 1k. If you kept going with only GPT, cache hit rate would be 200k/(200k+1k) = 99.5%. When you switch between two models with warm cache, hit rate instead becomes 200k/(200k+2k) = 99%. Model routers work the same way. Keep the cache warm, replicate it in two places. For this reason, when you set up your model pool for routing, you want to keep the model pool small and differentiated. First principles of model routing: https://try.works/first-principles-of-model-routing https://try.works/first-principles-of-model-routing role-model router and protocol: https://github.com/try-works/role-model https://github.com/try-works/role-model note: edited to keep the answer to the below message clearer
- thehamkercat 2mo agoCan you explain how does it work? like how is the previous K/V cache used when you switch to another model? Source?
- hedgehog 2mo agoSee sibling answer but essentially the effectiveness of cache is not diminished by having a separate one per model (relative to the win of doing more turns and generation with a cheaper model).
- try-working 2mo agoedit: updated the answer above to be more qualitative instead
- hedgehog 2mo agoTo elaborate, because I don't think some of the people reading this understand the reason, typically a lot or most of the cost in "agentic" API usage is cached read + generation. Cached read costs scale with turn count, which multi-model switching doesn't increase, and of course generation gets cheaper if you do some of it with a cheaper model. When you switch models the "catching up" batch of messages is just a single prefill and then that goes into cache. You don't even need to have the same chat history across models so long as the view from each model's perspective looks like a series of appends. The main problem with model routing in my experience is that to work well the router needs to be pretty strong, maybe even moreso than any of the actual models in service. There are probably clever solutions to this but I haven't seen any that look better than just using sub-agents.
- richwater 2mo ago> which multi-model switching doesn't increase Given model A with cache C(a) and model B with C(b) Isn't this not true because the moment you switch models from A to B, you need to provide C(b) the latest conversation diff since C(b) last updated, say many turns ago?
- eli 2mo agoI've seen ones that are configurable to pick a trade off point between lower cost (cache stickiness) and routing performance (best model for that turn). But yeah I'm skeptical all this overhead is worth it.
- quinncom 2mo agoCaching should be possible as long as all the models use the same shared cache. The models don't even need to be running on the same server if the shared cache is distributed. I have a feeling people reading this are thinking that a model router would be used to route between different providers. And in that case, a shared cache would be impossible, although some caching would still be effective. I think, ideally, a router like this is in front of a set of models hosted in one place.
- IanCal 2mo agoHow do caches work across models? I would have thought that was very model specific - if not I’ve really misunderstood what’s getting cached.
- armanckeser 2mo agoI am not sure the author of the comment you are replying to understands that LLM systems have prompt caches
- amluto 2mo agoHuh? Prompt caching isn’t about caching the literal text of the prompt. It’s about caching the result of running prefill on the prompt (or, equivalently, the result of generating the prompt one token at a time by autoregressive inference, or some combination of the above in the case of speculative decoding). This is often called the “KV” cache, and it is very model-specific.
- rufasterisco 2mo agohttps://github.com/NVIDIA-NeMo/Switchyard#routing-strategies https://github.com/NVIDIA-NeMo/Switchyard#routing-strategies Looks like your great question doesn’t have an answer, but looking at the routing strategies things get even more confused, since the proposed ones tend to rely on extra llm calls to determine which model to pick. The nice thing is that it makes sense for specific setups, less conversation oriented. As an example, you need to classify batches of data, and have many fine tuned models. Or you need to do speed to text and need to pick which whisper to use. You can write your own strategy, in that case an harness with subagents would be able to leverage this, picking the right model and then keeping its session sticky, but overall the lack of concern for caching points towards use cases where you do not gain much from it.