3 ms·
Great article, I think the most important aspect from it is the auto-routing. As humans laziness is in our nature, so having to think if the model is capable en
by imilev 2mo ago
Great article, I think the most important aspect from it is the auto-routing. As humans laziness is in our nature, so having to think if the model is capable enough is not something that most ppl will do - resulting in trying out smaller models which failed our task and then just giving up and running on the bigger model all the time.
- Terretta 2mo agoDISCLOSURE: I like Databricks. While cosplaying enterprise CTO, I've directed the purchase and heavy integration of their work for over a decade. That said, this type of post needs to be read with product marketing context in mind. > I think the most important aspect from it is the auto-routing On the contrary, in white paper studies auto-routing is shown to destroy the single largest token cost they found, curiously not shown in their opening graphic. As they put it in their own words near the end: "simple tuning of… caching settings… 50% reduction in … costs, with no observed quality degradation…" They also aren't showing the harness called ‘pi’ which generally tops results (not only with ‘open’ models), and is consistent with Anthropic's recent "works better when we don't stuff 100k system prompts into context" about Opus 5. Recent models do better with less jank trying to prescribe behavior. So then one wonders: why is this post written to say you should use a meta router but no mention that's offset by the cache busting, and show developers cost savings but no mention that Anthropic and OpenAI both have developer-facing utilization dashes already? Perhaps it's so engineers can show enterprise procurement why they want to buy exactly what Databricks happens to have for sale.
- yeswecatan 2mo agoDo you know they are using pi or are you inferring it?
- Terretta 2mo agoSorry, didn't mean they're using pi. They're itemizing ways to get better results but not showing pi, while other testing shows most models (including Opus) achieve better from within pi.
- aabdi 2mo agoIf that’s such a large problem then the clear solution is to do sticky sessions? What’s the problem then? Use a cheap ds4 or Luna and do the second model net net per one shot best case you save couple dollars per?
- Terretta 2mo agoExactly. There are plenty things they could suggest for that. Anyway you'll likely get a more effective and up to date set of advice by skipping marketing pieces and pointing your favorite frontier model at /r/LocalLLaMA in deep research mode and asking it to synthesize the latest advice, or even try ideas out for itself and let you know what works for your own session history.
- ankitmathur 2mo agoHey! Thanks for the feedback! I work on many of these things at Databricks, so figured I'd chime in on this. Firstly, while routing is important, simple things like observability into the token costs of various features, which can drive optimization of better default parameters, are low hanging fruit everyone should do. 1. As for the cache busting, we're very aware of this. Had said this in a thread above too, so copying it here. We're going to do a followup blog detailing our routing approach, but we're designing it to be cache aware. The router takes in the task description and infers what models and harnesses are available and makes a recommendation up-front. So essentially the routing decision is made when the harness + model is kicked off and it's only changed halfway through if there's a major delta in complexity from the initial judgment. Therefore, most of the time the cache is maintained just as it would be before. This does have disadvantages too, but the cache is the dominating cost reduction force. 2. We do love Pi (we published this too: https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase https://www.databricks.com/blog/benchmarking-coding-agents-d...), but we didn't mention it here because we're still in the process of making it available to devs internally. There's a lot that goes into this more than just cost (e.g. feature gaps vs. other common harnesses). 3. As for existing utilization dashboards, it's good that each tool has one, but we found that as we gave developers more freedom to use more tools, it was impossible to have a single source of truth. Without that, it's hard to really understand all-in AI spend and optimize it.
- Terretta 2mo ago> Firstly, while routing is important, simple things like observability STRONG agree. Most devs, if they can see their costs, try. Even just putting it in a starship prompt goes a long way! Observability supports effective, responsible, and accountable use, "cost controls" (as typically enterprise defined) supports bean counters. Ensuring Observability is recognized as a valid Control™ should be on every engineering executive's TODO list. > the cache is the dominating cost reduction force TY for the rich reply. As you're very aware of the cache issue, it might have been nice to have the word "cache" more prominent on your "Key Techniques" image (currently Lower Cost Models, Smart Routing, Spend Controls, Context Optimization; cache tuning buried at the end) or higher in the body sections. It was mostly the image that prompted (ha) me to call out the marketing driven angle.
- deleted 2mo ago[deleted]