3 ms·
Great question! Generally the neural network used for the router takes maybe ~20ms during inference. When deployed on prem, in your own cloud environment, then
by danlenton 2y ago
Great question! Generally the neural network used for the router takes maybe ~20ms during inference. When deployed on prem, in your own cloud environment, then this is the only latecy. When using the public endpoints with our own intermediate server, it might add ~150ms to the time-to-first-token, but inter-token-latency is not affected.
We generally see the router being useful when the LLM application is being scaled, and cost and speed start to matter a lot. However, in some cases the output quality actually improved, as we're able to squeeze the best of GPT4 and Claude etc.
Long-term plan for profitability would come from some future version of the router, where we save the user time and money, and then charge some overhead for the router, but with the user still paying less than they would be with a single endpoint. Hopefully that makes sense?
Happy to answer any other questions!
- jonahx 2y agoDo you save the user data, ie, the searches themselves? What do your TOS guarantee about the use of that data?
- danlenton 2y agoWe use this data to improve the base router by default. It's fully anonymized, and you can opt out.
- ColinHayhurst 2y agoWithout opt out it would be a no go, so that's great to hear. What's the downside of opting out?
- danlenton 2y agono down side