4 ms·
We run it at my org and it's never been a noticeable resource hog. It's actually the best performer between it, our AI observability stack and the front end.
by mpyne 22d ago
We run it at my org and it's never been a noticeable resource hog. It's actually the best performer between it, our AI observability stack and the front end.
- blazarquasar 21d agoIt may not be a huge resource hog, but it adds a ton of latency. https://www.getmaxim.ai/bifrost/resources/benchmarks https://www.getmaxim.ai/bifrost/resources/benchmarks Having ran both LiteLLM and Bifrost for months, I can largely confirm the numbers from those benchmarks for myself.
- mpyne 21d agoIt may, but the latency it contributes to the end-to-end AI processing has been not noticeable in practice for our users. That's not to say Bifrost wouldn't have been better, but the choice to use LiteLLM was arrived at after a fair bit of internal discussion (most of which predated my addition to the team), and so far we've seen nothing from LiteLLM that has been contradictory to the pros/cons they thought would be the case when LiteLLM was adopted. Or in other words, the org will be happy indeed when they have solved so many of the rest of the problems we've had in AI uptake that the difference in latency between one AI gateway or the other becomes a problem to be solved.