4 ms·
LLM inference has large economies of scale. Properly batched requests are tens of times cheaper than individually processed ones. And it's going to be quite har
by jsnell 1y ago
LLM inference has large economies of scale. Properly batched requests are tens of times cheaper than individually processed ones. And it's going to be quite hard for a self-hoster to have enough hardware + enough usage to benefit from high levels of batching.