4 ms·
If you self-host, you likely won't have anywhere near enough volume to do efficient batching, and end up bottlenecked on memory rather than compute. E.g. based
by jsnell 1y ago
If you self-host, you likely won't have anywhere near enough volume to do efficient batching, and end up bottlenecked on memory rather than compute.
E.g. based on the calculations in https://www.tensoreconomics.com/p/llm-inference-economics-from-first https://www.tensoreconomics.com/p/llm-inference-economics-fr..., increasing batch size from 1 to 64 cuts the cost per token to 1/16th.