3 ms·
One thing about LLMs is that from what I know, batching is very effective with them, because to generate a token you need to basically go through the entire net
by ColonelPhantom 3y ago
One thing about LLMs is that from what I know, batching is very effective with them, because to generate a token you need to basically go through the entire network for just that one sample. If you batch, you still need to stream the entire network, but that doesn't get any more expensive; you use each piece of data more often on hardware that used to just sit idle.
So I assume that at an OpenAI-scale, they're more than able to batch up requests to their models, which gives a tiny latency increase (assuming they're getting many requests per second), but massively improves compute utilization.