3 ms·
Scaling up compute can improve throughput, but can't easily improve latency between tokens. Generation is usually bottlenecked by the time it takes to go throug
by MasterScrat 3y ago
Scaling up compute can improve throughput, but can't easily improve latency between tokens. Generation is usually bottlenecked by the time it takes to go through the network for each token. To speed that up, you need to perform these computations faster, which is a hard problem after you've exhausted all the obvious options (use the fastest accelerator you can find, cache what you can etc).
- SeanAnderson 3y agoYeah. That makes sense, thank you for clarifying. I updated my original post with a chart from NVIDIA which highlights the H100's capabilities. It doesn't seem unreasonable to expect a 7B model to run at 500 tok/s on that hardware.
- snowfield 3y agoThis is a 50B model. (Mixtral 8x7b)
- SeanAnderson 3y agoOh, sorry, I assumed the 8 was for quantization. 8x7b is a new syntax for me. Still, the NVIDIA chart shows Llama v2 70B at 750 tok/s, no?
- tome 3y agoI guess that's total throughput, rather than per user? You can increase total throughput by scaling horizontally. You can't increase throughput per user that way.
- qeternity 3y agoAt batch size 1 LLMs are memory bandwidth bound, not compute bound…as in you spend most time waiting for model weights to load from vram. At higher batch sizes this flips. But this is why Groq is built around large numbers of chips with small amount of very fast sram.