3 ms·
Tokens per second is almost entirely memory bandwidth at inference time, training obviously needs more compute but you can add more chips for that.
by andy_ppp 15d ago
Tokens per second is almost entirely memory bandwidth at inference time, training obviously needs more compute but you can add more chips for that.
- cubefox 15d agoAccording to SemiAnalysis, both inference and post-training (RLVR) is mostly memory bandwidth bound. Only pre-training is compute bound, but it now only takes a small share of overall data center capacity. https://x.com/EugeneNg/status/2099315982959616369 https://x.com/EugeneNg/status/2099315982959616369
- darig 15d ago[dead]
- martinald 14d agoNot quite, it's got quite a bit more complicated with agentic use cases. Prefill (input tokens) is heavily compute bound. And the ratio of input to output continues to rise, as typically in agentic sessions you have a few tokens output for a tool call and (many) thousands of input from the tool result. Then you have cached input tokens, which is a totally different issue, system RAM or NVMe bound. Obviously output tokens is VRAM memory bandwidth bound, but this is less and less of the bottleneck these days for overall agentic speed.
- com2kid 14d agoI can easily use close to 100 million input tokens a day. A few million output tokens but at the end of the day maybe a thousand or so lines of code get written.
- andy_ppp 14d agoThis is why I was careful to specify tokens per second not time to first token which is the prefill step you’re talking about. Clearly to run these models well you need both but as I said adding more chips or compute units can give you more latency where as overall memory bandwidth (throughput) is limited by access to fast memory.