3 ms·
This is why I was careful to specify tokens per second not time to first token which is the prefill step you’re talking about. Clearly to run these models well
by andy_ppp 13d ago
This is why I was careful to specify tokens per second not time to first token which is the prefill step you’re talking about. Clearly to run these models well you need both but as I said adding more chips or compute units can give you more latency where as overall memory bandwidth (throughput) is limited by access to fast memory.