4 ms·
time to first token != tokens per second remember that EU -> US is ~150ms unavoidable latency, for example. then your comparison is local H100 vs Grok + 150ms
by huac 3y ago
time to first token != tokens per second
remember that EU -> US is ~150ms unavoidable latency, for example. then your comparison is local H100 vs Grok + 150ms latency to first token.
- LoganDark 3y ago> time to first token != tokens per second I said "and extremely low latency" because I know they are different. Groq's TTFT is still consistently competitive with any other provider, and lower than most of them. Here's some benchmarks: https://github.com/ray-project/llmperf-leaderboard#70b-models-1 https://github.com/ray-project/llmperf-leaderboard#70b-model...
- nl 3y agoI'm in Australia. I have 249ms of unavoidable latency and I'd still use the groq API if I could. It's that much faster than other inference solutions.