3 ms·
Things don't have to be exactly in the same category to usefully compare them. Groq has emerged as the premiere way to host LLMs at scale at speeds that are dra
by eigenvalue 2y ago
Things don't have to be exactly in the same category to usefully compare them. Groq has emerged as the premiere way to host LLMs at scale at speeds that are dramatically faster than anyone else, which allows for all kinds of new and exciting applications that wouldn't work if you had to wait for 50tok/sec responses, but which feel magical at 500tok/sec. And although the exciting with LLMs seems to be around the training of them, I think if you look out a few years, vastly more FLOPs will be expended on inference than on training.
- lostmsu 2y agoWhere can I find technical performance comparison? What models does Groq run at 500tok/sec? In what mode (e.g. batch size)? UPD. found the model, it is Mixtral 8x7B-32k. AFAIK 8xH100 will do 100+tok/sec with batch size 1. But that does not look too impressive for batch sizes higher than 1: https://www.baseten.co/blog/faster-mixtral-inference-with-tensorrt-llm-and-quantization/ https://www.baseten.co/blog/faster-mixtral-inference-with-te...