4 ms·
I assume it queues up requests if there are too many in flight to handle all of them at the same time, but it shows you the tokens per second (T/s) for every re
by coder543 3y ago
I assume it queues up requests if there are too many in flight to handle all of them at the same time, but it shows you the tokens per second (T/s) for every response, which is the number that matters (and presumably won't include the time spent in the queue).
- stavros 3y agoYou can be confused by why people are interested in the outputs, or you can think that the explanation is adequate. You can't do both at once. Here, in fact, the explanation is not adequate. Let's analyze: > Welcome Groq® Prompster! Are you ready to experience the world's fastest Large Language Model (LLM)? "The world's fastest LLM? They must have made an LLM, I guess" > We'd suggest asking about a piece of history, requesting a recipe for the holiday season, or copy and pasting in some text to be translated; "Make it French." "Instructions, ok". > This alpha demo lets you experience ultra-low latency performance using the foundational LLM, Llama 2, 70B created by Meta AI - running on Groq. "So it's not their own LLM? It's just LLaMa? What's interesting about that? And what's Groq, the thing this is 'running on'? The website? I kind of guessed that, by virtue of being here." > Like any AI demo, accuracy, correctness, or appropriateness cannot be guaranteed. "Ok, sure". Nowhere does it say "we've built new, very fast processing hardware for LLMs. Here's a demo a bog-standard LLM, LLaMa 70B, running on our hardware. Notice how fast it is".
- deleted 3y ago[deleted]