4 ms·
It’s just running bog standard Llama2-70B by all appearances. I don’t know why so many people here are interested in the outputs. The whole point of this demo
by coder543 3y ago
It’s just running bog standard Llama2-70B by all appearances.
I don’t know why so many people here are interested in the outputs. The whole point of this demo is that the company is trying to show off how fast their hardware could host one of your models, not the model itself.
- stavros 3y agoI'd assume it's because that's not explained very well in the linked page.
- coder543 3y agoThere is quite literally a modal pop-up that explains it, which you must dismiss before you can begin interacting with the demo. Quoting the pop-up: "This alpha demo lets you experience ultra-low latency performance using the foundational LLM, Llama 2, 70B created by Meta AI - running on Groq." Towards the bottom of the page, it also says "Model: Llama 2 70B/4096". Below that, it says "This is a Llama2-based chatbot."
- stavros 3y agoI posted this here, but it somehow moved, and I can't delete that one, so I'll repost. You can be wonder why people are interested in the outputs, or you can think that the explanation is adequate. You can't do both at once. Here, in fact, the explanation is not adequate. Let's analyze: > Welcome Groq® Prompster! Are you ready to experience the world's fastest Large Language Model (LLM)? "The world's fastest LLM? They must have made an LLM, I guess" > We'd suggest asking about a piece of history, requesting a recipe for the holiday season, or copy and pasting in some text to be translated; "Make it French." "Instructions, ok". > This alpha demo lets you experience ultra-low latency performance using the foundational LLM, Llama 2, 70B created by Meta AI - running on Groq. "So it's not their own LLM? It's just LLaMa? What's interesting about that? And what's Groq, the thing this is 'running on'? The website? I kind of guessed that, by virtue of being here." > Like any AI demo, accuracy, correctness, or appropriateness cannot be guaranteed. "Ok, sure". Nowhere does it say "we've built new, very fast processing hardware for LLMs. Here's a demo a bog-standard LLM, LLaMa 70B, running on our hardware. Notice how fast it is".
- addandsubtract 3y agoDoesn't the speed also depend on the number of people currently accessing it?
- coder543 3y agoI assume it queues up requests if there are too many in flight to handle all of them at the same time, but it shows you the tokens per second (T/s) for every response, which is the number that matters (and presumably won't include the time spent in the queue).
- stavros 3y agoYou can be confused by why people are interested in the outputs, or you can think that the explanation is adequate. You can't do both at once. Here, in fact, the explanation is not adequate. Let's analyze: > Welcome Groq® Prompster! Are you ready to experience the world's fastest Large Language Model (LLM)? "The world's fastest LLM? They must have made an LLM, I guess" > We'd suggest asking about a piece of history, requesting a recipe for the holiday season, or copy and pasting in some text to be translated; "Make it French." "Instructions, ok". > This alpha demo lets you experience ultra-low latency performance using the foundational LLM, Llama 2, 70B created by Meta AI - running on Groq. "So it's not their own LLM? It's just LLaMa? What's interesting about that? And what's Groq, the thing this is 'running on'? The website? I kind of guessed that, by virtue of being here." > Like any AI demo, accuracy, correctness, or appropriateness cannot be guaranteed. "Ok, sure". Nowhere does it say "we've built new, very fast processing hardware for LLMs. Here's a demo a bog-standard LLM, LLaMa 70B, running on our hardware. Notice how fast it is".
- deleted 3y ago[deleted]