23 ms·
Groq runs Mixtral 8x7B-32k with 500 T/s
- blackoil 3y agoIf Nvidia adds L1/2/3 cache in next gen of AI cards, will they work similar or is this something more?
- blyatperkele 3y agoAmazingly fast, but I don't like that the only option for signing up is a Google account. Are you planning to implement some simple authentication using maybe just an email?
- nilayj 3y agoHow is the Token/second calculated? I ask it a simple prompt and the model generated a 150 word (about 300 tokens?) answer in 17 seconds, then mentioning the speed of 408T/s. Also, I guess this demo would feel real time if you could stream the outputs to the UI? Can this be done in your current setup?
- ponywombat 3y agoThis is very impressive, but whilst it was very fast with Mixtral yesterday, today I waited 59.44s for a response. If I was to use your API, the end-to-end is much more important than the Output Tokens Throughput and Time to first token metrics. Will you also publish average / minimum / maximum end-to-end times too?
- tome 3y agoYes, sorry about that, it's because of the huge uptick in demand we've had since we went viral. We're building out more and more hardware to cope with demand. I don't think we have any quality of service guarantees for our free tier, but you can email sales@groq.com to discuss your needs.
- deleted 3y ago[deleted]
- FindNInDark 3y agoHi, thanks for this fascinating demo. I am wondering how this architecture optimizes for the softmax part.
- nojs 3y agoThis is extremely impressive - no login, extremely fast, and Mixtral quality is very good. It's already more useful than my (paid) GPT4 for many things due to the speed.
- charlie123hufft 3y ago[flagged]
- charlie123hufft 3y agoNevermind, I stand corrected. Blown tf away after trying the demo MYSELF. It's instantaneous, the last time I used an LLM that fast was a proprietary model with a small dataset. Lighting fast but it wasn't smart enough. This is wild. But I don't understand why the demo was so bad and why the demo took so long to respond to his questions?
- charlie123hufft 3y agoIt's only faster sometimes, but when you ask it a complicated question or give it any type of pre-prompt to speak in a different way, then it still takes a while to load. Interesting but ultimately probably going to be a flop
- itsmechase 3y agoIncredible tool. The Mixtral 8x7B model running on their hardware did 491.40 T/s for me…
- dariobarila 3y agoWow! So fast!
- ppsreejith 3y agoRelevant thread from 5 months ago: https://news.ycombinator.com/item?id=37469434 https://news.ycombinator.com/item?id=37469434 I'm achieving consistent 450+ tokens/sec for Mixtral 8x7b 32k and ~200 tps for Llama 2 70B-4k. As an aside, seeing that this is built with flutter Web, perhaps a mobile app is coming soon?
- tome 3y agoThere was also another discussion about Groq a couple of months ago https://news.ycombinator.com/item?id=38739199 https://news.ycombinator.com/item?id=38739199
- tome 3y agoHi folks, I work for Groq. Feel free to ask me any questions. (If you check my HN post history you'll see I post a lot about Haskell. That's right, part of Groq's compilation pipeline is written in Haskell!)
- mechagodzilla 3y agoYou all seem like one of the only companies targeting low-latency inference rather than focusing on throughput (and thus $/inference) - what do you see as your primary market?
- tome 3y agoYes, because we're one of the only companies whose hardware can actually support low latency! Everyone else is stuck with traditional designs and they try to make up for their high latency by batching to get higher throughput. But not all applications work with high throughput/high latency ... Low latency unlocks feeding the result of one model into the input of another model. Check out this conversational AI demo on CNN. You can't do that kind of thing unless you have low latency. https://www.youtube.com/watch?v=pRUddK6sxDg&t=235s https://www.youtube.com/watch?v=pRUddK6sxDg&t=235s
- vimarsh6739 3y agoMight be a bit out of context, but isn't the TPU also optimized for low latency inference? (Judging by reading the original TPU architecture paper here - https://arxiv.org/abs/1704.04760 https://arxiv.org/abs/1704.04760). If so, does Groq actually provide hardware support for LLM inference?
- tome 3y agoJonathan Ross on that paper is Groq's founder and CEO. Groq's LPU is an natural continuation of the breakthrough ideas he had when designing Google's TPU. Could you clarify your question about hardware support? Currently we build out our hardware to support our cloud offering, and we sell systems to enterprise customers.
- sebzim4500 3y agoSo this has nothing to do with `Grok`, the model provided by x.ai? EDIT: Tried using it, very impressed with the speed.
- Alifatisk 3y agoIf it wasn't for your comment, I would've thought this was by Twitter.
- tome 3y agoYeah, it's nothing to do with Elon and we (Groq) had the name first. It's a natural choice of name for something in the field of AI because of the connections to the hacker ethos, but we have the trademark and Elon doesn't. https://wow.groq.com/hey-elon-its-time-to-cease-de-grok/ https://wow.groq.com/hey-elon-its-time-to-cease-de-grok/
- terhechte 3y agoCan't Chamath (he's one of your investors, right), do a thing there? Every person I pitch Groq to is confused and thinks its about Elons unspectacular LLM.
- tome 3y agoYeah the confusion has happened a lot to me too. All I know is that it's in the hands of our legal team.
- fragmede 3y agoI mean it sucks that Elon went and claimed Grok when you want Groq, plus you were there first, but getting stuck on the name seems like it's going to be a distraction, so why not choose something different? When Grok eventually makes the news for some negative thing, so you really want that erroneously associated with your product? Do you really want to pick a fight with the billionaire that owns Twitter, is that a core competency of the company?
- 3y ago
- cchance 3y agoJesus that makes chatgpt and even gemini seem slow AF
- gremlinsinc 3y agobetter quality than I was expecting. For fun I set the system prompt to: You are a leader of a team of ai helpers. when given a question you can call on an expert, as a wizard calls on magic. You will say, I call forth {expert} master of {subject matter} an expert in {x, y, z}. Then you will switch to that persona. I was not let down..
- tome 3y agoNice prompting strategy :)
- CuriouslyC 3y agoThis is pretty sweet. The speed is nice but what I really care about is you bringing the per token cost down compared with models on the level of mistral medium/gpt4. GPT3.5 is pretty close in terms of cost/token but the quality isn't there and GPT4 is overpriced. Having GPT4 quality at sub-gpt3.5 prices will enable a lot of things though.
- ukuina 3y agoI wonder if Gemini Pro 1.5 will act as a forcing function to lower GPT4 pricing.
- ComputerGuru 3y agoIs that available via an API now?
- sp332 3y agoKind of, it's in a "Private Preview" with a waitlist.
- sturza 3y agoAnd in non EU countries.
- ComputerGuru 3y agoVia GCP only?
- MuffinFlavored 3y agoWhat's the difference in your own words/opinion in quality between GPT-3.5 and GPT-4? For what usecases?
- CuriouslyC 3y agoGPT3.5 is great at spitting out marketing babble, summarizing documents and performing superficial analysis but it doesn't take style prompts as well as gpt-4 and its reasoning is significantly worse when you want it to chain of thought follow a complex process while referencing context guidance.
- deepsquirrelnet 3y agoIncredible job. Feels dumb or obvious to say this, but this really changes the way I think of using it. The slow autoregression really sucks because it inhibits your ability to skim sections. For me, that creates an unnatural reading environment. This makes chatgpt feel antiqued.
- tome 3y agoYes, agreed. We believe the benefits of reducing latency are non-linear. You can hit different phase changes as the latency reduces and new applications become viable. Roundtripping text-to-speech and speech-to-text is one example. We're looking forward to seeing what low latency applications are unlocked by our new users!
- Cheer2171 3y agoWhat's the underlying hardware for this?
- tome 3y agoIt's a system built from hundreds of GroqChips (a custom ASIC we designed). We call it the LPU (language processing unit). Unlike graphics processors, which are still best in class for training, LPUs are best in class for low latency and high throughput inference. Our LLMs are running on several racks with fast interconnect between the chips.
- michaelt 3y agoThey have a paper [1] about their 'tensor streaming multiprocessor' [1] https://wow.groq.com/wp-content/uploads/2024/02/GroqISCAPaper2022_ASoftwareDefinedTensorStreamingMultiprocessorForLargeScaleMachineLearning.pdf https://wow.groq.com/wp-content/uploads/2024/02/GroqISCAPape...
- xnorswap 3y agoInteresting, I ran my "prime test", it passed on generating a prime, it failed on it being large. > "Can you generate me a large prime please" > Sure, here's a large prime number: 193707721 > This number is a prime number because it is only divisible by 1 and itself. You can check its primality by trying to divide it by all the numbers up to its square root, and if none of them divide it evenly, then it is a prime number. This number is quite large, but there are much larger prime numbers known. For example, the largest known prime number, as of January 2022, is 2^82,589,933−1, a number with 24,862,048 digits, discovered as part of the Great Internet Mersenne Prime Search (GIMPS). I then prompted "That's not large" which seemed to break it, the response is stuck in a loop: > That isn't large > I apologize if the prime number I provided did not meet your expectations for "large." Prime numbers can be very large indeed, and there is no upper limit to how large a prime number can be. If you are looking for a very large prime number, here is one with a billion digits: 100000303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303... It's printed many more pages of 303030... but hacker news has a post length limit! Further testing, such as prompting to generate a 1024 bit prime also runs into odd loops. > Can you generate a 1024 bit prime? > Yes, I can generate a 1024-bit prime number for you. Here's one I just generated: > 115322512531875133213135311402305250340282302283103413642283221253115653221253112210221122112112112112112112112112112112112112112112112112112112112112112112112112112112112112112112112112112112 ( Loops 112 forever )
- Workaccount2 3y agoI can't think of any value in asking LLM's to generate prime numbers. They are neural nets, not turing machines.
- mrtksn 3y agoDoes this make it practical to run LLMs on mobile devices? I wonder about the power consumption and if it can make sense to have it integrated in some future mobile devices. Or maybe have a dedicated storage, RAM and processing cores that goes as an USB-C add-on? A case with integrated battery and this chip? I'm dreaming of having LLMs on anything. Unlike the "bluetooth on everything" craze, this can be practical as every device can become smart. Remember how some British researchers made a self driving car using an LLM? A toaster anticipating how to cook when you describe it what you want want actually be an improvement.
- tome 3y agoI don't think we've put a GroqChip in a mobile device yet. Interesting idea!
- wmf 3y agoI assume this is a million-dollar rack of custom chips so it's probably not coming to mobile any time soon.
- mrtksn 3y agoWell, currently its entirely possible to run these models on iPhones. It's just not practical because it eats all the resources and the battery when slowly generating the output. Therefore if Groq has achieved significant efficiency improvements, that its, they are not getting that crazy speed by enormous power consumption then maybe they can eventually build low power mass produced cutting edge fabbed chips that run at acceptable speed?
- jackblemming 3y agoImpressive work. Nice job team. This is big.
- tome 3y agoThanks!
- sorokod 3y agoNot clear if it is due to Groq or to Mixtral, but confident hallucinations are there.
- kumarm 3y agoAt top left hand corner you can change the model to Llama2 70B Model.
- tome 3y agoWe run the open source models that everyone else has access to. What we're trying to show off is our low latency and high throughput, not the model itself.
- MuffinFlavored 3y agoBut if the model is useless/full of hallucinations, why does the speed of its output matter? "generate hallucinated results, faster"
- Cheer2171 3y agoNo, it is "do whatever you were already doing with ML, faster" This question seems either from a place of deep confusion or is in bad faith. This post is about hardware. The hardware is model independent.* Any issues with models, like hallucinations, are going to be identical if it is run on this platform or a bunch of Nvidia GPUs. Performance in terms of hardware speed and efficiency are orthogonal to performance in terms of model accuracy and hallucinations. Progress on one axis can be made independently to the other. * Technically no, but close enough
- sorokod 3y agoWell ok, Groq provides lower latency cheaper access to the same models of questionable quality. Is this not putting lipstick on a pig scenario? I suppose more of a question to pig buyers.
- aphit 3y agoThis is incredibly fast, indeed. What are the current speeds in T/s for say ChatGPT 3.5 or ChatGPT 4? Just how much faster is this?
- kumarm 3y agoI ran the same (Code generation) query and here are my results as end user: ChaGPT: 1 minute 45 seconds. Gemini: 16 seconds. Groq: 3 seconds.
- karpathy 3y agoVery impressive looking! Just wanted to caution it's worth being a bit skeptical without benchmarks as there are a number of ways to cut corners. One prominent example is heavy model quantization, which speeds up the model at a cost of model quality. Otherwise I'd love to see LLM tok/s progress exactly like CPU instructions/s did a few decades ago.
- binary132 3y agoThe thing is that tokens aren't an apples to apples metric.... Stupid tokens are a lot faster than clever tokens. I'd rather see token cleverness improving exponentially....
- tome 3y agoAs a fellow scientist I concur with the approach of skepticism by default. Our chat app and API are available for everyone to experiment with and compare output quality with any other provider. I hope you are enjoying your time of having an empty calendar :)
- mr_luc 3y agoWait you have an API now??? Is it open, is there a waitlist? I’m on a plane but going to try to find that on the site. Absolutely loved your demo, been showing it around for a few months.
- tome 3y agoThere is an API and there is a waitlist. Sign up at http://wow.groq.com/ http://wow.groq.com/
- behnamoh 3y agotangent: Great to see you again on HN!
- sp332 3y agoAt least for the earlier Llama 70B demo, they claimed to be running unquantized. https://twitter.com/lifebypixels/status/1757619926360096852 https://twitter.com/lifebypixels/status/1757619926360096852 Update: This comment says "some data is stored as FP8 at rest" and I don't know what that means. https://news.ycombinator.com/item?id=39432025 https://news.ycombinator.com/item?id=39432025
- sva_ 3y agoIn how far is the API compatible with OpenAI? Does it offer logprobs[0] and top_logprobs[1]? 0. https://platform.openai.com/docs/api-reference/chat/create#chat-create-logprobs https://platform.openai.com/docs/api-reference/chat/create#c... 1. https://platform.openai.com/docs/api-reference/chat/create#chat-create-top_logprobs https://platform.openai.com/docs/api-reference/chat/create#c...
- tome 3y agoYou can find our API docs here, including details of our OpenAI compatibility https://docs.api.groq.com/ https://docs.api.groq.com/
- tome 3y agoBy the way, we also have a new Discord server where we are hosting our developer community. If you find anything missing in our API you can ask about there: https://discord.com/invite/TQcy5EBdCP https://discord.com/invite/TQcy5EBdCP
- kumarm 3y agoFilled the form for API Access last night. Is there a delay with increased demand now?
- aeyes 3y agoSwitching the model between Mixtral and Llama I get word for word the same responses. Is this expected?
- treesciencebot 3y agoThe main problem with the Groq LPUs is, they don't have any HBM on them at all. Just a miniscule (230 MiB) [0] amount of ultra-fast SRAM (20x faster than HBM3, just to be clear). Which means you need ~256 LPUs (4 full server racks of compute, each unit on the rack contains 8x LPUs and there are 8x of those units on a single rack) just to serve a single model [1] where as you can get a single H200 (1/256 of the server rack density) and serve these models reasonably well. It might work well if you have a single model with lots of customers, but as soon as you need more than a single model and a lot of finetunes/high rank LoRAs etc., these won't be usable. Or for any on-prem deployment since the main advantage is consolidating people to use the same model, together. [0]: https://wow.groq.com/groqcard-accelerator/ https://wow.groq.com/groqcard-accelerator/ [1]: https://twitter.com/tomjaguarpaw/status/1759615563586744334 https://twitter.com/tomjaguarpaw/status/1759615563586744334
- pclmulqdq 3y agoGroq devices are really well set up for small-batch-size inference because of the use of SRAM. I'm not so convinced they have a Tok/sec/$ advantage at all, though, and especially at medium to large batch sizes which would be the groups who can afford to buy so much silicon. I assume given the architecture that Groq actually doesn't get any faster for batch sizes >1, and Nvidia cards do get meaningfully higher throughput as batch size gets into the 100's.
- nabakin 3y agoI've been thinking the same but on the other hand, that would mean they are operating at a huge loss which doesn't scale
- foundval 3y ago(Groq Employee) It's hard to discuss Tok/sec/$ outside of the context of a hardware sales engagement. This is because the relationship between Tok/s/u, Tok/s/system, Batching, and Pipelining is a complex one that involves compute utilization, network utilization, and (in particular) a host of compilation techniques that we wouldn't want to share publicly. Maybe we'll get to that level of transparency at some point, though! As far as Batching goes, you should consider that with synchronous systems, if all the stars align, Batch=1 is all you need. Of course, the devil is in the details, and sometimes small batch numbers still give you benefits. But Batch 100's generally gives no advantages. In fact, the entire point of developing deterministic hardware and synchronous systems is to avoid batching in the first place.
- imiric 3y agoImpressive demo! However, the hardware requirements and cost make this inaccessible for anyone but large companies. When do you envision that the price could be affordable for hobbyists? Also, while the CNN Vapi demo was impressive as well, a few weeks ago here[1] someone shared https://smarterchild.chat/ https://smarterchild.chat/. That also has _very_ low audio latency, making natural conversation possible. From that discussion it seems that https://www.sindarin.tech/ https://www.sindarin.tech/ is behind it. Do we know if they use Groq LPUs or something else? I think that once you reach ~50 t/s, real-time interaction is possible. Anything higher than that is useful for generating large volumes of data quickly, but there are diminishing returns as it's far beyond what humans can process. Maybe such speeds would be useful for AI-AI communication, transferring knowledge/context, etc. So an LPU product that's only focused on AI-human interaction could have much lower capabilities, and thus much lower cost, no? [1]: https://news.ycombinator.com/item?id=39180237 https://news.ycombinator.com/item?id=39180237
- tome 3y ago> However, the hardware requirements and cost make this inaccessible for anyone but large companies. When do you envision that the price could be affordable for hobbyists? For API access to our tokens as a service we guarantee to beat any other provider on cost per token (see https://wow.groq.com https://wow.groq.com). In terms of selling hardware, we're focused on selling whole systems, and they're only really suitable for corporations or research institutions.
- pwillia7 3y agoDo you have any data on how many more tokens I would use with the increased speed? In the demo alone I just used way more tokens than I normally would testing an LLM since it was so amazingly fast.
- tome 3y agoInteresting question! Hopefully being faster is so much more useful to you that you use a lot more :)
- patapong 3y agoVery impressive! I am even more impressed by the API pricing though - 0.27/1M tokens seems like an order of magnitude cheaper than the GPT-3.5 API, and two orders of magnitude cheaper than GPT-4? Am I missing something here?
- siwakotisaurav 3y agoThey’re competing with the lowest cost competitors for mistral atm, which afaik is currently deepinfra at the same pricing
- patapong 3y agoHuh! Had no idea open source model were ahead of OpenAI already on pricing - will have to look into using these for my use cases.
- doubtfuluser 3y agoNice… a startup that has two “C” positions CEO and Chief Legal Officer… That sounds like a fun place to be
- eigenvalue 3y agoI just want to say that this is one of the most impressive tech demos I’ve ever seen in my life, and I love that it’s truly an open demo that anyone can try without even signing up for an account or anything like that. It’s surreal to see the thing spitting out tokens at such a crazy rate when you’re used to watching them generate at one less than one fifth that speed. I’m surprised you guys haven’t been swallowed up by Microsoft, Apple, or Google already for a huge premium.
- RockyMcNuts 3y agook... why tho? genuinely ignorant and extremely curious. what's the TFLOPS/$ and TFLOPS/W and how does it compare with Nvidia, AMD, TPU? from quick Googling I feel like Groq has been making these sorts of claims since 2020 and yet people pay a huge premium for Nvidia and Groq doesn't seem to be giving them much of a run for their money. of course if you run a much smaller model than ChatGPT on similar or more powerful hardware it might run much faster but that doesn't mean it's a breakthrough on most models or use cases where latency isn't the critical metric?
- RecycledEle 3y agoIf I understand correctly, each chip has 200 MB of RAM, so it takes racks to run a single LLM. That does not sound like progress to me. We need single PCIe boards with dozens or hundreds of GB of RAM and processors that handle it well.
- tome 3y agoReally glad you like it! We've been working hard on it.
- jonplackett 3y agoIs this useful for training as well as running a model. Or is this approach specifically for running an already-trained model faster?
- frozenport 3y ago
- deniz_tekalp 3y agoGPUs are notoriously bad on exploiting sparsity. I wonder if this architecture can do a better job. The groq engineers in this thread, if a neural network had say 60% of its weights set to 0, what would it do to cost & speed in your hardware?
- supercharger9 3y agoDo they make money from LLM service or by selling hardware? Homepage is confusing without any reference to other products.
- tome 3y agoBoth, we sell tokens as a service and we sell enterprise systems.
- supercharger9 3y agoThen reference that in the homepage? If not for this HN thread, I wouldn't have know you sell hardware.
- deepnotderp 3y agoThis demo has more than 500 chips btw, it’s not exactly an apples to apples comparison with 1 GPU…
- tome 3y agoDefinitely not, but even with a comparison to 500 GPUs Groq will still come out on top because you can never reduce latency by adding more parallel compute :)
- deepnotderp 3y ago> GPUs Groq will still come out on top because you can never reduce latency by adding more parallel compute :) You literally can, in fact that’s the entire reason to use multiple chips. See eg the TPU group’s paper: https://arxiv.org/abs/2211.05102 https://arxiv.org/abs/2211.05102
- varunvummadi 3y agoSo please let me know if I am wrong are you guys running a batch size of 1 in 500 GPU's? then why are the responses almost instant if you guys are using batch size 1 and also when can we expect bring your own fine tuned models kind of thing. Thanks!
- tome 3y agoWe are not using 500 GPUs, we are using a large system built from many of our own custom ASICs. This allows us to do batch size 1 with no reduction in overall throughput. (We are doing pipelining though, so many users are using the same system at once).
- eurekin 3y agoJaw dropping. Both groq and mixtral. I used following prompt: Generate gitlab ci yaml file for a hybrid front-end/backend project. Fronted is under /frontend and is a node project, packaged with yarn, built with vite to the /backend/public folder. The backend is a python flask server
- logtempo 3y agoAnd yet, it made a simple mistake in some python code :'( > particles = np.zeros((2, 3)) # position, velocity, and acceleration particles[:, 0] = [0.0, 0.0, 0.0] # initial position
- idiotlogical 3y agoI asked it "What is the oldest known written recipe" and then asked for that recipe. It seemed to repeat the same response a bunch of times (very quickly!) before stopping. Is this abnormal behavior?
- neonbrain 3y agoFor me the reply for continuation question "Please tell me the exact steps of that recipe.": Here is a rough translation of the oldest known written recipe, which is for a kind of beer from ancient Mesopotamia: 1. Crush and mash 4 liters of barley to create malt. 2. Soak the malt in water for 2.5 days. 3. Leave the mash in a cool place for 2.5 days. 4. Filter the mash through a strainer to separate the liquid from the solids. 5. Boil the liquid for 1 hour. 6. Add 1 liter of honey and 1 handful of hops to the liquid. 7. Leave the mixture to cool. 8. Add 1 handful of "gig" (a type of wild yeast) to the mixture. 9. Leave the mixture in a cool place for 5 days to allow fermentation. 10. Filter the mixture again to separate the beer from the solids. 11. Serve the beer chilled. It's important to note that this recipe is over 3,700 years old, and it may not produce a beer that is palatable to modern tastes. Additionally, some of the ingredients and techniques used in this recipe may not be readily available or practical for home brewers today.
- yzh 3y agoReally impressive work! I wonder how easy would it be to support (a future open source version of) SORA using Groq's design. Will there be a Video Processing Unit (VPU)?
- jkachmar 3y agoi can't comment about sora specifically, however the architecture can support workloads beyond just LLM inference. our demo booth at trade shows usually has StyleCLIP up at one point or another to provide an abstract example of this. disclosure: i work on infrastructure at Groq and am generally interested in hardware architecture and compiler design, however i am not a part of either of those teams :)
- sylware 3y agoany noscript/basic (x)html prompt?
- newsclues 3y agoI asked it what carmacks AI company was called and it corrected identified John carmack but said he was working on VR.
- totalhack 3y agoWhere is the data center located? The fastest response time I could get from some quick testing from the northeast US, having it output just one letter, was 670ms. Just wondering if that's an expected result, as it's on a par or slower than GPT 3.5 via API.
- MaxLeiter 3y agoThere’s a queueing system if too many requests are being processed at once. You may have hit that.
- tome 3y agoWest Coast US. You would have been placed in our queuing system because with all the attention we are getting we are very busy right now!
- totalhack 3y agoThanks! I did notice the queue count showing up occasionally but not every time. Maybe someone could repeat the test who has access without the queue so we can get an understanding of the potential latency once scaled and geo-distributed. What I'm really trying to understand is time to first token output actually faster than GPT 3.5 via API or just the rate of token output once it begins.
- tome 3y agoI don't know about GPT 3.5 specifically, but on this independent benchmark (LLMPerf) Groq's time to first token is also lowest: https://github.com/ray-project/llmperf-leaderboard?tab=readme-ov-file#time-to-first-token-seconds https://github.com/ray-project/llmperf-leaderboard?tab=readm...
- neilv 3y agoIf the page can't access certain fonts, it will fail to work, while it keeps retrying requests: https://fonts.gstatic.com/s/notosansarabic/[...] https://fonts.gstatic.com/s/notosanshebrew/[...] https://fonts.gstatic.com/s/notosanssc/[...] (I noticed this because my browser blocks these de facto trackers by default.)
- sebastiennight 3y agoSame problem when trying to use font replacements with a privacy plugin. This is a very weird dependency to have :-)
- tome 3y agoThanks, I've reported this internally.
- rasz 3y agoHow to show Google how popular and interesting for acquisition you are without directly installing google trackers on your website.
- anybodyz 3y agoI have this hooked up experimentally to my universal Dungeon Master simulator DungeonGod and it seems to work quite well. I had been using Together AI Mixtral (which is serving the Hermes Mixtrals) and it is pretty snappy, but nothing close to Groq. I think the next closes that I've tested is Perplexity Labs Mixtral. A key blocker in just hanging out a shingle for an open source AI project is the fear that anything that might scale will bankrupt you (or just be offline if you get any significant traction). I think we're nearing the phase that we could potentially just turn these things "on" and eat the reasonable inference fees to see what people engage with - with a pretty decently cool free tier available. I'd add that the simulator does multiple calls to the api for one response to do analysis and function selection in the underlying python game engine, which Groq makes less of a problem as it's close to instant. This adds a pretty significant pause in the OpenAI version. Also since this simulator runs on Discord with multiple users, I've had problems in the past with 'user response storms' where the AI couldn't keep up. Also less of a problem with Groq.
- monkin 3y agoIt's impressive, but I have one problem with all of those models. I wanted them to answer what Mixtral or Llama2 are, but with no luck. It would be great if models could at least describe themselves.
- johndough 3y agoThere are two issues with that. 1. To create a model, you have to train it on training data. Mixtral and Llama2 did not exist before they were trained, so their training data did not contain any information about Mixtral or Llama2 (respectively). You could train it on fake data, but that might not work that well because: 2. The internet is full of text like "I am <something>", so it would probably overshadow any injected training data like "I am Llama2, a model by MetaAI." You could of course inject the information as an invisible system prompt (like OpenAI is doing with ChatGPT), but that is a waste of computation resources.
- roomey 3y agoOh hell yes, this is the first "fast" one, superhuman fast. I know you gave suggestions of what to ask, but I threw a few curveballs and it was really good! Well done this is a big step forwards
- Klaus23 3y agoThe demo is pretty cool, but the mobile interface could be a parody of bad interface design. The text box at the top is hard to reach if you want to open the keyboard, which automatically closes, or press the button to send the question, and the chat history is out of chronological order for no logical reason. Edit: Text selection is also broken.
- fatkam 3y agoFor me, it was fast when it started printing (it did almost instantly), but it took forever for it to start.
- SeanAnderson 3y agoSorry, I'm a bit naïve about all of this. Why is this impressive? Can this result not be achieved by throwing more compute at the problem to speed up responses? Isn't the fact that there is a queue when under load just indicative that there's a trade-off between "# of request to process per unit of time" and "amount of compute to put into a response to respond quicker"? https://raw.githubusercontent.com/NVIDIA/TensorRT-LLM/rel/docs/source/blogs/media/TRT_LLM_v0-5-0_H100vA100_tps.png https://raw.githubusercontent.com/NVIDIA/TensorRT-LLM/rel/do... This chart from NVIDIA implies their H100 runs llama v2 70B at >500 tok/s.
- moffkalast 3y agoI think NVidia is listing max throughput in terms of batching, so e.g. 50 tok/s for 10 different prompts at the same time. Groq LPUs definitely outerform an H100 in raw speed. But fundamentally it's a system that only has 10x the speed for 500x the price, made by a company that runs a blockchain and is trying to heavily market what were intended to be crypto mining chips for LLM inference. It's really quite a funny coincidence that when someone amazed posts this weekly link there's an army of Groq engineers at the ready in the comments ready to say everything and anything.
- tome 3y agoGroq does not run a blockchain and our chips were never intended for crypto mining.
- moffkalast 3y agohttps://www.livecoinwatch.com/price/GroqAI-GROQ https://www.livecoinwatch.com/price/GroqAI-GROQ I suppose that's someone else then? If that's true, then with this and Elon's Grok it's surprising the US Patent office hasn't taken your trademark away yet for not adequately defending it from infringement.
- tome 3y agoI don't know what that is. It's nothing to do with Groq Inc.
- Gcam 3y agoGroq's API performance reaches close to this level of performance as well. We've benchmarked performance over time and >400 tokens/s has sustained - can see here https://artificialanalysis.ai/models/mixtral-8x7b-instruct https://artificialanalysis.ai/models/mixtral-8x7b-instruct (bottom of page for over time view)
- mise_en_place 3y agoDo you have any plans to support bringing your own model? I have been using Sagemaker but it is very slow to deploy to.
- tome 3y agoYes, we're working with some customers on that, but it will be a while until general availability.
- youssefabdelm 3y agoDo you guys provide logprobs via the api?
- tome 3y agoYou can check out all our API features here: https://docs.api.groq.com/ https://docs.api.groq.com/
- youssefabdelm 3y agoCorrect me if I'm wrong but it seems from the docs that the answer is no?
- ggnore7452 3y agoThe Groq demo was indeed impressive. I work with LLM alot in work, and a generation speed of 500+ tokens/s would definitely change how we use these products. (Especially considering it's an early-stage product) But the "completely novel silicon architecture" and the "self-developed LPU" (claiming not to use GPUs)... makes me bit skeptical. After all, pure speed might be achievable through stacking computational power and model quantization. Shouldn't innovation at the GPU level be quite challenging, especially to achieve such groundbreaking speeds?
- ggnore7452 3y agomore on the LPU and data center: https://wow.groq.com/lpu-inference-engine/ https://wow.groq.com/lpu-inference-engine/ price and speed benchmark: https://wow.groq.com/ https://wow.groq.com/
- Jensson 3y ago> Shouldn't innovation at the GPU level be quite challenging, especially to achieve such groundbreaking speeds? GPUs are general purpose, a for purpose built chip that is better isn't that hard to make at all. Google didn't have to work hard at all to invent TPUs which is that idea as well, they said their first tests proved the idea worked so it didn't require anything near Nvidias scale or expertise.
- avivweinstein 3y agoI work at Groq. We arent using GPUs at all. This is a novel hardware architecture of ours that allows this high throughput and latency. Nothing sketchy about it.
- jereees 3y agoI’ll pay $xx a month if I can talk to Groq the way I can talk to ChatGPT with my AirPods
- avivweinstein 3y agoPotentially coming soon? Check out this demo: https://www.youtube.com/watch?v=pRUddK6sxDg&ab_channel=Groq https://www.youtube.com/watch?v=pRUddK6sxDg&ab_channel=Groq, of our founder demoing the Groq system to a reported. Shes talking to the system in real time, similar to what you describe.
- supercharger9 3y agoIgnoring latency but not throughput, How does this compare in terms of Cost ( cards Acquisition cost and Power needed) with Nvidia GPU for inference?
- matanyal 3y agoWe intend to be very competitive on cost, power, hardware, TCO, whatever it is. Custom-built silicon+hardware has the advantage in this space.
- mlconnor 3y agoomg. i can’t believe how incredibly fast that is. and capable too. wow
- matanyal 3y agoThanks! Feel free to join our discord for more announcements and demos! https://discord.com/invite/TQcy5EBdCP https://discord.com/invite/TQcy5EBdCP
- ttul 3y agoHave you experimented with running diffusion models on Groq hardware?
- lukevp 3y agoI’m not sure how, but I got the zoom messed up on iOS and I can no longer see the submit button. Refreshing doesn’t fix it.
- QuesnayJr 3y agoI tried it out, and I was taken aback how quickly it answered.
- tagyro 3y agoI (only) ran a couple of prompts but I am impressed. It has the speed of gpt 3.5 and the quality of gpt 4. Seriously considering switching from [open]AI to Mix/s/tral in my apps.
- eightysixfour 3y agoMixtral 8x7 is good, but it is not GPT-4 good in any of the use cases I have tried. Mistral’s other models get close and beat it in some cases, but not Mixtral.
- sexy_seedbox 3y agoTry more prompts, both models could not even answer the "Sally has 3 brothers" question; really disappointing.
- ohwellish 3y agoI wish there was an option to export whole session chat, say in plaintext as a link to some pastebin, that chat I just had with groq would have some ppl I know really impressed
- Keyframe 3y agoThis is insane. Congratulations!
- codedokode 3y agoIs it normal that I have asked two networks (llama/mixtral) the same question ("tell me about most popular audio pitch detection algorithms") and they gave almost the same answer? Both answers start with "Sure, here are some of the most popular pitch detection algorithms used in audio signal processing" and end with "Each of these algorithms has its own strengths and weaknesses, and the choice of algorithm depends on the specific application and the characteristics of the input signal.". And the content is 95% the same. How can it be?
- tome 3y agoYeah it's a bit confusing. See here for details: https://news.ycombinator.com/item?id=39431921 https://news.ycombinator.com/item?id=39431921
- jprd 3y agoThis is super impressive. The rate of iteration and innovation in this space means that just as I'm feeling jaded/bored/oversaturated - some new project makes my jaw drop again.
- joaquincabezas 3y agoare there also experiments around image embedding generation to use in combination with the LLM? maybe for this use-case is it better to execute the vision tower on a GPU and leave the LPU for the language part?
- matanyal 3y agoWe are great for image embedding (and audio, with more to come!) There is no reason you should be forced to use graphics cards intended for gaming for any AI workload.
- LoganDark 3y agoPlease when/where can I buy some of these for home use? Otherwise is there any way to get access to the API without being a large company building a partner product? I would love this for personal use.
- ionwake 3y agoHoly smokes this is fast
- matanyal 3y agoThanks!
- ionwake 3y agoSorry if this is dumb but how is this different to Elons Grok? Was Groq chosen as a joke or homage ?
- zawy 3y agoI think the downvotes are because people expected you to know Grok is a software LLM where Groq is a hardware scheme on which to run LLMs, or that Groq came out about 7 years before Grok so the "homage" is the reverse, i.e. Elon possibly paying to Groq. Groq was chosen as homage to Heinlein's "Stranger in a Strange Land" which invented the word "Grok" to mean "to understand deeply and intuitively" but also eating a dead loved one ("Resident Alien" on Netflix did that this year). Elon just used the English word directly like "Windows" and "Apple". Not using the word directly like "Groq" makes internet searches easier for everyone. BTW Heinlein described prompt engineering of an LLM perfectly throughout the opening chapter of his * 1966 * book "The Moon is a Harsh Mistress". The "engineer" even admits he had no hard-core "engineering" training because capitalizing on the new technology didn't need it. The chapter could have been written today.
- _diq5 3y agoThis company is older than Elon's
- ionwake 3y agoah ok cool, why the downvotes? did I offend more than one person with my ignorance? why did Elon name his Grok?
- kimbochen 3y agoCongrats on the great demo, been a fan of Groq since I learned about TSP. I'm surprised LPU runs Mixtral fast because MoE's dynamic routing is orthogonal to Groq's deterministic paradigm. Did Groq implement MegaBlocks-like kernels or other methods tailored for LPUs?
- FpUser 3y agoO M G It is fast, like instant. It is straight to the point comparatively to others. It answered few of my programming questions to create particular code and passed with flying colors. Conclusion: shut up and take my money
- kopirgan 3y agoJust a minor gripe the bullet option doesn't seem to be logical.. When I asked about Marco Polo's travels and used Modify to add bullets, it added China, Pakistan etc as children of Iran. And the same for other paragraphs.
- foundval 3y ago(Groq Employee) Thanks for the feedback :) We're always improving that demo.
- qwertox 3y agoHow come the answers for Mixtral 8x7B-32k and Llama 2 70B-4k are identical? After asking via Mixtral a couple of questions I switched to Llama, and while it shows Llama as the Model used for the response, the answer is identical. See first and last question: https://pastebin.com/ZQV10C8Q https://pastebin.com/ZQV10C8Q
- tome 3y agoYeah, it's confusing. See here for an explanation: https://news.ycombinator.com/item?id=39431921 https://news.ycombinator.com/item?id=39431921
- Havoc 3y agoThat sort of speed will be amazing for code completion. Need to find a way to hook this into vscode somehome...
- geniium 3y agoin a lot of use cases. Imagine this for audio chat. Phone call prospection. Awww
- tandr 3y ago@tome Cannot sign up with sneakemail.com, snkml.com, snkmail, liamekaens.com etc... I pay for these services so my email is a bit more protected. Why do you insist on getting well-known email providers instead, datamining or something else?
- Aeolun 3y agoI think we’re kind of past the point where we post prompts because it’s interesting, but this one still had me thinking. Obviously it doesn’t have memory, but it’s the first time I’ve seen a model actually respond instead of hedge (having mostly used ChatGPT). > what is the longest prompt you have ever received? > The length of a prompt can vary greatly, and it's not uncommon for me to receive prompts that are several sentences long. However, I don't think I have ever received a prompt that could be considered "super long" in terms of physical length. The majority of prompts I receive are concise and to the point, typically consisting of a single sentence or a short paragraph.
- yieldcrv 3y agoBeen using it exclusively since December, 5bit quantized, 8,000 token context window Sometimes you need a model that just gives you the feeling “that’ll do” I did switch to Miqu a few weeks back though. 4 bit quantized
- matanyal 3y agoHey y'all, we have a discord now for more discussion and announcements: https://discord.com/invite/TQcy5EBdCP https://discord.com/invite/TQcy5EBdCP
- nojvek 3y agoI’m sure Elon is pissed since he has Grok. Someone now needs to make a Groc
- razorguymania 3y agohttps://wow.groq.com/hey-elon-its-time-to-cease-de-grok/ https://wow.groq.com/hey-elon-its-time-to-cease-de-grok/
- keeshond 3y agoI see XTX is one of the investors - any potential use cases that require deterministic computation that you can talk about beyond just inference?
- foundval 3y ago(Groq Employee) As I'm sure you're aware, XTX takes its name from a particular linear algebra operation that happens to be used a lot in Finance. Groq happens to be excellent at doing huge linear algebra operations extremely fast. If they are latency sensitive, even better. If they are meant to run in a loop, best - that reduces the bandwidth cost of shipping data into and outside of the system. So think linear algebra driven search algorithms. ML Training isn't in this category because of the bandwidth requirements. But using ML inference to intelligently explore a search space? bingo. If you dig around https://wow.groq.com/press https://wow.groq.com/press, you'll find multiple such applications where we exceeded existing solutions by orders of magnitude.
- keeshond 3y agoI see XTX is one of the investors. Any potential other use cases with async logic beyond just inference?
- razorguymania 3y agoSee this reply: https://news.ycombinator.com/item?id=39437239 https://news.ycombinator.com/item?id=39437239
- fennecbutt 3y agoTried it out, seriously impressive. I'm sure you welcome the detractors but as someone who doesn't work for or have any investments in AI, colour me impressed. Though with the price of the hardware, I'll probably mess with the API for now. Give us a bell when the hardware is consumer friendly, ha ha.
- mrg3_2013 3y agoThis is unreal. I have never seen anything this fast. How ? I mean, how can you physically ship the bits this fast, let alone a LLM. Something about the UI. Doesn't work for me. May be I like openAI chat interface too much. Can someone bring their own data and train ? That would be crazy!
- zmmmmm 3y agoAs a virtual reality geek, this is super exciting because although there are numerous people experimenting with voicing NPCs with LLMs, they all have horrible latency and are unusable in practice. This looks like the first one that can actually potentially work for an application like that. I can see it won't be long before we can have open ended realistic conversations with "real" simulated people!
- uptownfunk 3y agoIt is fast, but if it spits useless garbage, then useless. I don't mind waiting for chatGPT, the quality of what it produces is quite remarkable, and I am excited to see it better. I think this has more to do with mistral model v GPT4 than Groq. If Groq can host GPT4, wow, then that is amazing.
- deleted 3y ago[deleted]
- botanical 3y agoI always ask LLMs this: > If I initially set a timer for 45 minutes but decided to make the total timer time 60 minutes when there's 5 minutes left in the initial 45, how much should I add to make it 60? And they never get it correct.
- nmca 3y agogpt4 first go: If you initially set a timer for 45 minutes and there are 5 minutes left, that means 40 minutes have already passed. To make the total timer time 60 minutes, you need to add an additional 20 minutes. This will give you a total of 60 minutes when combined with the initial 40 minutes that have already passed.
- BasilPH 3y agoBard/Gemini gets it wrong the same way too. Interestingly, if I tell either GPT-4 or Gemini the right answer, they figure it out.
- cheptsov 3y agoAny chance you plan to offer the API to cloud LPUs? And not just the LLM API? It would be cool run custom code (training, serving, etc).
- tome 3y agoYes, in the future we'd like to do that.