17 ms·
Groq CEO: 'We No Longer Sell Hardware'
- BoorishBears 3y agoRead: We're forcing someone's hand in acquiring us. Groq is still under a 30 request per minute rate-limit, which drops to 10 requests per minute if you have all day usage. Billing has been "coming soon" this whole time, and while they've built out hype enabling features like function calling, somehow they can't setup a Stripe webhook to collect money for realistic rate limits. They couldn't scream "we can't service the tiniest bit of our demand" any louder at this point. _ Edit: For anyone looking for fast inference without the smoke and mirrors, I've been using Fireworks.ai in production and it's great. 200 tk/s - 300 tk/s is closer to Groq than it is to OpenAI and co. And as a bonus they support PEFT with serverless pricing.
- arthurcolle 3y agothey don't even let us pay them, it's insane I just have free API access with no ability to add a credit card.
- deleted 3y ago[deleted]
- brcmthrowaway 3y agoWhat are you using all this for? Whats the product?
- BoorishBears 3y agoI run an AI story telling site and an AI ideation platform. The story telling site alone averaged 27k requests a day this week, so about double what their current request limit is, and honestly not even that popular of a site. You can't run much more than a toy project on their current rate limits.
- scosman 3y agoWildly fast inference. And current chips are 14nm so headroom to get a lot better.
- jsheard 3y agoNote that SRAM density doesn't scale at the same rate as logic density, and Groqs "secret sauce" is putting a ton of SRAM on their chips. Their stuff won't necessarily see the full benefits of switching to denser nodes if the bottleneck is how much SRAM they can pack onto each chip. https://www.tomshardware.com/news/no-sram-scaling-implies-on-more-expensive-cpus-and-gpus https://www.tomshardware.com/news/no-sram-scaling-implies-on... IIRC the last big jump for SRAM density was at 7nm, so they do still have that card to play, but progress has slowed to a crawl beyond that. TSMC 3nm SRAM is barely denser than TSMC 7nm SRAM.
- gandalfgeek 3y agoThey're calling the lie on needing bleeding edge hardware for performance. 5 yr old silicon (14 nm!!) and no hbm. Their secret sauce seems to be an ahead-of-time compiler that statically lays out entire computation, enabling zero contention at runtime. Basically, they stamp out all non-determinism. https://wow.groq.com/isca-2022-paper https://wow.groq.com/isca-2022-paper
- halflings 3y agoNo HBM because they use tons of fast SRAM instead. Isn't that the main driver for performance here? (the way I understood it => it's still cost effective at scale due to throughput increase this brings)
- germanjoey 3y agocost effective in what sense? groq doesn't achieve high efficiency, only low latency. but that's not done in a cost-effective way. compare sambanova achieving the same performance with 8 chips instead of 568, and with higher precision.
- halflings 3y agoThe # of chips is not the most important metric. Most important, even ignoring latency, is throughput (tokens) per $$$. And according to their own benchmark [1] (famous last words :)), they're quite cost efficient. [1] https://www.semianalysis.com/p/groq-inference-tokenomics-speed-but https://www.semianalysis.com/p/groq-inference-tokenomics-spe...
- gandalfgeek 3y ago> No HBM because they use tons of fast SRAM instead. Isn't that the main driver for performance here? No doubt fast SRAM helps, but from a computation pov imho its that they've statically planned computation and eliminated all locks. Short explainer here: https://www.youtube.com/watch?v=H77tV1KcWIE https://www.youtube.com/watch?v=H77tV1KcWIE (Based on their paper).
- dsrtslnd23 3y agoSo unless there are new Croq datacenters coming, this is only interesting for North American users. Otherwise H100 based latency optimized solutions would be faster - in particular for time-to-first-token sensitive applications.
- LoganDark 3y ago> latency optimized solutions would be faster - in particular for time-to-first-token sensitive applications Do you have any idea how fast Groq is? Go try it. Consistently over 400 t/s for most of the models that they support, and extremely low latency.
- huac 3y agotime to first token != tokens per second remember that EU -> US is ~150ms unavoidable latency, for example. then your comparison is local H100 vs Grok + 150ms latency to first token.
- LoganDark 3y ago> time to first token != tokens per second I said "and extremely low latency" because I know they are different. Groq's TTFT is still consistently competitive with any other provider, and lower than most of them. Here's some benchmarks: https://github.com/ray-project/llmperf-leaderboard#70b-models-1 https://github.com/ray-project/llmperf-leaderboard#70b-model...
- nl 3y agoI'm in Australia. I have 249ms of unavoidable latency and I'd still use the groq API if I could. It's that much faster than other inference solutions.
- LoganDark 3y agoThat sucks. I wanted to save up for a couple years and get some hardware for home, but I guess the "AI" space moves so fast you barely get a couple months
- wmf 3y agoSave up for Tenstorrent instead.
- LoganDark 3y agoI'll look into it, though seeing "contact us" always makes me think they're not going to sell a single unit to a home user. (With that said, Groq probably wouldn't either. You can technically buy LPUs for 20k each, without an expectation of support, but it takes tens of them to run Mixtral.) Tenstorrent also looks incredibly Python-specific (as in, everything including their SMI seems mostly Python-based) which doesn't seem promising?
- sipjca 3y agofwiw as a consumer I have a tenstorrent card in my machine
- latchkey 3y agothat is great. please hit me up privately, i'd love to chat with you about it.
- ojn 3y agoMost of the low-level pieces are in Rust, the TUI is written in Python and most of the remaining pieces are getting lowered down to the Rust libraries over time. (It was all Python up until ~6 months ago) EDIT: Oh, and you can buy the Grayskull cards online now, without contacting anyone.
- latchkey 3y agoEven talking to people there, my experience is that they are super nice!
- zetazzed 3y agoMan, I want to appreciate a nice new hardware approach, but they say such BS that it is hard to read about them: > “There might need to be a new term, because by the end of next year we’re going to deploy enough LPUs that compute-wise, it’s going to be the equivalent of all the hyperscalers combined,” he said. “We already have a non-trivial portion of that.” Really? Does anyone seriously believe they are going to be the equivalent of all hyperscalers in compute next year? (Where Meta alone is at 1 million H100 equivalents.) In the same article where they say it's too hard for them to sell chips? And when they literally don't have a setup to even accept a credit card today?
- wmf 3y agoYou don't put a million-dollar rack on a credit card. I'm not sure they want retail customers for their API either.
- latchkey 3y agoInteresting, I guess that is why I never got a response back from them about buying their stuff. My guess is that they realized that just selling hardware is a lot harder than running it themselves. Deploying this level of compute is non-trivial, with very high rates of failure, as well as huge supply chain issues. If you have to sell the hardware and support people buying it, that is a world of trouble. > no-one wants to take the risk of buying a whole bunch of hardware I do! Nobody has stated it yet, but this is probably great news for tenstorrent. Disclosure: building a cloud compute provider starting with AMD MI300x, and eventually any other high end hardware that our customers are asking for.
- rnts08 3y agoThat's interesting, something that I've been really wanting to get into as well, but where I am there is literally no venture capital to raise for this at the moment. I'd be interested to know more and/or bounce some ideas though.
- latchkey 3y agoExtremely capital intensive, but also requires relationships in the industry at many levels. Luckily, I happen to have both and they are crazy enough to put their trust in me. I feel very grateful for that.
- htrp 3y agothat or if you put chips in the hands of your customers, they may start to benchmark it against other equivalent solutions
- latchkey 3y agoFunny you should mention that. ;-) https://www.reddit.com/r/LocalLLaMA/comments/1bpgrdf/wanted_amd_mi300x_benchmarks/ https://www.reddit.com/r/LocalLLaMA/comments/1bpgrdf/wanted_... I've got about a dozen people signed up. Just working through some hardware issues right now (see above about high rate of failures), and hope to have this resolved next week, so that I can get people onto them and doing their testing.
- creato 3y ago> If customers come with requests for high volumes of chips for very large installations, Groq will instead propose partnering on data center deployment. Ross said that Groq has “signed a deal” with Saudi state-owned oil company Aramco, though he declined to give further details, saying only that the deal involved “a very large deployment of [Groq] LPUs.” What? How does this make sense?
- theturtletalks 3y agoIf you read on, Groq said they would only sell hardware to US companies and outside companies would get cloud services, not the LPUs. I think the US government told them to keep the LPUs in-house since they could be the secret sauce for scale.
- creato 3y agoI'm not questioning the deployment strategy, I'm wondering why Saudi Aramco wants to access so much compute power that is highly specialized(?) for generative AI workloads. Or is it more general than that?
- aleph_minus_one 3y ago> I'm wondering why Saudi Aramco wants to access so much compute power that is highly specialized(?) for generative AI workloads. Or is it more general than that? For vanity reasons and because AI is the future (not every company acts that rationally for huge buying decisions).
- brookst 3y agoAlso diversification. They’re smart people; oil isn’t the future. Comparatively small investments to hedge make perfect sense.
- theturtletalks 3y agoMaybe Saudi doesn’t want to rely on OpenAI and other APIs and wants to run a fine-tuned Mixtral model on the cloud or their hardware. International companies will probably opt for an open source model since the data is sensitive and OpenAI could pass that to intelligence.
- rnts08 3y agoSounds like they're looking to get bought up to me. I'm sure they could monetize their current hardware, and build to sell just like other niche hardware vendors. Anyone remember the hype around big "cloud" storage boxes 10 years back?
- Havoc 3y agoGiven that their hardware is different I can kinda see how they don’t want to deal with supporting customers. > what do you mean I can’t just drop a CUDA docker image in?
- htrp 3y agoif you're a hardware startup that doesn't sell hardware, what are you?
- aleph_minus_one 3y ago> if you're a hardware startup that doesn't sell hardware, what are you? A hardware startup that sells cloud access to its hardware. :-)
- Havoc 3y agoHardware setup that produces superior hardware and extracts the benefit in house ?
- alted 3y agoCustom state-of-the-art silicon is ridiculously expensive. For a minimum 100 wafers = 10k chips, Groq may have paid $100M = $10k/chip purely in amortizing design costs. Chip design (software + engineer time) and fabrication setup (lithography masks) grow exponentially [1][2] with smaller nodes, e.g., maybe $100M for Groq's current 14nm chips to ~$500M for their planned 4nm tapeout. Once you reach mass production (>>1000 wafers, which have ~150 large chips each), wafers are $10k each. On top of this, it takes ~1 year to design then have prototypes made. (These same issues still exist on older slower nodes, albeit not as bad.) This could be reduced somewhat if chip design software were cheaper and margins were lower, but maybe 20% of this cost is due to fundamental manufacturing difficulty. (disclosure: I don't work with recent tech nodes myself; this is my best guess) [1] https://www.semianalysis.com/p/the-dark-side-of-the-semiconductor https://www.semianalysis.com/p/the-dark-side-of-the-semicond... [2] https://www.extremetech.com/computing/272096-3nm-process-node https://www.extremetech.com/computing/272096-3nm-process-nod...
- shrubble 3y agoThe report I read said that latest TSMC is 17K per wafer. How much less it is for 14nm I don't know.
- karma_pharmer 3y agoThe masks are the expensive part, not the wafers.
- thrtythreeforty 3y agoThey are both fabulously expensive.
- latchkey 3y ago> Custom state-of-the-art silicon is ridiculously expensive. Think about the amount of money being dumped into "AI" at this point. If you've got the technology and people to make stuff faster/better/cheaper, finding investors to dump money into your chip making business is probably not as hard as it was 2 years ago. Groq is making this change for other reasons than the expense of tapping out chips.
- mlazos 3y agoThe smoke and mirrors around groq are finally clearing. Truth is that their system is insanely expensive to maintain. hundreds (> 500 iirc) of chips to get wild tokens/s but the power and maintenance expense is crazy high for that number of chips. TCO just isn’t worth it
- ein0p 3y agoYou don't know that. For one thing, their silicon costs are going to be relatively cheap. It's an old reliable, 14nm process, and compared to even Google's TPU this is a relatively simple chip. For another they _could_ be putting all that silicon to a good use, and by all indications they are. Because there's far less local memory movement, and weights are distributed throughout the system, even this 14nm system could be energy efficient. 9/10ths of all power in a conventional system does not go towards compute - it's wasted in moving data back and forth. This is especially bad in transformers, which, because of their size, largely defeat the memory hierarchies the architects worked so hard to perfect. IOW, all your caches are useless and you're unnecessarily wasting 90% of your energy while also getting worse latency and worse throughput (due to memory bus bandwidth constraints). Oops. These folks seem to be offering something that nobody else does - a feasible, proven way to get out of jail free. I wish them all the success they can get, because all the other currently available architectures are largely unsuitable for high throughput transformer inference, and they work in spite, instead of because, of their design.
- mlazos 3y agoPeak H100 power consumption is 700W. Average power consumption of the groq card (from their own website) is 240W. With 576 chips it just doesn’t look good. How much is that millisecond perf gain worth it to end users? That said I think their arch is super interesting. I just think that demo was way too hype when the actual system is pretty impractical.
- ein0p 3y agoSo? They aren't performing the same computation. You can't compare the two. What you can compare is power draw at an equivalent tokens/sec on the same model for the entire system. But you don't have that number.
- geor9e 3y agoI don't understand why the comments are trash-talking Groq. They are the fastest LLM inference provider by a big margin. Why would they sell their hardware to any other company for any price? Keep it all for themselves and take over the market. 95% of my LLM requests go to Groq these days because it's 0.25 seconds round trip for a complete answer. In comparison, "Claude Instant" takes about 4 seconds. The other 5% of my requests go to Claude Opus and GPT-4, when I'm willing to wait an excruciating 5+ seconds for a better answer. I hate waiting. Latency is king. Groq wins.
- huac 3y agowhy don't you stream the results?
- tpetry 3y agoYou still have to wait for the end of the streamed response until you can continue with your task.
- fitzn 3y agoWhat open source model are you using when you hit groq? I just benchmarked some perf for some of my larger context window queries last week and groq's API took 1.6 seconds versus 1.8 to 2.2 for OpenAI GPT-3.5-turbo. So, it wasn't much faster. I almost emailed their support to see if I was doing something wrong. Would love to hear any details about your workload or the complexity of your queries.
- laserbeam 3y ago> 1.6 vs 1.8-2.2 seconds I believe certain companies would kill for 20% performance improvements on their main product.
- metadat 3y ago"kill", .. why would anyone kill for a fraction of a second in this case? Informed folks know that LLM hosters aren't raking in the big bucks. They're selling dreams and aspirations, and those are what's driving the funding.
- karma_pharmer 3y agoAnother casualty of AI KYC.
- ilaksh 3y agoI'm not able to get consistent replies from the API. It's lightening fast for like ten minutes and then starts freezing up for several seconds. I want to use it, but it's been very unreliable. I have been using Claude 3 and thinking about together.ai with Mixtral.
- QuadrupleA 3y agoSame, it's great when it's quick / available, but they seem underprovisioned for busy times and I often get long 10-30 second stalls.
- pha392 3y agoIMHO, Groq is being shadow acquired by Google
- vinay_ys 3y agoThis business model is bound to get attacked and suffer a painful exit soon. Here's why: First, the whole systems of chips architecture that everyone is talking about will solve for increasing overall SRAM available to keep more model state on super fast memory and avoid going to slow memory. Secondly, anyone serious about their data (enterprises) won't be okay with making API calls to Groq. Anyone serious about their data and have a lot scale (consumer internet) won't also be okay with making expensive API calls to Groq at scale. Their cloud is attractive only if I can use their API for experimentation toy apps to continue developing in this direction while the rest of the major industry players systems of chip architecture catches up and solves for SRAM size bottleneck and manufacturing process bottleneck, and once that's solved, I get more powerful compute for cheaper $$ to deploy on-prem. So, this cloud strategy is short-lived. I see another pivot on the horizon.
- sebastiennight 3y agoThe same has been said of OpenAI for a couple of years now (that they're just a platform to prototype on before moving on to open source models)... ... and yet, they're still leading the field. I think it's a bit early to think the field is getting commoditized yet.
- frozenport 3y ago>> won't be okay with making API calls to Groq Linked article: If customers come with requests for high volumes of chips for very large installations, Groq will instead propose partnering on data center deployment
- zachbee 3y agoTotally saw this one coming! [1] I think one major challenge they'll face is that their architecture is incredibly fast at running the ~10-100B parameter open-source models, but starts hitting scaling issues with state-of-the-art models. They need 10k+ chips for a GPT-4-class model, but their optical interconnect only supports a few hundred chips. [1] https://www.zach.be/p/why-is-everybody-talking-about-groq https://www.zach.be/p/why-is-everybody-talking-about-groq
- TamiDeines99 3y ago[dead]