5 ms·
You can chat with these new models at ultra-low latency at groq.com. 8B and 70B API access is available at console.groq.com. 405B API access for select customer
by foundval 2y ago
You can chat with these new models at ultra-low latency at groq.com. 8B and 70B API access is available at console.groq.com. 405B API access for select customers only – GA and 3rd party speed benchmarks soon.
If you want to learn more, there is a writeup at https://wow.groq.com/now-available-on-groq-the-largest-and-most-capable-openly-available-foundation-model-to-date-llama-3-1-405b/ https://wow.groq.com/now-available-on-groq-the-largest-and-m....
(disclaimer, I am a Groq employee)
- sagz 2y ago405B is already being served on WhatsApp! https://ibb.co/kQ2tKX5 https://ibb.co/kQ2tKX5
- Workaccount2 2y agoHow do you get that option?
- e12e 2y agoAnd available via poe: https://poe.com/s/LCAyUbAgUx8UcVMhM3Re https://poe.com/s/LCAyUbAgUx8UcVMhM3Re
- geepytee 2y agoWe also added Llama 3.1 405B to our VSCode copilot extension for anyone to try coding with it. Free trial gets you 50 messages, no credit card required - https://double.bot https://double.bot (disclaimer, I am the co-founder)
- noble-lombax 2y agowould be great if there was a page showing benchmarks compared to other auto completion tools
- quotemstr 2y agoGroq's TSP architecture is one of the weirder and more wonderful ISAs I've seen lately. The choice of SRAM in fascinating. Are you guys planning on publishing anything about how you bridged the gap between your order-hundreds-megabytes SRAM TSP main memory and multi-TB model sizes?
- foundval 2y agoThere is a lot out here. I gave a seminar about the overall approach recently, abstract: https://shorturl.at/E7TcA https://shorturl.at/E7TcA, recording: https://shorturl.at/zBcoL https://shorturl.at/zBcoL. This two-part AMA has a lot more detail if you're already familiar with what we do: https://www.youtube.com/watch?v=UztfweS-7MU https://www.youtube.com/watch?v=UztfweS-7MU https://www.youtube.com/watch?v=GOGuSJe2C6U https://www.youtube.com/watch?v=GOGuSJe2C6U
- quotemstr 2y agoThanks!
- listic 2y agoJust checked it out. Is pay-as-you-go API access available at all? It says 'Coming Soon' https://console.groq.com/settings/billing https://console.groq.com/settings/billing
- Alifatisk 2y agoI think you answered it yourself? It’s coming soon, so it is not available now, but soon.
- senko 2y agoIt's been coming soon for a couple of months now, meanwhile Groq churns out a lot of other improvements, so to an outsider like me it looks like it's not terribly high on their list of priorities. I'm really impressed by what (&how) they're doing and would like to pay for a higher rate limit, or failing that at least know if "soon" means "weeks" or "months" or "eventually". I remember TravisCI did something similar back in the day, and then Circle and GitHub ate their lunch.
- weberer 2y agoI've found Bedrock to be nice with pay-as-you-go, but they take a long time to adopt new models.
- d13 2y agoAnd twice as expensive in comparison to the source providers’ APIs
- d13 2y agoAt what quantisation are you running these?
- serverlessmania 2y agoYou can chat with all these models for free and ultra-low latency using this hosted website https://nat.dev/chat https://nat.dev/chat for free by GitHub Founder