7 ms·
I just want to say that this is one of the most impressive tech demos I’ve ever seen in my life, and I love that it’s truly an open demo that anyone can try wit
by eigenvalue 3y ago
I just want to say that this is one of the most impressive tech demos I’ve ever seen in my life, and I love that it’s truly an open demo that anyone can try without even signing up for an account or anything like that. It’s surreal to see the thing spitting out tokens at such a crazy rate when you’re used to watching them generate at one less than one fifth that speed. I’m surprised you guys haven’t been swallowed up by Microsoft, Apple, or Google already for a huge premium.
- RockyMcNuts 3y agook... why tho? genuinely ignorant and extremely curious. what's the TFLOPS/$ and TFLOPS/W and how does it compare with Nvidia, AMD, TPU? from quick Googling I feel like Groq has been making these sorts of claims since 2020 and yet people pay a huge premium for Nvidia and Groq doesn't seem to be giving them much of a run for their money. of course if you run a much smaller model than ChatGPT on similar or more powerful hardware it might run much faster but that doesn't mean it's a breakthrough on most models or use cases where latency isn't the critical metric?
- RecycledEle 3y agoIf I understand correctly, each chip has 200 MB of RAM, so it takes racks to run a single LLM. That does not sound like progress to me. We need single PCIe boards with dozens or hundreds of GB of RAM and processors that handle it well.
- tome 3y agoReally glad you like it! We've been working hard on it.
- jonplackett 3y agoIs this useful for training as well as running a model. Or is this approach specifically for running an already-trained model faster?
- frozenport 3y agoIn principle, training is basically the same as running inference but iteratively, in practice training would use a different software stack.
- robrenaud 3y agoTraining requires a lot more memory to keep gradients + gradient stats for the optimizer, and needs higher precision weights for the optimization. It's also much more parallelizable. But inference is kind of a subroutine of training.
- tome 3y agoCurrently graphics processors work well for training. Language processors (LPUs) excel at inference.
- lokimedes 3y agoThe speed part or the being swallowed part?
- tome 3y agoThe speed part. We're not interested in being swallowed. The aim is to be bigger than Nvidia in three years :)
- timomaxgalvin 3y agoSure, but the responses are very poor compared to MS tools.
- brcmthrowaway 3y agoI have it on good authority Apple was very closing to acquiring Groq
- baq 3y agoIf this is true, expect a call from the SEC...
- 317070 3y agoEven if it isn't true. Disclosing inside information is illegal, _even if it is false and fabricated_, if it leads to personal gains.
- varelse 3y ago[dead]
- KRAKRISMOTT 3y agoYou have to prove the OP had personal gains. If he's just a troll, it will be difficult.
- frognumber 3y agoYou also have to be an insider. If I go to a bar, and overhear a pair of Googlers discussing something secret and overhear it, I can: 1) Trade on it. 2) Talk about it. Because I'm not an insider. On the other hand, if I'm sleeping with the CEO, I become an insider. Not a lawyer. Above is not legal advice. Just a comment that the line is much more complex, and talking about a potential acquisition is usually okay (if you're not under NDA).
- throwawayurlife 3y agoIt doesn't matter if you overheard it at a bar or if you're just some HN commenter posting completely incorrect legal advice; the law prohibits trading on material nonpublic information. I would pay a lot to see you try your ridiculous legal hokey-pokey on how to define an "insider."
- elorant 3y agoPerplexity Labs also has an open demo of Mixtral 8x7b although it's nowhere near as fast as this. https://labs.perplexity.ai/ https://labs.perplexity.ai/
- vitorgrs 3y agoPoe has a bunch of them, including Groq as well!
- larodi 3y agowhy sell? it would be much more delightful to beat them on their own game?