5 ms·
Output from Cerebras with GPT model is 750 tokens per second. Don’t blink. (Chatjimmy has 14,200 TPS.)
by sixtyj 2mo ago
Output from Cerebras with GPT model is 750 tokens per second.
Don’t blink.
(Chatjimmy has 14,200 TPS.)
- tomrod 2mo agoChatJimmy is a much smaller model and, AFAIK, has no reasoning capability. Absolutely insane raw speed, like a supercar, while Sol is more like a freight truck.
- sixtyj 2mo agoAt such output speed, I wouldn’t expect reasoning. (But I didn’t know it, thanks.) 700 TPS with reasoning is awesome and it speeds things up. Cerebras as public traded company is worth keeping an eye what they produce.
- msdz 2mo ago> At such output speed, I wouldn’t expect reasoning. As the sibling comment to yours mentioned, if they had a reasoning model “hardware-ified” onto a custom chip (as is their plan for IIRC this or next year, a new ASIC), it’d output fast decode speeds for the regular output as well as reasoning sections. Both would be ≈equally fast.
- senordevnyc 2mo agoYeah, I thought reasoning was literally just chain of thought in the output token stream, with the model itself adding delimiters to indicate what part of the output is internal reasoning, and what part is an answer to the user. Is that wrong?
- beering 2mo agoYou are right, reasoning is unrelated to tokens per second.
- msdz 2mo agoNo, as the sibling comment mentioned, your understanding was correct there. What’s more, the only technical difference in speeds could be, and likely also is with the HC1 chip, between prefill (prompt processing) and decode (text generation) speeds. I don’t know whether it’s the case with Taalas’ chip, but in the “software-based” LLMs we typically see and use so far, those two stages hit different parts of a computer (processing/compute-bound vs. memory/bandwidth-bound).
- notfromhere 2mo agoAnything will be fast if you etch it straight to silicon
- dzhiurgis 2mo agoThe knowledge of ChatJimmy is terrible. Even Qwen on my iPhone is better.
- headPoet 2mo agoChatJimmy isn't a model, it's Llama 3.1 8B hardwired into silicon. The point isn't to be a good llm, but to showcase the speedup that's possible
- senordevnyc 2mo agoHaha, at first I thought you meant that the knowledge of the existence of an LLM that’s so fast is terrible because it’s ruined every other LLM for you!
- walrus01 2mo agowell, yeah, it's based on a 2+ year old tiny model. It's very much an alpha proof of concept that they can perma-bake an LLM into silicon. https://huggingface.co/meta-llama/Llama-3.1-8B https://huggingface.co/meta-llama/Llama-3.1-8B
- mips_avatar 2mo agoUnfortunately AMD bought them, so I don't think we will get to see another release from them.
- sixtyj 2mo agoAha, thanks, that’s fresh; press release is from Aug 6 https://ir.amd.com/news-events/press-releases/detail/1296/amd-acquires-taalas-to-advance-compute-solutions-for-rapidly-growing-ai-inference-market https://ir.amd.com/news-events/press-releases/detail/1296/am...
- perching_aix 2mo agoNever heard of it before, that's fucking insane. Apparently they baked the Llama 3.1 8B model weights [0] into silicon (the actual hardware is called Taalas HC1). I guess for the trillion parameter models this would not scale due to cost? Imagine buying GPT 6 in the form of a PCI-E card, pulling these speeds, with up to 120 cct agent sessions. It'd be beyond wild. [0] the weights are also using some cut down small format, but HC2 will have regular FP4 supposedly, and support for 20B params on one die
- TacticalCoder 2mo ago> Never heard of it before, that's fucking insane. They've been acquired by AMD. Those saying the model sucks are completely missing the point: it was a proof-of-concept. The question is: what happens to a model like Anthropic's Fable 5 that does, what, 70 tokens/s (and requires lots of output tokens) when the latest open-weights model is etched on silicon and does 14 000 tokens/s? Shall the better model still have the upper hand or will the raw speed compensate?
- JohnBooty 2mo agoShall the better model still have the upper hand or will the raw speed compensate? At 14,000 tokens/sec there's just so much ridiculous stuff that might be possible. Let's assume that this POC proves they can take the next step, and can eventually etch a capable ~27B model into silicon. Let's call it Fred. Ralph loops automatically get real real interesting again. 200x the iteration speed. This is such a clear win I feel like there's hardly anything to talk about. Instead of one stubborn iterating idiot, you could have dozens of idiots competing in parallel, genetic algorithm style. The other common orchestration pattern I see is "big model for planning, small parallel subagents implementing, big model reviewing" Today it's Sol dispatching a handful of Luna subagents. Tomorrow maybe it's Sol dispatching as many Fred subagents as it could possibly want. But what patterns have we not even thought about yet in a world where subagents are 200x faster/cheaper? What if instead of dispatching single Haiku/Luna/etc subagents, we dispatched "teams" of Fred agents? Maybe each team is 8 Freds. Five come up with competing ideas and the other three vote on a winner. Or what if they were heterogenous teams? One Luna and a bunch of Freds. What if instead of a two-tier orchestration system (Sol->Luna) it was three-tier or n-tier? (Sol->Luna->Fred->...Fred) Those ideas overlap a bit, and crazy shit like Gas Town has already explored even wilder ideas I guess. But man, 14000 tk/sec opens up so much stuff.