9 ms·
Llama 3.1 405B now runs at 969 tokens/s on Cerebras Inference
- brcmthrowaway 2y agoSo out of all AI chip startups, Cerebras is probably the real deal
- gdiamos 2y agojust in time for their ipo
- ipsum2 2y agoIt got cancelled/postponed.
- icelancer 2y agoGroq is legitimate. Cerebras so far doesn't scale (wide) nearly as good as Groq. We'll see how it goes.
- hendler 2y agoGoogle TPUs, Amazon, a YC funded ASIC/FPGA company, a Chinese Co. all have custom hardware too that might scale well.
- throwawaymaths 2y agoHow exactly does groq scale wide well? Last I heard it was 9 racks!! to run llama-2 70b Which is why they throttle your requests
- pama 2y agoWell, Cerebras pretty much needs a data center to simply fit the 405B model for inference.
- throwawaymaths 2y agoI guess this just shows the insanity of venture led AI hardware hype and shady startup messaging practices
- easeout 2y agoHow does binning work when your chip is the entire wafer?
- shrubble 2y agoThey expect that some of the cores on the wafer will fail, so they have redundant links all throughout the chip, so they can seal off/turn off any cores that fail and still have enough cores to do useful work.
- why_only_15 2y agoMy understanding is that they mask off or otherwise disable a whole row+column of cores when one dies
- wtallis 2y agoThat's way too wasteful. Take a look at https://fuse.wikichip.org/news/3010/a-look-at-cerebras-wafer-scale-engine-half-square-foot-silicon-chip/2/ https://fuse.wikichip.org/news/3010/a-look-at-cerebras-wafer... and specifically the diagram https://fuse.wikichip.org/wp-content/uploads/2019/11/hc31-cerebras-mem-yield.png https://fuse.wikichip.org/wp-content/uploads/2019/11/hc31-ce... The fabric can effectively route signals diagonally to work around an individual defective core, with a displacement of one position for cores in the same row from that defect over to the nearest spare core. That's how they get away with a claimed "1–1.5%" of spare cores.
- kuprel 2y agoI wonder if Cerebras could generate video decent quality in real time
- jadbox 2y agoNot open beta until Q1 2025
- zackangelo 2y agoThis is astonishingly fast. I’m struggling to get over 100 tok/s on my own Llama 3.1 70b implementation on an 8x H100 cluster. I’m curious how they’re doing it. Obviously the standard bag of tricks (eg, speculative decoding, flash attention) won’t get you close. It seems like at a minimum you’d have to do multi-node inference and maybe some kind of sparse attention mechanism?
- yalok 2y agohow much memory do you need to run fp8 llama 3 70b - can it potentially fit 1 H100 GPU with 96GB RAM? In other words, if you wanted to run 8 separate 70b models on your cluster, each of which would fit into 1 GPU, how much larger your overall token output could be than parallelizing 1 model per 8 GPUs and having things slowed down a bit due to NVLink?
- zackangelo 2y agoIt’s been a minute so my memory might be off but I think when I ran 70b at fp16 it just barely fit on a 2x A100 80GB cluster but quickly OOMed as the context/kv cache grew. So if I had to guess a 96GB H100 could probably run it at fp8 as long as you didn’t need a big context window. If you’re doing speculative decoding it probably won’t fit because you also need weights and kv cache for the draft model.
- qingcharles 2y agoIt should work, I believe. And anything that doesn't fit you can leave on your system RAM. Looks like an H100 runs about $30K online for one. Are there any issues with just sticking one of these in a stock desktop PC and running llama.cpp?
- joha4270 2y ago> Are there any issues with just sticking one of these in a stock desktop PC and running llama.cpp? Cooling might be a challenge. The H100 has a heatsink designed to make use of the case fans. So you need a fairly high airflow through a part which is itself passive. On a server this isn't too big a problem, you have fans in one end and GPU's blocking the exit on the other end, but in a desktop you probably need to get creative with cardboard/3d printed shrouds to force enough air through it.
- WiSaGaN 2y agoI am wondering how much cost is needed for serving at such a latency. Of course for customers, static cost depends on the pricing strategy. But still, the cost really determines how widely this can be adopted. Is it only for those business that really need the latency, or this can be generally deployed.
- deleted 2y ago[deleted]
- ilaksh 2y agoMaybe it could become standard for everyone to make giant chips and use SRAM? How many SRAM manufacturers are there? Or does it somehow need to be fully integrated into the chip?
- AlotOfReading 2y agoSRAM is usually produced on the same wafer as the rest of the logic. SRAM on an external chip would lose many of the advantages without being significantly cheaper.
- YetAnotherNick 2y agoYes, the limiting factor for bandwidth is generally the number of pins which are not cheap and you can only have few 1000s in a chip. The absolute state of the art is 36 Gb/s/pin[1], and your $30 RAM could have 6 Gb/s/pin[2]. [1]: https://en.wikipedia.org/wiki/GDDR7_SDRAM https://en.wikipedia.org/wiki/GDDR7_SDRAM [2]: https://en.wikipedia.org/wiki/DDR5_SDRAM https://en.wikipedia.org/wiki/DDR5_SDRAM
- why_only_15 2y agothe cost is not the memory technology per se but primarily the wires. SRAM is fast because it's directly inside the chip and so the connections with the logic that does the work is cheap because it's close.
- danpalmer 2y agoI'm not sure if they're comparing apples to apples on the latency here. There are roughly three parts to the latency: the throughput of the context/prompt, the time spent queueing for hardware access, and the other standard API overheads (network, etc). From what I understand, several, maybe all, of the comparison services are not based on provisioned capacity, which means that the measurements include the queue time. For LLMs this can be significant. The Cerebras number on the other hand almost certainly doesn't have some unbounded amount of queue time included, as I expect they had guaranteed hardware access. The throughput here is amazing, but to get that throughput at a good latency for end-users means over-provisioning, and it's unclear what queueing will do to this. Additionally, does that latency depend on the machine being ready with the model, or does that include loading the model if necessary? If using a fine-tuned model does this change the latency? I'm sure it's a clear win for batch workloads where you can keep Cerebras machines running at 100% utilisation and get 1k tokens/s constantly.
- qeternity 2y agoEveryone presumes this is under ideal conditions...and it's incredible. It's bs=1. At 1,000 t/s. Of a 405B parameter model. Wild.
- colordrops 2y agoRight, I'd assume most LLM benchmarks are run on dedicated hardware.
- danpalmer 2y agoCerebras' benchmark is most likely under ideal conditions, but I'm not sure it's possible to test public cloud APIs under ideal conditions as it's shared infrastructure so you just don't know if a request is "ideal". I think you can only test these things across significant numbers of requests, and that still assumes that shared resource usage doesn't change much.
- qeternity 2y agoI'm not talking about that. I and many others here have spun up 8x or more H100 clusters and run this exact model. Zero other traffic. You won't come anywhere close to this.
- LASR 2y agoWhat you can do with current-gen models, along with RAG, multi-agent & code interpreters, the wall is very much model latency, and not accuracy any more. There are so many interactive experiences that could be made possible at this level of token throughput from 405B class models.
- TechDebtDevin 2y agoLike what..
- davidfiala 2y agoImagine increasing the quality and FPS of those AI-generated minecraft clones and experiencing even more high-quality, realtime AI-generated gameplay (yeah, I know they are doing textual tokens. but just sayin..) edit: context is https://oasisaiminecraft.com/ https://oasisaiminecraft.com/
- vineyardmike 2y agoYou can create massive variants of OpenAI's 01 model. The "Chain of Thought" tools become way more useful when you can get when you can iterate 100x faster. Right now, flagship LLMs stream responses back, and barely beat the speed a human can read, so adding CoT makes it really slow for human-in-the-loop experiences. You can really get a lot more interesting "thoughts" (or workflow steps, or whatever) when it can do more, without slowing down the human experience of using the tool. You can also get a lot fancier with tool-usage when you can start getting an LLM to use and reply to tools at a speed closer to the speed of a normal network service. I've never timed it, but I'm guessing current LLMs don't handle "live video" type applications well. Imagine an LLM you could actually video chat with - it'd be useful for walking someone through a procedure, or advanced automation of GUI applications, etc. AND the holy-grail of AI applications that would combine all of this - Robotics. Today, Cerebras chips are probably too power hungry for battery powered robotic assistants, but one could imagine a Star-Wars style robot assistant many years from now. You can have a robot that can navigate some space (home setting, or work setting) and it can see its environment and behavior, processing the video in real-time. Then, can reason about the world and its given task, by explicitly thinking through steps, and critically self-challenging the steps.
- germanjoey 2y agoPretty amazing speed, especially considering this is bf16. But how many racks is this using? The used 4 racks for 70B, so this, what, at least 24? A whole data center for one model?!
- aurareturn 2y agoEach Cerebras wafer scale chip has 44GB of SRAM. You need 972 GB of memory to run Llama 405b at fp16. So you need 22 of these. I assume they're using SRAM only to achieve this speed and not HBM.
- gdiamos 2y agoI'm so curious to see some multi-agent systems running with inference this fast.
- ipsum2 2y agoThere's no good open source agent models at the moment unfortunately.
- fillskills 2y agoNo mention of their direct competitor Groq?
- icelancer 2y agoI'm a happily-paying customer of Groq but they aren't competitive against Cerebras in the 405b space (literally at all). Groq has paying customers below the enterprise-level and actually serves all their models to everyone in a wide berth, unlike Cerebras who is very selective, so they have that going for them. But in terms of sheer speed and in the largest models, Groq doesn't really compare.
- guyomes 2y agoSambanova is not often mentioned either [0]. One of his co-founder is known as “father of the multi-core processor” [1]. [0]: https://sambanova.ai/ https://sambanova.ai/ [1]: https://en.wikipedia.org/wiki/Kunle_Olukotun https://en.wikipedia.org/wiki/Kunle_Olukotun
- bargle0 2y agoTheir hardware is cool and bizarre. It has to be seen in person to be believed. It reminds me of the old days when supercomputers were weird.
- IAmNotACellist 2y agoDon't leave us hanging, show us a weird computer!
- campers 2y agohttps://web.archive.org/web/20230812020202/https://www.youtube.com/watch?v=pzyZpauU3Ig https://web.archive.org/web/20230812020202/https://www.youtu...
- aurareturn 2y agoNormally, I don't think 1000 tokens/s is that much more useful than 50 tokens/s. However, given that CoT makes models a lot smarter, I think Cerebras chips will be in huge demand from now on. You can have a lot more CoT runs when the inference is 20x faster. Also, I assume financial applications such as hedge funds would be buying these things in bulk now.
- deadmutex 2y ago> Also, I assume financial applications such as hedge funds would be buying these things in bulk now. Please elaborate.. why?
- aurareturn 2y agoI'm assuming hedge funds are using LLMs to dissect information from company news, SEC reports as soon as possible then make a decision on trading. Having faster inference would be a huge advantage.
- dgfitz 2y agoHoly bananas, the title alone is almost its own language.
- maryndisouza 2y ago[flagged]
- arthurcolle 2y agoDamn that's a big model and that's really fast inference.
- owenpalmer 2y agoThe fact that such a boost is possible with new hardware, I wonder what the ceiling is for improving performance for training via hardware as well.
- bufferoverflow 2y agoThe ultimate solution would be to convert an LLM to a pure ASIC. My guess is that would 10X the performance. But then it's a very very expensive solution.
- why_only_15 2y agoWhy would converting a specific LLM to an ASIC help you? LLMs are like 99% matrix multiplications by work and we already have things that amount to ASICs for matrix multiplications (e.g. TPU) that aren't cheaper than e.g. H100
- mikewarot 2y agoAn ASIC could have all of the weights baked into the design, completely eliminating the von Neumann bottleneck that plagues computation. They are inherently parallel, so you might be able to get a token per clock cycle. A billion tokens per second opens quite a few possibilities. It could also eliminate all of the multiplication or addition of bits that are 0 from the design, making each multiply smaller by 50 percent silicon area, on average. However, an ASIC is a speculation that all the design tools work. It may require multiple rounds to get it right.
- ryao 2y agoI doubt you could have a token per clock cycle unless it is very low clocked. In practice, even dedicated hardware for matrix-matrix multiplication does not perform the multiplication in a single clock cycle. Presumably, the circuit paths would be so large that you would need to have a very slow clock to make that work, and there are many matrix multiplications done per token. Furthermore, things are layered and must run through each layer. Presumably if you implement this you would aim for 1 layer per clock cycle, but even that seems like it would be quite long as far as circuit paths go. I have some local code running llama 3 8B and matrix multiplications in it are being done by 2D matrices with dimensions ranging from 1024 to 4096. Let’s just go with a nice 1024x1024 matrix and do matrix-vector multiplication, which is the minimum needed to implement llama3. That is 1048576 elements. If you try to do matrix-vector multiplication in 1 cycle, you will need 1048576 fmadd units. I am by no means a chip designer, so I asked ChatGPT to estimate how many transistors are needed for a bf16 fmadd unit. It said 100,000 to 200,000. Let’s go with 100,000 transistors per unit. Thus to implement a single matrix multiplication according to your idea, we would need over 100 billion transistors, and this is only a small part of the llama 3 8b model’s calculations. You would probably be well into the trillions of transistors if you implemented all of it in an ASIC and did 1 layer per cycle (don’t even think of 1 token per cycle). For reference, Nvidia’s H100 has 80 billion transistors. The CSE-3 has 4 trillion transistors and I am not sure if even that would be enough. It is a nice idea, but I do not think it is feasible with current technology. That said, I do like your out of box thinking. This might be a bit too far out of the box, but there is probably a middle ground somewhere.
- xwww 2y agoTransistor(GPU)-> Integrated Circuit (WSE-3)
- qwertox 2y agoI'd like to see a tokens / second / watt comparison.
- shreezus 2y agoThis is seriously impressive performance. I think there's a high probability Nvidia attempts to acquire Cerebras.
- gorkempacaci 2y agoThey're considering an IPO. I'd say an acquisition is unlikely. Even then, they'd be worth more to Facebook or MS.
- szundi 2y agoNo, they would make a capital infusion on paper and then make Cerebras buy more hw from that money on paper, thus showing huge revenues on Nvidia books. Makes sense… or whatever
- dustypotato 2y agoThey're making custom chips right? Why would Cerebras buy hardware from Nvidia?
- gorkempacaci 2y agonvidia hates this one little trick
- zurfer 2y agoI laughed and upvoted, but if anything I bet they put their best people on it to replicate this offering. What I take away from this is: we are just getting started. I remember in 2023 begging OpenAI to give us more than 7 tokens/second on GPT-4.
- ryao 2y agoNvidia’s target is performance across concurrent users and they are likely already outperforming Cerebras there as far as costs are concerned. They have no reason to try to beat the single user performance of this.
- frogfish 2y agoGenuinely curious and willing to learn: what are the different inference approaches broadly? Is there any difference in the approach between Cerebras and simplismart.ai which claims to be the fastest?
- leobg 2y agoCerebras features in the internal OpenAI emails that recently came out. One example: Ilya Sutskever to Elon Musk, Sam Altman, (cc: Greg Brockman, Sam Teller, Shivon Zilis) - Sep 20, 2017 2:08 PM > In the event we decide to buy Cerebras, my strong sense is that it'll be done through Tesla. But why do it this way if we could also do it from within OpenAI?
- adhambadr 2y agois it just me or isn't the most important contender in speed, Groq, missing from the comparison ? not sure why does it matter to put azure there, no one uses it for speed.
- sumedh 2y agoThey have a waitlist for trying their API. You have to be a but skeptical when a company makes claims but does not offer their services to buy.
- perfobotto 2y agoTo be clear a cerebras chip is consuming a whole wafer and has only 44 GB of SRAM on it. To fit a 405B model in bf16 precision (excluding kv cache and activation memory usage) you need 19 of these “chips” (and the requirement will grow as the sequence length increases for the kvcache). Looking online it seems on one wafer one can fit between 60 to 80 H100 chips, so it’s equivalent to using >1500 H100 using wafer manufacturing cost as a metric
- latchkey 2y agoThis gets tons of press and discussion here on HN, but frankly AMD has a better overall product with the upcoming MI325x [0]. I love to see the development and activity, but companies like Cerebras are trying to compete on a single usecase and doing a poor job of it because they can only offer a tightly controlled API. Ask yourself how much capex + power/space/cooling (opex) it requires to run that model (and how many people it can really serve) and then compare that against what AMD is offering. [0] https://www.amd.com/en/products/accelerators/instinct/mi300/mi325x.html https://www.amd.com/en/products/accelerators/instinct/mi300/...
- deleted 2y ago[deleted]
- KETpXDDzR 2y agoI once saw the founder with a wafer sitting in an In n Out in the bay area. I almost worked for them, but they were short on funding. IMO the cost efficient for Cerebras "chips" is still a problem. I can't imagine how many defects they have on each one.