6 ms·
This is astonishingly fast. I’m struggling to get over 100 tok/s on my own Llama 3.1 70b implementation on an 8x H100 cluster. I’m curious how they’re doing it
by zackangelo 2y ago
This is astonishingly fast. I’m struggling to get over 100 tok/s on my own Llama 3.1 70b implementation on an 8x H100 cluster.
I’m curious how they’re doing it. Obviously the standard bag of tricks (eg, speculative decoding, flash attention) won’t get you close. It seems like at a minimum you’d have to do multi-node inference and maybe some kind of sparse attention mechanism?
- yalok 2y agohow much memory do you need to run fp8 llama 3 70b - can it potentially fit 1 H100 GPU with 96GB RAM? In other words, if you wanted to run 8 separate 70b models on your cluster, each of which would fit into 1 GPU, how much larger your overall token output could be than parallelizing 1 model per 8 GPUs and having things slowed down a bit due to NVLink?
- zackangelo 2y agoIt’s been a minute so my memory might be off but I think when I ran 70b at fp16 it just barely fit on a 2x A100 80GB cluster but quickly OOMed as the context/kv cache grew. So if I had to guess a 96GB H100 could probably run it at fp8 as long as you didn’t need a big context window. If you’re doing speculative decoding it probably won’t fit because you also need weights and kv cache for the draft model.
- qingcharles 2y agoIt should work, I believe. And anything that doesn't fit you can leave on your system RAM. Looks like an H100 runs about $30K online for one. Are there any issues with just sticking one of these in a stock desktop PC and running llama.cpp?
- joha4270 2y ago> Are there any issues with just sticking one of these in a stock desktop PC and running llama.cpp? Cooling might be a challenge. The H100 has a heatsink designed to make use of the case fans. So you need a fairly high airflow through a part which is itself passive. On a server this isn't too big a problem, you have fans in one end and GPU's blocking the exit on the other end, but in a desktop you probably need to get creative with cardboard/3d printed shrouds to force enough air through it.
- modeless 2y agoCerebras is a chip company. They are not using GPUs. Their chip uses wafer scale integration which means it's the physical size of a whole wafer, dozens of GPUs in one. They have limited memory on chip (all SRAM) and it's not clear how much HBM bandwidth they have per wafer. It's a completely different optimization problem than running on GPU clusters.
- why_only_15 2y agothey have about 125GB/s of off-chip bandwidth
- saagarjha 2y agoDo they just not do HBM at all or
- why_only_15 2y agoI'm not too up to date but as I recall there are a lot of weirdnesses because of how big their chip is (e.g. thermal expansion being a problem). I believe they have a single giant line in the middle of the chip for this reason. maybe this makes HBM etc. hard? certainly their chip would be more appealing if they cut down the # of cores by 10x, added matrix units and added HBM but looks like they're not going to go this way.
- ryao 2y agoThey do not use HBM. Offchip memory is accessible at 150GB/sec.
- danpalmer 2y agoCerebras makes CPUs with ~1 million cores, and they're inferring on that not on GPUs. It's an entirely different architecture which means no network involved. It's possible they're doing this significantly from CPU caches rather than HBM as well. I recommend the TechTechPotato YouTube videos on Cerebras to understand more of their chip design.
- zackangelo 2y agoAh, makes a lot more sense now.
- StrangeDoctor 2y agoalso the WSE3 pulls 15kw. https://www.eetimes.com/cerebras-third-gen-wafer-scale-chip-doubles-performance https://www.eetimes.com/cerebras-third-gen-wafer-scale-chip-... but 8x h100 are ~2.6-5.2kw (I get conflicting info, I think based on pice vs smx) so anywhere between roughly even and up to 2x efficient.
- swyx 2y ago> TechTechPotato YouTube videos on Cerebras https://www.youtube.com/@TechTechPotato/search?query=cerebras https://www.youtube.com/@TechTechPotato/search?query=cerebra... for anyone also looking. there are quite a lot of them.
- accrual 2y agoI hope we can buy Cerebras cards one day. Imagine buying a ~$500 AI card for your desktop and having easy access to 70B+ models (the price is speculative/made up).
- chessgecko 2y agoOne day is doing some heavy heavy lifting here, we’re currently off by ~3-4 orders of magnitude…
- accrual 2y ago
- parsimo2010 2y agoThey are doing it with custom silicon with several times more area than 8x H100s. I’m sure they are doing some sort of optimization at execution/runtime, but the primary difference is the sheer transistor count. https://cerebras.ai/product-chip/ https://cerebras.ai/product-chip/
- coder543 2y agoTo be specific, a single WSE-3 has the same die area as about 57 H100s. It's a big chip.
- deleted 2y ago[deleted]
- cma 2y agoIt is worth splitting out the stacked memory silicon layers on both too (if Cerebras is set up with external DRAM memory). HBM is over 10 layers now so the die area is a good bit more than the chip area, but different process nodes are involved.
- tomrod 2y agoAmazing!
- deleted 2y ago[deleted]
- boroboro4 2y agoThere are two big tricks: their chips are enormous, and they use sram as their memory, which is vastly faster than hbm ram being used by GPUs. In fact this is main reason it’s so fast. Groq has the speed because of the same reason.
- mmaunder 2y agoNah. Try vLLM and 405B FP8 on that hardware. And make sure you’re benchmarking with some concurrency for max TPS.
- zackangelo 2y agoRelated recent discussion on twitter: https://x.com/Teknium1/status/1858987850739728635 https://x.com/Teknium1/status/1858987850739728635 Looks like other folks get 80 tok/s with max batch size, that's surprising to me but vLLM is definitely more optimized than my implementation.
- simonw 2y agoThey have a chip the size of a dinner plate. Take a look at the pictures: https://cerebras.ai/product-chip/ https://cerebras.ai/product-chip/
- pram 2y agoI'd love to see the heatsink for this lol
- futureshock 2y agoThey call it the “engine block”! https://www.servethehome.com/a-cerebras-cs-2-engine-block-bare-on-the-sc22-show-floor/ https://www.servethehome.com/a-cerebras-cs-2-engine-block-ba...
- Aeolun 2y ago21 petabytes per second. Can push the whole internet over that chip xD
- why_only_15 2y agoThe number for that is I believe 1 terabit or 125GB/s -- 21 petabytes is the speed from the SRAM (~registers) to the cores (~ALU) for the whole chip. It's not especially impressive for SRAM speeds. The impressive thing is that they have an enormous amount of SRAM
- KeplerBoy 2y agoThat's their on chip cache bandwidth. Usually that stuff isn't even measured in bandwidth but latency.
- ekianjo 2y agowhat kind of yield do they get on that size?
- bufferoverflow 2y agoIt's near 100%. Discussed here: https://youtu.be/f4Dly8I8lMY?t=95 https://youtu.be/f4Dly8I8lMY?t=95
- hendler 2y agoCheck out BaseTen for performant use of GPUs
- mikewarot 2y agoImagine if you could take Llama 3.1 405B and break it down to a tree of logical gates, optimizing out all the things like multiplies by 0 in one of the bits, etc... then load it into a massive FPGA like chip that had no von Neumann bottleneck, was just pure compute without memory access latency with a conservative 1 Ghz clock rate. Such a system would be limited by the latency across the reported 126 layers worth of math involved, before it could generate the next token, which might be as much as 100 uSec. So it would be 10x faster, but you could have thousands of other independent streams pipelined through in parallel because you'd get a token per clock cycle out the end. In summary, 1 Gigatoken/second, divided into 100,000 separate users each getting 10k tokens/second. This is the future I want to build.
- seangrogg 2y agoI'm actively trying to learn how to do exactly this, though I'm just getting started with FPGAs now so probably a very long range goal.
- ryao 2y agoThere is not enough memory attached to FPGAs to do this. Some FOGAs come with 16GB of HBM attached, but that is not enough and the bandwidth provided is not as high as it is on GPUs. You would need to work out how to connect enough memory chips simultaneously to get high bandwidth and enough capacity in order for a FPGA solution to be performance competitive with a GPU solution.
- mikewarot 2y agoInstead of separate memory/compute, I propose to fuse them.
- jacobgorm 2y agoSee Convolutional Differentiable Logic Gate Networks https://arxiv.org/abs/2411.04732 https://arxiv.org/abs/2411.04732 , which is a small step in that direction.