3 ms·
Because of memory bandwidth. H100 has 3350gB/s of bandwidth, more gpus will give you more memory but not bandwidth. If you load 175b parameters in 8bit then you
by slimesli 3y ago
Because of memory bandwidth.
H100 has 3350gB/s of bandwidth, more gpus will give you more memory but not bandwidth.
If you load 175b parameters in 8bit then you can get theoretically
3350/175=19 tokens/second.
In MoE you need to process only one expert at a time so sparse 8x220b model would be only slightly slower than dense 220b model.
- fancyfredbot 3y agoOkay, memory bandwidth certainly matters, but 19 tokens a second is not some fundamental lower limit on the speed of a language model and so this doesn't really explain why the limit would be 220b rather than say 440b or 800b?
- slimesli 3y agoIt's not a fundamental limit. Google palm had 540B parameters as dense model. But it's a practical limit because models with over 1T would be extremely slow even on newest gpus. Even now, OpenAI has limit of 25 messages. You can read more here: https://bounded-regret.ghost.io/how-fast-can-we-perform-a-forward-pass/ https://bounded-regret.ghost.io/how-fast-can-we-perform-a-fo...
- fancyfredbot 3y agoI'm not trying to say memory bandwidth isn't a bottleneck for very large models. I'm wondering why he picked 220b which is weirdly specific. (To be honest although I completely agree the costs would be very high, I think there are people who would pay for and wait for answers at seconds or even minutes per token if they were good enough, so not completely sure I even agree it's a practical limit)