5 ms·
« We can’t really make models bigger than 220B parameters » Can someone explains why?
by it_citizen 3y ago
« We can’t really make models bigger than 220B parameters »
Can someone explains why?
- sp332 3y agoIt doesn't fit in VRAM.
- hesdeadjim 3y agoI’ve been a bit surprised that Nvidia hasn’t gone to extreme lengths to fit 1tb of memory on a card just for this reason.
- andrewstuart2 3y agohttps://nvidianews.nvidia.com/news/nvidia-announces-dgx-gh200-ai-supercomputer https://nvidianews.nvidia.com/news/nvidia-announces-dgx-gh20... I think they _are_ going pretty extreme now.
- samplatt 3y agoOfftopic, but as a VR gamer that article just made me very sad. I was really hoping to see NVidia produce some decent cards in the near future, but looks like their main revenue is really going to be gargantuan number-crunchers. They'll likely only keep increasing the VRAM of gaming cards by arbitrarily-small numbers once every few years :-(
- deleted 3y ago[deleted]
- speedgoose 3y agoYes, the future of VR gaming looks closer to the Sony Playstation or even the Apple Vision than NVIDIA's products.
- Tepix 3y agoGaming seems a lot less important than AI, in particular the graphical fidelity. Even games with crappy graphics can be fun. Crappy AI, not so much.
- Chamix 3y agoThe issue, as pointed above, is primarily bandwidth (at inference), not addressable memory. Put simply, the best bandwidth stack we currently have is on-package HBM -> NVLink, -> Mellanox InfiniBand, and for inference speed you really can't leave the NVLink bandwidth (read, 8x DGX pod) for >100b parameters. And stacking HBM dies is much harder (read, expensive) than GDDR dies which is harder than DDR etc. Cost aside, HMB dies themselves aren't getting significantly denser anytime soon, and there just simply isn't enough package space with current manufacturing methods to pack a significantly increased number of dies on the gpu. So I suspect the major hardware jumps will continue to be with NVLink/NVSwitch. Nvlink 4 + NVSwitch 3 actually already allows for up 256x GPUs https://resources.nvidia.com/en-us-grace-cpu/nvidia-grace-hopper https://resources.nvidia.com/en-us-grace-cpu/nvidia-grace-ho... ; increased numbers of links will let ever increasing numbers of GPUs pool with sufficient bandwidth for inference on larger models. As already mentioned, see this HN post about the GH200 https://news.ycombinator.com/item?id=36133226 https://news.ycombinator.com/item?id=36133226, which has some further discussion about the cutting edge of bandwidth for Nvidia DGX and Google TPU pods.
- hesdeadjim 3y agoThanks for this info!
- slimesli 3y agoBecause of memory bandwidth. H100 has 3350gB/s of bandwidth, more gpus will give you more memory but not bandwidth. If you load 175b parameters in 8bit then you can get theoretically 3350/175=19 tokens/second. In MoE you need to process only one expert at a time so sparse 8x220b model would be only slightly slower than dense 220b model.
- fancyfredbot 3y agoOkay, memory bandwidth certainly matters, but 19 tokens a second is not some fundamental lower limit on the speed of a language model and so this doesn't really explain why the limit would be 220b rather than say 440b or 800b?
- slimesli 3y agoIt's not a fundamental limit. Google palm had 540B parameters as dense model. But it's a practical limit because models with over 1T would be extremely slow even on newest gpus. Even now, OpenAI has limit of 25 messages. You can read more here: https://bounded-regret.ghost.io/how-fast-can-we-perform-a-forward-pass/ https://bounded-regret.ghost.io/how-fast-can-we-perform-a-fo...
- fancyfredbot 3y agoI'm not trying to say memory bandwidth isn't a bottleneck for very large models. I'm wondering why he picked 220b which is weirdly specific. (To be honest although I completely agree the costs would be very high, I think there are people who would pay for and wait for answers at seconds or even minutes per token if they were good enough, so not completely sure I even agree it's a practical limit)