4 ms·
Afaik you actually need much more. About 400GB just to load the trained model according to this tweet thread: https://twitter.com/tomgoldsteincs/status/16001969
by dxuh 4y ago
Afaik you actually need much more. About 400GB just to load the trained model according to this tweet thread: https://twitter.com/tomgoldsteincs/status/1600196981955100694?lang=de https://twitter.com/tomgoldsteincs/status/160019698195510069...
I am not quite sure how reliable the source is, but it makes sense that you at least need to store the 175 billion parameters that define the model in VRAM. I know that for a short-ish while GPUDirect storage is a thing, so that could help for sure, but it would definitely impact execution time as well.
- moyix 4y agoThat thread makes a bunch of assumptions that seem a bit dubious to me. We've known since Chinchilla that you don't need 175B parameters to get GPT-3 quality – a 70B model can outperform GPT3 [1]. And his numbers assume the model is loaded into GPU memory in FP16 (175B*2 = 350GB), but people have shown you can quantize down to 8-bit (and in some cases 4 bit) with almost no performance loss. So in 8-bit precision with a 70B model you need ~70GB of VRAM, which you can get with two A6000s on a desktop (each 48GB). And finally there are lots of other ways to get this down. Aside from quantization, people have also shown that you can do pruning – getting rid of many of the weights – again without much perf loss. You can also offload the weights to CPU RAM or an NVME and stream them in as needed [2]; it's slower but if you arrange things right the performance is not too bad. There are also ways to speed up inference using techniques like early exit [3], where you can skip running the whole model for some tokens that are easy to predict. Overall it feels like within a year or two a combination of better quantization/pruning, improved understanding of how to train smaller LLMs, and hardware improvements will put inference for ChatGPT-style models within reach of the average user. [1] https://towardsdatascience.com/a-new-ai-trend-chinchilla-70b-greatly-outperforms-gpt-3-175b-and-gopher-280b-408b9b4510 https://towardsdatascience.com/a-new-ai-trend-chinchilla-70b... [2] https://github.com/FMInference/FlexGen https://github.com/FMInference/FlexGen [3] https://ai.googleblog.com/2022/12/accelerating-text-generation-with.html https://ai.googleblog.com/2022/12/accelerating-text-generati...