2 ms·
That thread makes a bunch of assumptions that seem a bit dubious to me. We've known since Chinchilla that you don't need 175B parameters to get GPT-3 quality –
by moyix 4y ago
That thread makes a bunch of assumptions that seem a bit dubious to me. We've known since Chinchilla that you don't need 175B parameters to get GPT-3 quality – a 70B model can outperform GPT3 [1]. And his numbers assume the model is loaded into GPU memory in FP16 (175B*2 = 350GB), but people have shown you can quantize down to 8-bit (and in some cases 4 bit) with almost no performance loss. So in 8-bit precision with a 70B model you need ~70GB of VRAM, which you can get with two A6000s on a desktop (each 48GB).
And finally there are lots of other ways to get this down. Aside from quantization, people have also shown that you can do pruning – getting rid of many of the weights – again without much perf loss. You can also offload the weights to CPU RAM or an NVME and stream them in as needed [2]; it's slower but if you arrange things right the performance is not too bad. There are also ways to speed up inference using techniques like early exit [3], where you can skip running the whole model for some tokens that are easy to predict.
Overall it feels like within a year or two a combination of better quantization/pruning, improved understanding of how to train smaller LLMs, and hardware improvements will put inference for ChatGPT-style models within reach of the average user.
[1] https://towardsdatascience.com/a-new-ai-trend-chinchilla-70b-greatly-outperforms-gpt-3-175b-and-gopher-280b-408b9b4510 https://towardsdatascience.com/a-new-ai-trend-chinchilla-70b...
[2] https://github.com/FMInference/FlexGen https://github.com/FMInference/FlexGen
[3] https://ai.googleblog.com/2022/12/accelerating-text-generation-with.html https://ai.googleblog.com/2022/12/accelerating-text-generati...