3 ms·
Could someone with experience explain: what's the theoretical minimum hardware requirement for llama 7B, 15B, etc, that still provides output on the order of <1
by 2bitencryption 4y ago
Could someone with experience explain: what's the theoretical minimum hardware requirement for llama 7B, 15B, etc, that still provides output on the order of <1sec/token?
It seems like we can pull some tricks, like using F16, and some kind of quantization, etc.
At the end of the day, how much overhead is left that can be reduced? What can I expect to have running on 16gb ram with a 3080 and a midrange AMD processor?
- TaylorAlexander 4y agoWell I was able to run the original code with the 7B model on 16GB vram: https://news.ycombinator.com/item?id=35013604 https://news.ycombinator.com/item?id=35013604 The output I got was underwhelming, though I did not attempt any tuning.
- fnbr 4y agoparameter tuning is pretty necessary, according to anecdotes. People on twitter have got good results by changing the default parameters.
- int_19h 4y agoFor 13b and 30b, it really needs high temperature to produce good outputs.
- fdb 4y agoThe author just made an update that makes the generation much better, even with the 7B model: https://twitter.com/ggerganov/status/1634310199170179075 https://twitter.com/ggerganov/status/1634310199170179075 I tried it out myself (git pull && make) and the difference in results are day and night! It's amazing to play with, although you should prompt it differently than ChatGPT (more like the GPT-3 API).
- thewataccount 4y ago16GB of vram can run the 7B for sure, I'm not sure what the most cutting-edge memory optimization but the 15B is going to be pretty tight I'm not sure that'll fit with what I know of at least, I've got it working at a bit over 20gb of vram I think at 8bit. If you can't fit it all in vram you can still run it but it'll be slooooow, at least that's been my experience with the 30b.
- 0xbadc0de5 4y agoThe 4-bit GPTQ LLaMA models are the current top-performers. This site has done a lot of the heavy lifting: https://github.com/qwopqwop200/GPTQ-for-LLaMa https://github.com/qwopqwop200/GPTQ-for-LLaMa With 30b-4bit on a RTX 4090, I'm seeing numbers like: Output generated in 4.17 seconds (4.03 tokens/s, 21 tokens) Output generated in 4.38 seconds (4.25 tokens/s, 23 tokens) Output generated in 4.57 seconds (4.25 tokens/s, 24 tokens) Output generated in 3.86 seconds (3.40 tokens/s, 17 tokens) The lower size (7b, 13b) are even faster with lower memory use. A 16GB 3080 should be able to run the 13b at 4-bit just fine with reasonable (>1 token/s) latency.
- loufe 4y agoAt 4 bits the 13B LLaMa model can run on a 10GB card!