3 ms·
I have a 5950x with 64 gb ram and they are quantized to 4 bit yes :) The weights are stored on a samsung 980 pro so the load time is very fast too. I get about
by tyfon 4y ago
I have a 5950x with 64 gb ram and they are quantized to 4 bit yes :)
The weights are stored on a samsung 980 pro so the load time is very fast too.
I get about 2 tokens/second with this setup.
edit: forgot to confirm, it is llama.cpp
edit2: I am going to try the FP16 version after easter as I ordered 64 GB of additional ram. But I suspect the speed will be abyssal with the 5950x having to calculate through 120 gb of weights. Hopefully some smart person will come up with a way to allow the GPU to run off system memory via the amd infinity fabric or something.
- barbariangrunge 4y agoI thought it needed 64gb of vram. 64gb of ram is easy to obtain
- sbierwagen 4y ago5950x is a CPU model. Integer-quantized models are generally run with CPU inference. For the larger models the problem then becomes generation time per token.
- int_19h 4y agoQuantized models are used aplenty with GPUs as well - 4-bit quantization is the only way you can squeeze llama-30b into 24Gb of VRAM (i.e. RTX 3090 or 4090). In fact, I would say that, at this point, most people running LLaMA locally are likely using 4-bit quantization regardless of model size and hardware, just to get the most out of the latter.
- whimsicalism 4y agoMost people running llama locally are doing CPU inference, period.
- barbariangrunge 4y agoIf your desktop had 256gb of ram, could you train a far larger model? Some motherboards support that