3 ms·
Half a terabyte could run 8 bit quantized versions of some of those full size llama and deepseek models. Looking forward to seeing some benchmarks on that.
by api 2y ago
Half a terabyte could run 8 bit quantized versions of some of those full size llama and deepseek models. Looking forward to seeing some benchmarks on that.
- zamadatix 2y agoDeepseek would need Q5ish level quantization to fit.