6 ms·
This seems saying that the 1bit model is a bit better than GPTQ Q2. However, I find there are few situations where you would want to use GPTQ Q2 in the first pl
by WiSaGaN 2y ago
This seems saying that the 1bit model is a bit better than GPTQ Q2. However, I find there are few situations where you would want to use GPTQ Q2 in the first places. You would want to run the F16 version if you want quality, and if you want to have a sweet spot, you usually find something like Q5_K_M of the biggest model you can run.
- qeternity 2y agoNobody is running llama.cpp in production…
- luke-stanley 2y agoWhat do you mean? What makes you think that?
- qeternity 2y agoBecause for anything other than CPU inference it is inferior to TensorRT and vLLM.
- luke-stanley 2y agoAre you sure about that? what are the benchmarks it fails on when setup like-for-like with GPU drivers? Even still, it can do constrained grammar and is really easy to setup with wrappers like Ollama and the Python server too. I struggled to find support for that with other inference engines - though things change fast!
- szundi 2y agoProbably more often than not
- qeternity 2y agoHighly doubtful. If anyone is running it here in a production setting, please post and prove me wrong.