3 ms·
I run 8b on my Macbook air m2 with 16gb's of vram at decent speed of 12 t/s (mlx can probably get 15 t/s). I use the Q5_K_M GGUF, which is >99% the same as the
by MyFirstSass 2y ago
I run 8b on my Macbook air m2 with 16gb's of vram at decent speed of 12 t/s (mlx can probably get 15 t/s).
I use the Q5_K_M GGUF, which is >99% the same as the original.
I've seen tests that there is close to no divergence with these quantisations, but it rises steeply going lower:
https://www.reddit.com/r/LocalLLaMA/comments/1816h1x/how_much_does_quantization_actually_impact_models/ https://www.reddit.com/r/LocalLLaMA/comments/1816h1x/how_muc...