3 ms·
I wonder if the new mbp can run it at q4.
by coconut08 2y ago
I wonder if the new mbp can run it at q4.
- throwdbaaway 2y agoUsing https://github.com/kvcache-ai/ktransformers/ https://github.com/kvcache-ai/ktransformers/, an intel/amd laptop with 128GB RAM and 16GB VRAM can run the IQ4_XS quant and decode about 4-7 token/s, depending on RAM speed and context size. Using llama.cpp, the decoding speed is about half of that. Mac with 128GB RAM should be able to run the Q3 quant, with faster decoding speed but slower prefilling speed.