3 ms·
5090 is plenty for the Q4_K_M quantized version of 3.6 27B with reduced context size. I run it on a 3090(24GB) and 64k context using GGUF format and llama-cpp.
by gessha 2mo ago
5090 is plenty for the Q4_K_M quantized version of 3.6 27B with reduced context size.
I run it on a 3090(24GB) and 64k context using GGUF format and llama-cpp. Double 3090 gives you 128k, quad 3090 gets you to full context - 256k.
- FeepingCreature 2mo agoOr you can run quantized context, there's some degradation but it fits in a lot less.