4 ms·
At 4-bit quantization it should already fit quite nicely.
by GaggiX 6mo ago
At 4-bit quantization it should already fit quite nicely.
- Aurornis 6mo agoUnfortunately not with a reasonable context length.
- kkzz99 6mo agoIt really depends on what you think a reasonable context length is, but I can get 50k-60k on a 4090.
- GaggiX 6mo agoThe model uses Gated DeltaNet and Gated Attention so the memory usage of the KV cache is very low, even at BF16 precision.
- regularfry 6mo agoI've got 139k context with the UD-Q4_K_XL on a 4090, q8_0 ctk/v. Could probably squeeze a little more but that's enough for me for the moment.
- corysama 6mo agoHey, buddy! Can I bum a command line arg list off ya?