3 ms·
I am running it on a single RTX 3090 (24GB VRAM). Some folks on Reddit are having the same experience: https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/mus
by delicious_apple 2mo ago
I am running it on a single RTX 3090 (24GB VRAM).
Some folks on Reddit are having the same experience: https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_glimmer_actually_fits_on_a_single_rtx_3090/ https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_gl...
It uses an order of magnitude less VRAM at longer contexts which is a huge advantage over Qwen 3.6 27B
- r0b05 2mo agoSeems like that's the tradeoff with this model. Close to 27b intelligence while using less vram.