3 ms·
I think if they made this for Qwen3.8-Next it could fit in a single 5090?
by 2001zhaozhao 15d ago
I think if they made this for Qwen3.8-Next it could fit in a single 5090?
- kennywinker 15d ago180b * 1.76 bits per weight = 39.6 gigabytes. Best you could realistically run in 32gb is like 28gb, or a 127B param model
- jokethrowaway 14d agoQwen3.8-Next, thanks to its new architecture, is quite fast even if part of it is streaming from disk
- kennywinker 14d agoTotally. Any MoE model can have experts swapped in and out from disk or system ram. I only framed it this way because the question was about the model fitting in vram.