3 ms·
I am surprised at the lack of open-weights models in the >35B, but <200B range. I keep thinking about devices like the NVIDIA Spark and AMD Ryzen Halo, which ha
by Schlagbohrer 9d ago
I am surprised at the lack of open-weights models in the >35B, but <200B range. I keep thinking about devices like the NVIDIA Spark and AMD Ryzen Halo, which have their 128GB of combined memory, but there are so few models made for that range. Nearly all the open weights distillations are for larger customer bases with <24GB VRAM.
- MaKey 9d agoThe market is too small.
- ElectricalUnion 9d agoThe only (still in prototype stage!) "competitor" for those GB10/Ryzen Al Max+ 395 (in my region, borderline unobtainable) systems seems to be the Xiaomi AI Cube.
- Schlagbohrer 8d agoYes exactly, I am hoping that when the AMD Ryzen Halo gets wider release and more consumers have these 128GB devices, the market will justify a wider range of quantized model sizes. And then when I win the lottery and can buy one I'll have lots of nice options!
- bitexploder 9d agoQwen Flash Next 3.8 … even at 3 bit quant it is very solid.
- Catloafdev 9d agoQwen3.8 Flash Next just released which hits that range. Also, Deepseek V4 Flash can be run relatively well in hybrid 2-bit quantization on 128gb devices, with way better results than you'd expect for a typical 2-bit quant. Those are currently the 'smartest' options for that memory level.