3 ms·
This situation has improved quite a bit recently, Qwen Flash Next will run on a $4000 PC and can reliably implement small features on its own (feels comparable
by hedgehog 12d ago
This situation has improved quite a bit recently, Qwen Flash Next will run on a $4000 PC and can reliably implement small features on its own (feels comparable to Opus 4.5). It's a bit slow but pretty effective.
- bitexploder 12d agoLess. Probably 2-3K if you build right. Qwen 3.8 27B on constrained tasks is Opus 4.6-ish to me, it just doesn’t know enough, but when task is laid out just gets it done. Comes down to how much of the ambiguity we expect out of the model.
- hedgehog 12d agoWhat would you build? Flash Next on 128GB works ok, 3.8 27B might technicially work on less but it's just too slow at least on current APU style chips. I'm curious about the non-NVIDIA 32GB discrete GPU options but I haven't tried yet.
- bitexploder 11d agoFlash next only needs 64GB for core model inference. If you really wanted it. Need 64GB vram, 900+ GB/s speed, and a lot of system ram (128GB). Seems feasible. Hmm. Maybe older GPUs work.
- hedgehog 11d ago128GB on a Ryzen 395 is enough but not by a lot. Around 90GB for weights plus 10GB to 20GB for KV cache and checkpoints. I don't know about 64GB of VRAM for less than around $2500 by itself.
- bitexploder 11d agoTesla V100? Would have to be a 3-ish bit quant in 64GB ram