2 ms·
$2600 will buy you two AMD 9700 gpus with 32Gb ram per card running about 285 Watts per card. Less than a 5090 in both cost and power. A VLLM build patched for
by androiddrew 4mo ago
$2600 will buy you two AMD 9700 gpus with 32Gb ram per card running about 285 Watts per card. Less than a 5090 in both cost and power. A VLLM build patched for AITER and you can run Qwen3.6 27B FP8 at roughly 45-50TPS during real coding sessions with Opencode or PI with a full context window. I really hope more 30B dense models continue to be released, but Qwen3.6 should get you a lot of agentic mileage.
ROCm stack is not for people though who aren’t willing to dig in and patch things themselves.