3 ms·
> ...one of the truly baffling things from either their team or AMDs team, is how much the workstation grade AMD r9700 has been ignored. It makes a huge amount
by roenxi 1mo ago
> ...one of the truly baffling things from either their team or AMDs team, is how much the workstation grade AMD r9700 has been ignored.
It makes a huge amount of sense after considering AMD's approach to graphics cards from around 2010 to 2025. They just didn't see graphics cards as viable compute platform and many who made the mistake of believing that good specs would translate into in-practice performance got badly burned. I'd have been involved in the AI boom but for an expensive AMD graphics card, I'm not going to forget that for a while.
George Hotz was interesting as a public example, but I think his story probably repeated a few times outside the public eye. People tried to make AMD work and ended up the worse for it.
People who had an interest in using AMD cards to get things done are probably by and large waiting for a new generation of hopefuls to prove this time is different. The mutterings out of AMD are promising, but that isn't persuasive enough given the scale of the failures.
- hgoel 1mo agoGCN was such a promising compute architecture, AMD even pioneered stuff like async compute and compute shader heavy rendering pipelines, only to never seriously go beyond that on consumer gear. I agree with your assessment that the story of supporting the competition, only to get burned, has repeated many times with AMD outside the public eye. It's why I don't put much stock in claims that things work great as long as specific flags are used.
- lrvick 1mo agoI have 4x r9700s as my coding daily drivers running qwen 3.8 27b at 80tps each. Zero complaints. Especially at $1200 each.
- SomeHacker44 1mo agoPlease share what operating system and model runtime you use? I have two and don't get close to that with AMD's own Lemonade. Thanks!
- lrvick 1mo agoLemonade is one of the worst performing options. Run any modern Linux distro and ask your current LLM to setup llama.cpp with dflash2 for you as an unprivileged container running from a systemd user unit. Obviously only on a system you do not trust at all.
- pyrolistical 1mo agoNo way you can get that much each without extremely quants For a single r9700 you have 637 GB/s and for qwen 3.8 27b q4_k_xl the maximum tg/s is 33 before mtp Now if you meant 4xr9700 tensor parallelism with mtp, 80 tg/s starts to make sense
- naasking 1mo agoQ4 isn't an extreme quant, and I average 75 toks/s on code, 45 tok/s on prose with MTP.
- latentsea 1mo agoYou can get those numbers with https://codeberg.org/ggz14/radiance-vllm-mxfp4 https://codeberg.org/ggz14/radiance-vllm-mxfp4. I also can get it on a single R9700 but the 75 ~ 80 t/s is only peak acceptance of very predictable tokens like coding or json, and averages lower for prose. It's still much faster than regular llama.cpp.
- julianlam 1mo agoCan confirm. I've a single R9700 and have maxed out at 45 tokens/second on llama.cpp with Q4 Qwen 3.6 27B (with MTP)