3 ms·
People in my server are running it on Strix Halo 128GB using RoCmFP4 and reporting 35tok/s, without much optimization, with proper MTP, better kernel, expecting
by OutlawHusbando 1mo ago
People in my server are running it on Strix Halo 128GB using RoCmFP4 and reporting 35tok/s, without much optimization, with proper MTP, better kernel, expecting about 50-60tok/s.
- manmal 1mo agoHow do they like it, compared with 3.8 and DS4?
- walrus01 1mo agoThe largest unsloth published gguf also fits and runs just fine on a cpu-only machine with 128GB RAM, using llama-server PR 27742 https://github.com/ggml-org/llama.cpp/pull/27742 https://github.com/ggml-org/llama.cpp/pull/27742