3 ms·
I'm running qwen3.6, gpt-oss and embeddinggemma with Ollama on Ryzen 7 8700G + 96 GB DDR5 with 12-14 tokens/sec. Consider this as a floor. On Strix Halo it will
by denn-gubsky 2mo ago
I'm running qwen3.6, gpt-oss and embeddinggemma with Ollama on Ryzen 7 8700G + 96 GB DDR5 with 12-14 tokens/sec. Consider this as a floor. On Strix Halo it will run 4-6 times faster depending on memory bandwidth. For my local task 14 tok/s is quite enough for the price I paid.
- sloaken 2mo agoSo no GPU? Just a strong processor? I am considering building a 'in my house' system. Timing of this post is very fortuitous.
- denn-gubsky 2mo agoRyzen 7 8700G is an APU, and it has built-in Radeon 780M GPU (12 CUs, ~12.6 TFLOPS, plus an NPU). This is not very strong, but this is floor which runs local models for me and saves a lot of paid tokens. Potentially this setup may be upgraded to AI MAX when AM5 APUs will become available in box packages (now they are OEM-availabe only). The most expensive part of my system is DDR5, and it may be reused in case of upgrade. Actually, this is my self-hosted NAS (TrueNAS Scale) which also runs Ollama and serves as my local inference machine.