3 ms·
It will work fine but it’s not necessarily insane performance. I can run a q4 of gpt-oss-120b on my Epyc Milan box that has similar specs and get something like
by oceanplexian 7mo ago
It will work fine but it’s not necessarily insane performance. I can run a q4 of gpt-oss-120b on my Epyc Milan box that has similar specs and get something like 30-50 Tok/sec by splitting it across RAM and GPU.
The thing that’s less useful is the 64G VRAM/128G System RAM config, even the large MoE models only need 20B for the router, the rest of the VRAM is essentially wasted (Mixing experts between VRAM and/System RAM has basically no performance benefit).
- syntaxing 7mo agoSplit RAM and GPU impacts it more than you think. I would be surprised if the red box doesn’t outperform you by 2-3X for both PP and TG
- androiddrew 7mo agoCould you share what you are using for inference and how you are running it? I have a 64G VRAM/128G system RAM setup.
- sosodev 7mo agoMost people are using something in the llama family for inference. Llama server is my go to. Unsloth guides describe how to configure inference for your model of choice.
- datadrivenangel 7mo agoYeah I've got the q4 gpt-oss-120b running at ~40-60 tokens per second on an M5 Pro.