3 ms·
GPT-OSS-120B runs like hell on my DGX Spark
by vkaufmann 8mo ago
GPT-OSS-120B runs like hell on my DGX Spark
- embedding-shape 8mo agoThe MXFP4 variant I suppose? My setup (RTX Pro 6000) does around ~140 tok/s with llama.cpp, around 160 tok/s with vLLM.
- vkaufmann 8mo agoyep MXFP4 really fast :D