4 ms·
I like this calculator but it’s really wrong at least for dgx spark. I have one and I get 4x the tokens/s .
by shadowpho 19d ago
I like this calculator but it’s really wrong at least for dgx spark. I have one and I get 4x the tokens/s .
- rlindsey123 19d agoAw very interesting! This is great feedback - what model are you running? I'm keen to do more crowdsourced data as time goes on.
- shadowpho 19d agoThe big three :) Qwen3.8-flash-next Deepseek4-0731-flash Glm5.3 The latest unsloth llama.cpp has a lot of nice features that runs them faster than before. I’ll have to double check which one runs how fast, but it’s generally 20-40 t/s. (And infil is fast but not sure how that’s counted)