3 ms·
We need a rule that if your LLM benchmark is running under 20t/s it's simply unusable in any real workflow. 6t/s is unbearable, if you used it with OpenCode yo
by bearjaws 7mo ago
We need a rule that if your LLM benchmark is running under 20t/s it's simply unusable in any real workflow.
6t/s is unbearable, if you used it with OpenCode you would be waiting 20+ minutes per turn.
- zozbot234 7mo agoThis is not an ordinary LLM benchmark, it's streaming experts' weights from storage. It opens up running very large (near-SOTA, potentially SOTA) MoE models on very limited hardware, since you no longer need enough RAM for the entirety of the model's parameters. The comparison to 20 t/s local AI models is simply not fair.
- bearjaws 7mo agoI understand that. I am saying there is a clear cliff where the value of an LLM reaches 0. At 1t/s you are never going to get anything done. At 6t/s, it's significantly degrades the experience, one mistake setting you back 20-30 minutes. At ~20t/s it's much more usable.