3 ms·
It's running a very small, non-reasoning model at the moment. But more generally, almost all LLMs argue on the hardware/model they are/are on.
by shaewest 2mo ago
It's running a very small, non-reasoning model at the moment. But more generally, almost all LLMs argue on the hardware/model they are/are on.
- dumberquestions 2mo agoWhich model? Or how many active parameters?
- _whiteCaps_ 2mo agoLlama 3.1 8B model
- dumberquestions 2mo agoSo this demo is around 90 times faster than typical speeds for the same model at openrouter, and around 30 times faster than the absolute fastest option available (Groq).
- anthonypasq 2mo agoim assuming energy expenditure is substantially lower as well
- mdp2021 2mo agohttps://taalas.com/h-content/uploads/2026/02/graph.png https://taalas.com/h-content/uploads/2026/02/graph.png
- Gander5739 2mo agohttps://xkcd.com/1162/ https://xkcd.com/1162/
- deleted 2mo ago[deleted]
- metadat 2mo agoWhat would tokens/sec performance look like for a reasoning model? An order of magnitude slower?
- penagwin 2mo agoReasoning models are the same speed. They’re just post trained with RL to do CoT inside tags like <thinking></thinking> before a tag like <response></response> There’s no difference in the inference implementation, parameter count, or speed.
- paytonjjones 2mo agoThere's a difference in the latency distribution between when you submit a query and you see the response, which is what the comment is (clumsily) asking about. But yeah, there are a lot of factors, so it's hard to answer, and tokens/s isn't the right question.