4 ms·
It would be great if they made their inference capacity for this model available via OpenRouter; the fastest provider on OpenRouter right now is at ~80tps https
by jasongill 1mo ago
It would be great if they made their inference capacity for this model available via OpenRouter; the fastest provider on OpenRouter right now is at ~80tps https://openrouter.ai/qwen/qwen3.8-27b#providers https://openrouter.ai/qwen/qwen3.8-27b#providers
They do appear to host other models on OpenRouter so maybe Qwen3.8 will be there soon: https://openrouter.ai/provider/cerebras https://openrouter.ai/provider/cerebras
- zackangelo 1mo agoWe're serving it around 150-200tok/s (uses our new speculative decoding implementation on a DFlash2 draft model). https://mixlayer.com https://mixlayer.com, LAUNCH-Q38-27B gets you $5 in credits if you want to kick the tires.
- deleted 1mo ago[deleted]
- bookernath 1mo agoThis feels great
- danielklnstein 1mo agoI tried in your playground and got 14.2 tok/s?
- zackangelo 1mo agoapologies we just got a sudden burst of new users and traffic, it's scaling up now.
- zackangelo 1mo agojust added 8 more H200s to the cluster, if you (or anyone else) runs into issues please feel free to drop me a message: zack at mixlayer.com
- danielklnstein 1mo agoWorks much better now! Got 103.9 tok/s, not quite 200 - but still amazing! Thanks for sharing
- zackangelo 1mo agoSomething a lot of model providers don't talk about: any time an engine uses speculative decoding the throughput will depend on how much your output token distribution matches what the draft model was trained on. The DFlash2 draft model we're using was trained on a lot of code, so if you use it in a coding agent you'll probably notice it run a lot faster (we've seen it break 300 tok/s).
- danielklnstein 1mo agoFYI, I might be missing something but I think your billing system might not be working well - I'm not seeing any indication in the UI that my usage is being deducted from the $5 of free credits.
- chrisboulton 1mo agoHey Daniel! It's a bit hidden, but at the bottom of the billing page there's a "Credits" section which should show usage of any active credits and the balance remaining. The usage/billing metrics are batched/handled async so it might take a minute or so for usage to be reflected. Let us know if it feels off.
- RussianCow 1mo agoI don't see any kind of input cache discount listed on your pricing page. Do you offer that, or is all input priced the same?
- scratchyone 1mo agoany way to see the tok/s for all the models listed on your homepage? curious which has the best speed/quality tradeoff for me
- egorfine 1mo agoTrying now. History question: got 96 t/s. Code review task: 130 t/s. Nice!