3 ms·
Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU
- deleted 2mo ago[deleted]
- supermatt 2mo agoCan you please try and see how many tokens you get with some form of concurrency. Pretty much ALL the benchmarks I've seen on the more accessible cards are just single request.
- jermaustin1 2mo agoBecause concurrency with a single "accessible" card quickly diminishes. I have dual 3090s, and on Qwen 3.6 35B A3B at 80k a single card with max concurrent set to 4 will get 70-80tps single request, 50-60 TOTAL tps with 2, 45-50 with 3, and around 40tps with all 4 going. I know this doesn't exactly match your request, but I'm happy with my 3.6's performance. I've seen people claim they can get 120+tps on a 3090, but I'm not impatient when it comes to streaming text faster than I can read it.
- Tostino 2mo agoYou have something misconfigured then. Concurrency has never lowered my overall TPS. Also have dual 3090s. Generally use vllm though.
- jermaustin1 2mo agoI've had some rough time getting LM-Studio properly configured for multi-card. It exists, but I feel like it is kind of buggy. I will disable a card and it will still load the model into it. Sometimes it will split the model even though there is loads of room available. I might need to finally make the switch away from it, but it is so convenient, especially as a chat interface for system prompt experimentation.
- pich 2mo agovLLM is probably the key difference there… its scheduler is built around batching/concurrency, while this setup is heavily optimized llama.cpp for single-stream latency
- dannyw 2mo agoYour configuration is broken or wrong. What are you using? Hopefully not llama.cpp? I’ve sweeped concurrency across many models and many different kinds of hardware, and the only times I saw similar results to you were when I didn’t configure it correctly.
- mhitza 2mo agoWhat do you use instead of llama.cpp? With vllm for example most models don't seem to be supported out of the box.
- supermatt 2mo agoI haven’t tried any larger models, but a 12B model on my Ampere A5000 gets around 4-5x the aggregate throughput with concurrency. I have the maximum context configured to 32k, but my actual requests are usually around 2-4k tokens. No idea how that compares to running a larger model and context though.
- petu 2mo agoI guess it's due to testing on MoE. Different completions activate different experts, thus very little cache reuse and completions "steal" memory bandwidth from each other. As I understand (useful) concurrency for MoE requires very large batches, where about every expert gets activated per pass. With dense Qwen 27B on 3090/llama.cpp I get: - no MTP: 1x42, 2x33, 3x24, 4x19 t/s - MTP: 1x50, 2x30, 3x33, 4x30 t/s
- jermaustin1 2mo agoThat is interesting. I'll have to test that theory out today.
- pich 2mo ago[dead]
- nubg 2mo agoquantization level?
- MaxikCZ 2mo agoIts egregious the quant level isnt disclosed along the "Qwen" string. Everyone knows theres huge difference in speed/quality along the quant axis, I now attribute the ommision of such to deliberate choice to not curb the hype of the tittle.
- metadat 2mo agoThe article mentions Q4, Q5, Q8, and NVFP4. It's total AI slop though, tough read. In my testing I got 150 tokens/sec with a single 5090 RTX.
- Foobar8568 2mo agoWhich model/quant/command line did you use? I can barely get 100 token/secs and for sure clearly not a full context. With vllm, I am limited to 130k tokens with vllm + nvfp4.
- iv42 2mo agoIf you've got RTX 5090, maybe try ninfer (https://github.com/Neroued/ninfer https://github.com/Neroued/ninfer). Folks over on /r/localllama have been reporting wild prefill/token gen speeds with ninfer (NVFP4; 256k ctx).
- Foobar8568 2mo agoLooks like slope benchmarks and results, and as usually, people are mixing MTP numbers with non MTP numbers. Or just 100 token input benchmarks. Or just failed ones as actual measures. https://github.com/Neroued/ninfer/blob/master/docs/performance.md https://github.com/Neroued/ninfer/blob/master/docs/performan... Category MTP3 stochastic sampler DFlash stochastic sampler DFlash greedy Code 1/15 natural stops; 0/15 prompt-complete 2/15 natural stops; 0/15 prompt-complete 0/15 natural stops Story 9/15 natural stops; the nine Chinese outputs pass requested division and minimum length 8/15 natural stops; the eight Chinese outputs pass requested division and minimum length 10/15 natural stops; five Chinese dialogue outputs are under length Translation 15/15 natural stops; 15/15 pass structural checks 15/15 natural stops; 15/15 pass structural checks 15/15 natural stops; 15/15 pass structural checks Structured 0/15 satisfy the requested complete record/script contract 0/15 satisfy the requested complete record/script contract 0/15 satisfy the requested complete record/script contract And on my "own" "quick" benchmark, it's slower than vllm.
- nodja 2mo agoThe whole site looks like and reads like AI slop. The outcomes also don't make any sense and don't feel rigorously tested (no, having claude test for you doesn't count as rigorous).
- genxy 2mo agoThe person is having a AI induced manic episode, we have all been there.
- IncreasePosts 2mo agoPlease stop making this comment. The war is lost. Instead, you should be commenting that it looks like a human wrote this when you come across the rare brain-produced writing
- simonw 2mo ago"Combining them into one heroic speedup would make a better headline and a worse benchmark." "The machine immediately taught me that capacity estimates are just admission tickets." "Useful in production, poison in a kernel comparison." Please don't publish writing like this, it's exhausting to read. You can edit that stuff out. The best part of this is the "Five xhigh artifacts from the finished model" section at the bottom, I suggest either moving that up or at least prominently promoting it at the top of the article.
- PeterStuer 2mo agoIt has gotten so much worse over the last month. The default writing style of the Claude 5 model series in Claude Code is some sort of jiberish jargon.
- dannyw 2mo agoI find the ‘explanatory’ output style of Claude to be a bit more tolerable, but yes. Claude seems to speak and write more in Claude-speak with every release.
- cedws 2mo agoApparently we've blown way past the Turing test and approaching AGI and yet LLM-generated text still sticks out like a sore thumb. Maybe LLMs aren't that good at writing after all.
- wgd 2mo agoLLMs are great at writing, it's The Assistant who is a terrible writer. Sadly that one persona is all you get these days.
- Groxx 2mo agoI think it's fair to say they're better at a paragraph or so than most humans. And have been for quite some time, which is probably why their use in writing has exploded. Long form though? Still pretty bad. Probably getting worse in practice, as people have them write larger and larger chunks of text without paying any more attention to the result.
- spottedmarley 2mo ago[flagged]
- Tepix 2mo agoAlways put the quantisation in the title!
- pich 2mo agoIts not quite that simple here. The iMatrix-guided hybrid uses different quantization levels per tensor/layer, so there isnt one honest Q4/Q5/NVFP4 label I can put in the title
- gitowiec 2mo agoWhy it's flagged?
- jonaddb 2mo ago[flagged]