3 ms·
Because concurrency with a single "accessible" card quickly diminishes. I have dual 3090s, and on Qwen 3.6 35B A3B at 80k a single card with max concurrent set
by jermaustin1 2mo ago
Because concurrency with a single "accessible" card quickly diminishes. I have dual 3090s, and on Qwen 3.6 35B A3B at 80k a single card with max concurrent set to 4 will get 70-80tps single request, 50-60 TOTAL tps with 2, 45-50 with 3, and around 40tps with all 4 going.
I know this doesn't exactly match your request, but I'm happy with my 3.6's performance. I've seen people claim they can get 120+tps on a 3090, but I'm not impatient when it comes to streaming text faster than I can read it.
- Tostino 2mo agoYou have something misconfigured then. Concurrency has never lowered my overall TPS. Also have dual 3090s. Generally use vllm though.
- jermaustin1 2mo agoI've had some rough time getting LM-Studio properly configured for multi-card. It exists, but I feel like it is kind of buggy. I will disable a card and it will still load the model into it. Sometimes it will split the model even though there is loads of room available. I might need to finally make the switch away from it, but it is so convenient, especially as a chat interface for system prompt experimentation.
- pich 2mo agovLLM is probably the key difference there… its scheduler is built around batching/concurrency, while this setup is heavily optimized llama.cpp for single-stream latency
- dannyw 2mo agoYour configuration is broken or wrong. What are you using? Hopefully not llama.cpp? I’ve sweeped concurrency across many models and many different kinds of hardware, and the only times I saw similar results to you were when I didn’t configure it correctly.
- mhitza 2mo agoWhat do you use instead of llama.cpp? With vllm for example most models don't seem to be supported out of the box.
- supermatt 2mo agoI haven’t tried any larger models, but a 12B model on my Ampere A5000 gets around 4-5x the aggregate throughput with concurrency. I have the maximum context configured to 32k, but my actual requests are usually around 2-4k tokens. No idea how that compares to running a larger model and context though.
- petu 2mo agoI guess it's due to testing on MoE. Different completions activate different experts, thus very little cache reuse and completions "steal" memory bandwidth from each other. As I understand (useful) concurrency for MoE requires very large batches, where about every expert gets activated per pass. With dense Qwen 27B on 3090/llama.cpp I get: - no MTP: 1x42, 2x33, 3x24, 4x19 t/s - MTP: 1x50, 2x30, 3x33, 4x30 t/s
- jermaustin1 2mo agoThat is interesting. I'll have to test that theory out today.