3 ms·
I administer a simple AI server in the office, which just uses a single RTX 5090 but is able to serve ~80 people throughout the day. I'm impressed by Qwen3.6-27
by kgeist 5mo ago
I administer a simple AI server in the office, which just uses a single RTX 5090 but is able to serve ~80 people throughout the day. I'm impressed by Qwen3.6-27b's capabilities in agentic coding/tasks so far. Devs say it's not much different from Sonnet 4.6 on many tasks (sometimes it even outperformed it), 40-60 tok/sec, up to 260k context. The server cost about $10k with all the bells and whistles.
I spent a lot of time researching/adding/benchmarking many custom modifications to the software stack and its settings to make the server optimally handle the load with just 1 RTX 5090 without losing quality, but it's still not enough, and the wait times in the queue are getting longer. We're at the limits of the hardware, and I'm out of tricks.
The experiment was kind of a success, and the CTO agrees we should scale it. With our own infra, we could run agents 24/7 on everything. Currently, a lot of use cases for the cloud providers are completely blocked by PII/trade secret concerns (our infosec department doesn't buy the "zero retention" promise), plus you don't have to think about billing/budgets/etc. anymore.
Now I can't decide how to scale it. On one hand, I'd like to run larger models. And we have the budget to buy, say, 8xH200. But in many benchmarks, the larger models that do fit in 8xH200 comfortably and can serve many parallel requests with acceptable speed/quality don't seem to outperform Qwen3.6 that much in agentic coding/tasks to justify the price.
So another option is just to buy a bunch of RTX 6000s and scale horizontally instead: run a copy of a midrange LLM like Qwen3.6 on each GPU. It's cheaper and easier to scale/replace, but then we'll run into problems running larger models in the future if we have to, because of no NVLink support (say, if Alibaba & Co. stop releasing ~30b models and/or ~30b models start falling behind 400b+ models considerably)
Does anyone here have experience running large models in a multi-GPU setup with several RTX 6000s in a high-concurrency regime and with large context lengths? (something like Deepseek 4 Flash, Minimax 2.7 etc.)
- zozbot234 5mo agoWouldn't that be a fairly ideal setup for layer parallelism? That doesn't need the high-performance communication of tensor parallelism, and the high-concurrency regime would make it easy to keep the pipeline full with microbatches. You'd also be able to scale out your KV cache storage since that naturally splits layer-wise.
- CobaltFire 5mo agoI have a 5090 machine sitting idle that I'm considering turning into a machine for my own small team (3 devs). Are you willing to share any lessons learned, etc. that I could make use of? We are evaluating paying for a SOTA sub or trying this, and the talk about Qwen3.6-27B makes me want to try deploying this machine.
- gpt5 5mo agoSell the machine for $4K, use it to pay for Codex Pro for everyone for a year. Everyone will be significantly more productive and happy. It's not even a real comparison if they are actually using them for coding. If you are deploying always running agents (e.g. monitoring logs and services) then sure - a QWEN local server is a good choice. But for coding the cost in productivity of using a lower performing model is way too high.
- 59nadir 5mo agoAnyone who frivolously suggests throwing away possible independence in favor of dependence on a Silicon Valley company is either incredibly naïve or acting in bad faith.
- gpt5 5mo agoI'll choose not to respond to your personal attack. But in term of actually running a dev team - you are free to use QWEN or another quantized local model that can run on an RTX 5090 for coding if it makes you feel more independence. However you would struggle and spend many many more hours achieving the same thing, with a lot more debugging time, long delays before it's done, and many more prompts. It's just not the right approach. I use QWEN and other local models all the time, but for more clearly defined monitoring and classification tasks.
- nine_k 5mo agoNot necessarily so. I can see how a bid to predict how thing will be in 1 year in AI-based coding is likely a losing one. So the idea is to extract the maximum value now, and turn it into profits that would buy you whatever is adequate for the next steps. For comparison, the AI-based coding landscape a year ago, in May 2025, wasn't even close to what we have now, and half the key tools did not exist. OTOH, as we see, the larger models demonstrate diminishing returns, smaller models demonstrate improvements, and hardware does not show any signs of becoming cheaper, so holding on existing decent GPUs may, too, be a winning strategy in longer term.
- CamperBob2 5mo agoDoes anyone here have experience running large models in a multi-GPU setup with several RTX 6000s in a high-concurrency regime and with large context lengths? (something like Deepseek 4 Flash, Minimax 2.7 etc.) For what it's worth, I've been seeing ~100 tps with 4-bit MiniMax 2.7 on two RTX 6000 boards, just running under llama-server without any optimization effort at all. I have no serious long-context experience with that setup, but at 30K context it's still above 90 tps. If you are happy with Qwen 3.6 27B, I would personally switch the 5090 out for 2x RTX 6000s and keep running 27B. That will give you ~2x your current throughput with a lot more headroom for multiple users. More important, it would buy time to see how things develop over the next few months before you spend a whole lot of money.
- undefuser 5mo agoWith that amount of memory can you run 4-bit DeepSeek 4 Flash? It is way more efficient in the KV cache department so may be worth a try
- CamperBob2 5mo agoI haven't looked into DS4 yet but based on antirez's results on 128 GB Macbooks, it shouldn't be a problem to run it on a pair of RTX6000 Pros. Also see https://www.reddit.com/r/LocalLLaMA/comments/1sv649s/to_run_deepseek_v4_flash_how_much_max_vram_we/ https://www.reddit.com/r/LocalLLaMA/comments/1sv649s/to_run_... .
- reissbaker 5mo agoQwen 3.6 27B is fine but it's not in the same ballpark as GLM-5.1 or Kimi K2.6. If you truly want to scale up, you should get the 8xH200 with NVLink.
- anon373839 5mo ago> our infosec department doesn't buy the "zero retention" promise They are wise to be skeptical! It is neither a promise nor zero data retention. Look at Anthropic's Zero Data Retention policy -- and remember, this is the policy that applies to the exclusively eligible enterprise partners who can even qualify for a ZDR agreement with Anthropic: > When ZDR is enabled, prompts and model responses generated during Claude Code sessions are processed in real time and not stored by Anthropic after the response is returned, *except where needed to comply with law or combat misuse*. > Even with ZDR enabled, Anthropic may retain data where required by law or to address Usage Policy violations. If a session is flagged for a policy violation, *Anthropic may retain the associated inputs and outputs for up to 2 years*.... This means that Anthropic is actively inspecting all of your data with machine learning classifiers. When the usage is flagged for whatever reason as violating any aspect of Anthropic's Usage Policy, then they get to keep your data for 2 years, with no apparent limitation on what they can then use it for. Crucially, you have ZERO guarantees about the sensitivity or specificity of these classifiers. For all anyone knows, Anthropic is silently flagging 75% of queries and retaining the data. https://code.claude.com/docs/en/zero-data-retention https://code.claude.com/docs/en/zero-data-retention
- random3 5mo agoI think it’s a cost/opportunity tradeoff at best with any agreement, regardless. The rest of the contract may make it difficult to impossible to do anything about it, starting with basic arbitration clauses and ending in a ton of other provisions that can make any legal action futile. I doubt there’s much room to negotiate too. Given that all labs need to diversify to become profitable, they’ll end up competing with their customers and theres nothing that exposes a business more than having AI offload every job function for every account, every mail etc. Assuming this won’t be an issue is naive at best.
- zenapollo 5mo agoI wonder how aws handles this in bedrock. Do they use Anthropics classifiers? Or their own? Or none? Would their data policing be different in bedrock than their other services?
- ramshanker 5mo agoThank you for the insight. This makes me feel confident, the L40S we are about to acquire with 48GB VRAM for engineering application should be useful for agentic coding as well.
- biddit 5mo ago> Does anyone here have experience running large models in a multi-GPU setup with several RTX 6000s in a high-concurrency regime and with large context lengths? (something like Deepseek 4 Flash, Minimax 2.7 etc.) Join the RTX6kPRO tribe! - https://discord.gg/pYCvaQTf https://discord.gg/pYCvaQTf - https://github.com/local-inference-lab/rtx6kpro https://github.com/local-inference-lab/rtx6kpro
- r0b05 5mo agoHow can a single 5090 serve 80 people? Something doesn't add up here.
- hacker_homie 5mo agoThey are using it as an assistant, bot running multiple fully automated agents loops?
- kgeist 5mo agoThey don't use the server all at once. In the UI, users typically ask a question, get a response, and continue with their work. In the case of autonomous agentic loops, an agent simply waits its turn until the server is ready to accept the request. Agents don't hammer the server 24/7 every second either, because they either need to be triggered or are busy doing other work, such as compiling or running tests.
- r0b05 5mo agoIt would be more interesting to know how many simultaneous users this setup can serve. Otherwise I can just say it serves 500 users but not all of them use it at the same time which doesn't communicate the right level of detail.
- p1esk 5mo agoDepends on TTFT and tokens per second you want.
- mixermachine 5mo agoWith parallelism of 16 you can still get around 25 to 30 tokens per user when all 16 channels are running. Not everyone will use the model at the same time but it certainly will be tight, especially for agentic coding. For pure chat applications this should be quite fine.
- zozbot234 5mo agoThe problem with wide parallelism with most models is that it blows up your KV cache. There's open models with KV caches lean enough to parallelize inference or even to offload the KV cache itself to disk without immediately running into wearout concerns, but they're quite exceptional.
- nicman23 5mo ago> 260k context with a single 5090?
- throawayonthe 5mo ago> don't seem to outperform Qwen3.6 that much in agentic coding/tasks idk i imagine you'll hit less edges with a larger model just because.. more data if you think of them as a kind of NN compression, it's ~obvious that the larger model can have more stuff encoded in it and hopefully accessible i don't use LLMs much right now but using midrange models seems like an unnecessary compromise in most cases, especially since the big open models sound to be rivaling opus and not just sonnet :p
- iamtheworstdev 5mo agoI thought NVLINK didn't matter anymore because of the latest PCI-E speeds. Am I wrong there?
- brianwawok 5mo agoAre we talking 1 GPU or 8?