6 ms·
RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
- ComputerGuru 4mo agoI would have liked to see a bit more on the theory side of things, explaining optimal weight and inference splits, actual issues with existing drivers, etc instead of what’s essentially just a recipe.
- verdverm 4mo agoI've been using https://spark-arena.com/leaderboard https://spark-arena.com/leaderboard to glean this kind of information for DGX Spark, a sort of recipe book. The Nvidia forum has people talking about the things you wish to know. I see some on Discord/Reddit/et al, but less cohesive I've switched from using the spark as a way to run one model as best it can to running several support models for the md kb I'm working on
- atq2119 4mo agoAgreed. To put this in perspective, batch 1 token decode is bandwidth limited in theory. Memory bandwidth of RTX 3090 is listed as 936GB/s. The post isn't fully clear on which model they used and how big it is, but even assuming it perfectly filled the 24GB of that GPU, 30tok/s means the achieved bandwidth is only 720GB/s. There's a bunch of room for improvement here even without MTP, and those improvements should largely stack with MTP.
- deng 4mo agoI can understand the joy of running things yourself, and can also see the privacy aspect. However, I pay ~3$ per 1M/tokens for that model on Openrouter, and it's not even quantized. A refurbished 3090 and a 5080 will set you back well over 2k, not to mention the electricity to run them...
- TSiege 4mo agoIt’s a personal hobby project why should we care this is how someone chooses to spend their free time and money? Lots of hobbies are expensive and pointless if you think of commercially available offerings. That’s why it’s a hobby and not a small business
- redfloatplane 4mo ago> I pay ~3$ per 1M/tokens for that model on Openrouter I think the thing is, there's an unspoken "for now" at the end of that sentence and people running this locally are hedging against that "for now". Some people prefer to feel that they own the means rather than rent the means, even if the one they own is worse than the one they can rent. Especially with today's Fable news and the harsh realisation that the "for now" is dependent on very many unpredictable factors, where the one you have locally costs you capital today and a relatively predictable run-rate (made more predictable with on-prem solar for example), but should otherwise work predictably forever. I'm not saying that you're wrong to do what you're doing, just that many people have their own lines in the sand where renting vs buying makes sense, and it doesn't only boil down to a rational (or irrational) financial decision.
- jubilanti 4mo agoYou're treating open weight inference providers the same as proprietary ones. They're fundamentally different business models. Proprietary companies have an incentive to subsidize actual inference and training costs in order to gain market share. The few dozen or so companies selling Qwen models by the token on openrouter are in a commodities market. If suddenly the CCP declared a total digital embargo on Alibaba's Qwen models or even if for some reason all of mainland China (and Singapore) was completely unreachable from the rest of the world, the dozen or so companies selling Qwen by the token elsewhere in the world could continue business as usual.
- redfloatplane 4mo agoI was thinking of user-side regulations as well, not only provider-side ones. I could imagine a world where a government rules that you may not use LLMs for anything, which would be much easier to get around if you have local means.
- bee_rider 4mo agoI don’t know anything about the open weight host business model. Do we know for certain that the folks selling inference by the token are really selling them in an upfront and profitable way? No subsidies from harvesting the info, to sell to the model trainers or anything like that?
- Der_Einzige 4mo agoOpenrouter doesn't give you access to the models internals, i.e. complete control of logprobs, sampler stack, any PeFTs. Openrouter fking sucks and I don't know why people here act like it's so great. Stop using it if you care about local AI and accept that the cost you'll pay for tokens is higher than you will when consumed via any cloud. That's the price for privacy, control, and better quality via inference time optimizations that otherwise aren't available.
- jubilanti 4mo ago> Openrouter doesn't give you access to the models internals, i.e. complete control of logprobs, sampler stack, any PeFTs. Openrouter gives you access to whatever the inference provider gives. They're just the middleman. Many providers give logprobs if you ask, it's in their API. And yeah, no Peft or Lora, but that's an entirely different product. And some of the inference providers do that directly. > Openrouter fking sucks and I don't know why people here act like it's so great. Stop using it if you care about local AI But the whole point of openrouter is that you can run models by the token and you don't have to care about local AI? Sounds like you're more upset that people aren't making the same calculation on privacy and local control vs cost and ease of use.
- NicoJuicy 4mo agoRtx 3090 24 gb set me back 390€ a year ago ( 2nd hand)
- rirze 4mo agoWas it still in good condition? That price makes me wonder if it was used for crypto mining, which can wear down the hardware.
- toyg 4mo agoYeah but they can also be used to play games and do other stuff.
- ThunderSizzle 4mo agoAn R9700 is $1350 and can get 100 TPS running Qwen3.6-35B-A3B Q5 with 130k context window (with room to spare) with a bit of fine tuning llamacpp-vulkan, but llamacpp's repository instability and lack of real versioning frustrates me. In terms of electricity, if you aren't using it, even with all the vram loaded, at most your wasting about 30 watts or so. Prompt processing a large uncached context is annoying, which is why I forced a lower context window, but I don't know if it's any worse in performance than the cloud models I've used. There's a niceness, to me, knowing I don't have to rent it anymore. If you rent it, the terms can change regularly.
- bertili 4mo agoQwen 27b is a compute heavy dense model.
- rsync 4mo ago"An R9700 is $1350 and can get 100 TPS running Qwen3.6-35B-A3B Q5 with 130k context window ..." How would that change (improve) if you had two R9700 in a similar configuration ?
- vardalab 4mo agobetter prompt processing like 1.5x+ and more kv but tg most likely lower like 0.8x or so but I am just going by memory for Qwen3.5 without mtp.
- medfield 4mo agoI use local models to explore, hosted models to refine. I somewhat envy those who can sustain local models (q8 120b+) running as a hobby.... for me, the practical path is a better SearXNG setup and knowing my routes forward.
- PeterStuer 4mo agoWhen they declare open models a 'security risk', his setup will be running, yours will not and even that 3090 will be way outside of your reach.
- deleted 4mo ago[deleted]
- amelius 4mo agoYou are paying with your privacy ...
- alexjplant 4mo agoI've spent the past week trying to scheme a way to get affordable local inference of something useful (Qwen3.6-36B-A3B) for ~$500 and have come to the conclusion that it simply isn't viable. A pair of power-restricted P100s in a workstation gets close but the workstations themselves are expensive and rare as hen's teeth (not to mention loud and large). I think early '27 will be when things open up as the hardware market unclenches and further strides are made in small capable models.
- mappu 4mo agoI'm running Qwen3.6-35B-A3B on a very ordinary desktop PC (32GB DDR5, 8GB Radeon 6600XT) and getting a useful 15-20 tok/sec out of it. The MoE architecture and auto offloading from system to VRAM is just fantastic. Unsloth Q4_K_XL. The Qwen3.6-27B is unbearably slow as it doesn't fit in VRAM, though, i think the MoE is very easy to run. It is also extremely nice that you can just `apt install llama.cpp libggml0-backend-vulkan` now too.
- ozim 4mo agoI wonder what parent poster means with „useful” and what he actually tried? Feels like he was just comparing some benchmarks. Yesterday I downloaded Gemma4-26B with Ollama on quite rusty desktop with 1070 8gb and 32gb of ram and Core i5-9400. I drop photo of my water meter and tell it to read the value and serial number. It was far from instant but it was also easily under 3 minutes and result was correct. Earlier like in February I was trying the same photo with Gemma3 on the same hardware and results were bad.
- alexjplant 4mo ago> I drop photo of my water meter and tell it to read the value and serial number. It was far from instant but it was also easily under 3 minutes and result was correct. "Useful" as in "has a use that isn't just for show". It takes me two seconds to read a photo of a water meter. Having an LLM read it for me in 3 minutes isn't useful. Similarly small models are capable of tool use (e.g. web searches) but their synthesis leaves much to be desired. As an example I'd ask some small models to find examples of products with specific characteristics and they'd come back with only one or two because they discounted other possibilities incorrectly by reasoning themselves out of it. > Feels like he was just comparing some benchmarks. On what do you base this assertion?
- alexhans 4mo agoI think it's important to be able to do both so you can stay in control of the price to value created relationship. In last year, some people were publishing aider /ollama/open router [1] and now thankfully people are publishing all around about pi/qwen/llama.cpp/openrouter. It's widespread. [1] https://alexhans.github.io/posts/aider-with-open-router.html https://alexhans.github.io/posts/aider-with-open-router.html
- pier25 4mo ago> not to mention the electricity to run them... And noise.
- flowbarai 4mo ago[flagged]
- sixothree 4mo agoYou also aren't limited to LLMS. Vision, whisper, etc. You can even have claude farm out tasks to your local servers.
- avyeed_desa 4mo agoI just bought a $25 chinese 2x Oculink card and two Minis Forum DEG1, had some spare PSUs lying around, and just installed two cards on each. It works. I saw that there is also a 4x Oculink card, but i don't know it that will work, too.
- atlgator 4mo agoWhich "good quality PCIe 4 riser" did you buy?
- iMil 4mo agoThis one: https://es.aliexpress.com/item/1005010123289822.html?spm=a2g0o.order_list.order_list_main.23.21ef1802iUNZPA&gatewayAdapt=glo2esp https://es.aliexpress.com/item/1005010123289822.html?spm=a2g...
- sieste 4mo agoThat's almost exactly my setup and I'm very happy with its performance. I noticed recently that I started to prefer my local Qwen3.6 35B A3B and pi agent over Claude Code. Both fail at different tasks, and Qwen more so than Claude. But the way Qwen fails is much more straightforward. In writing tasks Qwens hallucinations and bullshitting are much easier to spot because it doesn't have the sleek vocabulary and wordsmithing skills to disguise its ignorance. In coding tasks that Qwen can't solve it often just goes into a tool calling doom loop that the pi harness can catch, whereas Claude attempts ever more convoluted and creative things just making more and more mess that takes forever to clean up. I think part of the story is that the tasks for which I use AI are fairly simple and maybe don't need a frontier model. But I wonder if "proper" developers had similar experience?
- eurekin 4mo agoI keep finding more and more usecases for Q3.6 27b (same league) and the best performance is, when answers to my question is already in the context. The moment I'm trying something open-ended or ambitious, Claude/ChatGPT clearly take you to the goal quicker. For things, where there's a way to build a knowledgebase though, the local llm definitely can be a true contender. Plus, having a big context and no worries about filling it over and over - you can get quite far. I'm writing this, literally in between cooking a pasta, that the local llm ordered products for me online. I've built a grocery shopping skill, so that it roughly knows what I have in fridge (losely), my last 10 representative orders (general preferences plus rich info about shops and skus around me) and actual real-time in stock info. The last part has been my personal pet peeve for every product that promised cooking ingredient delivery (that is not packaged specifically for that). This is what has been promised to us by every big tech company with an agent, and now a local llms actually solved that for me fully.
- matthewfcarlson 4mo agoI keep playing around with this exact concept. While I don’t always trust entirely AI generated recipe, more traditional setups are super rigid when it comes to ingredients
- ydj 4mo ago80tp/s with 5080 3090 combo is wild. I’ve been working with a 4090 and two Tenstorrent p150 cards, and manage only about 30 tps utilizing all three for qwen3.6 27b q8. Guess I got more optimization to do. Would like to see the perf of their setup with and without mtp and ngram speculative decoding though, as well as parallel decode performance (once llamacpp mtp plays well with multiple slots). Being in California electricity alone puts this non-competitive with just paying a cloud though.
- manbart 4mo agoHow is the software compatibilty with the Tenstorrent cards? Are you stuck using vendor supplied runtimes/models? It's surprising how little these things come up given the price they go for
- ydj 4mo agoThe software stack is pretty immature, definitely very DIY. Their officially supported models are pretty old at this point, though there’s community support for gemma4, and models with GDN like qwen3.6 is supposedly very close. The entire stack (minus some binary blobs in firmware) is open source, so if you have the time and persistence you can get whatever you want done. A few community members have been working on support with llamacpp, where we can have supported operations offloaded to the TT cards, while having unsupported ops running on GPU or CPU. Llamacpp is pretty good at that. The existing kernels could definitely be better, and I’ll try my hand at writing some kernels some time.
- arjie 4mo agoThat’s the cost of using a new hardware provider. A single RTX Pro 6000 Blackwell Max-Q will do better than that and be much more usable. I have 2 running DS4 Flash at 160 tok/s with max num seqs 4. Very interesting though, these Tenstorrent chips. Might get one to experiment with.
- ydj 4mo agoYeah that’s definitely the smarter buy if you want to just have models running quickly. But the cost of 2 p150 and a 4090 was <$5000 for me. The main issue is the immature software, and somewhat baroque way of writing kernels. Please, buy one and join us.
- varispeed 4mo agoCould 2x RTX5080 work just as well?
- triwats 4mo agoPotential specs: NVIDIA GeForce RTX 5080: https://flopper.io/gpu/nvidia-geforce-rtx-5080-16gb https://flopper.io/gpu/nvidia-geforce-rtx-5080-16gb NVIDIA GeForce RTX 3090: https://flopper.io/gpu/nvidia-geforce-rtx-3090-24gb https://flopper.io/gpu/nvidia-geforce-rtx-3090-24gb
- stared 4mo agoI really like Qwen 3.6 27B Q8. On Apple Silicon, with MLX-LM, I am getting 20 tok/s with Macbook Max M5. Not sure how it compares to llama.cpp performance. In any case, while it is noticeably slower than this Nvidia RTX setup, being able to run such models on laptop is wild. Though, it heats my laptop rapidly.
- well_ackshually 4mo agoIt does come with one tiny little issue: it now draws 700W on full load. Just a single 5080 is enough to measurably heat up a room when loaded (320W draw at the wall on mine), and with that amount of power flowing through, you better have a good PSU as well as checking your power plugs themselves, these are going to get HOT when your entire setup is basically drawing 1kW.
- iMil 4mo agoI am actually surprised with the power draw, the box itself idles at 20W, which already amazes me for a Ryzen; when computing, I barely pass the 600W bar, and as I am not really using it to vibecode an entire system, I don't even notice the spikes on the power monitor (Shelly + homeassistant).
- nsbk 4mo ago[dead]
- washadjeffmad 4mo agoI've got a 4090 and 3090 in a node that peaks at 600W. If you're not power limiting in nvidia-smi, start.
- cybertim 4mo agoI bought two 3080/20gb and one of those MACHINIST X99 mainboards as well (one with two full x16 pcie slots) those boards come with a xeon cpu included (for the pcie lane support) it set me back 800 euros total (had a spare psu, ssd and mem in a drawer) and now im also happily running 80tk/s Qwen 3.6 Q8 (MTP).
- iMil 4mo agoGood call, I really hesitated between the X570 and the X99, are you using P2P?
- cybertim 4mo ago$ nvidia-smi topo -p2p r GPU0 GPU1 GPU0 X CNS GPU1 CNS X i guess not, i use llama.cpp with: --spec-draft-n-max 3 --spec-type draft-mtp --split-mode tensor --tensor-split 1,1 and my (gen) tk/s are between 60-80 tk/s will test this uncensored model and ngram added as well this weekend btw, i also set my powerlimit to 220watt per card (with nvidia-smi) that will cost you around 1 tk/s but safe you a LOT of power and heat :)
- iMil 4mo agoCNS means Chipset not supported and I doubt it is the case, are you sure you are using the patched nvidia module? modinfo nvidia to check which one is loaded
- cybertim 4mo agoI'm using bazzite on my ai-rig just because it has the gpu-optimized things setup (also nvidia-open). Looking at P2P seems to be available only for 90-versions of the nvidia rtx gpu line, not 80, and some versions of 50xx? (apparently the 5080?). Anyways, i downloaded that uncensored model and tweaked those kv settings etc. still getting 60-80tk/s but im able to get my context on 180224 now, used to be 131072 which gave me some trouble, this is already a win :)
- cybertim 4mo ago
- tonyrice 4mo agoIf I had an eGPU right now, I'd 100% be using Qwen
- skhameneh 4mo agoWould you mind giving these a try and let me know how they work for you? I’d imagine you would get better results and the latter will fit on a single GPU. https://huggingface.co/easiest-ai-shawn/Qwen3.6-27B-ExCal-EXL3 https://huggingface.co/easiest-ai-shawn/Qwen3.6-27B-ExCal-EX... https://huggingface.co/easiest-ai-shawn/Qwen3.6-27B-ExCal-Micro-EXL3 https://huggingface.co/easiest-ai-shawn/Qwen3.6-27B-ExCal-Mi... Do be sure to use dflash and/or mtp for the draft: https://huggingface.co/turboderp/Qwen3.6-27B-MTP-exl3 https://huggingface.co/turboderp/Qwen3.6-27B-MTP-exl3 https://huggingface.co/turboderp/Qwen3.6-27B-DFlash-exl3 https://huggingface.co/turboderp/Qwen3.6-27B-DFlash-exl3
- DiabloD3 4mo agoThe recommended values for Qwen 3.6 in thinking mode is `--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.00`, and `--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00` for coding/tool calling tasks, and for non-thinking, `--temp 0.7 -top-p 0.8 --top-k 20 --presence-penalty 1.5 --min-p 0.00`. The options listed are none of these. Also, the recommended Qwen MTP settings are `--spec-type draft-mtp --spec-draft-n-max 2`. 3 is not good on Nvidia hardware under different workloads. You can also add `ngram-mod`, but after `draft-mtp`; however, default `ngram-mod` settings aren't well tuned, and you want `--spec-ngram-mod-n-min 12 --spec-ngram-mod-n-max 16 --spec-ngram-mod-n-match 6` (defaults are 48, 64, 24; the ratio is good, the magnitude is suboptimal). Of abliterated Qwen 3.6 27B models, huihui's ends up being the worst. Try heretic instead. https://huggingface.co/mradermacher/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-i1-GGUF https://huggingface.co/mradermacher/Qwen3.6-27B-uncensored-h...
- aand16 4mo ago> You can also add `ngram-mod`, but after `draft-mtp` It looks like there's a hardcoded preference, CLI order is not important. (speculative.cpp:1322-1381): common_get_enabled_speculative_configs converts the types vector to a bitmask (order-independent). Then configs are added in a hardcoded priority order: ngram-simple ngram-map-k ngram-map-k4v ngram-mod ngram-cache draft-simple draft-eagle3 draft-mtp (speculative.cpp:1557-1603): common_speculative_draft iterates impls in the hardcoded priority order. Once an impl produces a draft for a sequence, later impls skip that sequence.
- DiabloD3 4mo agoInteresting.
- WeylandDarkStar 4mo agoSits in silence, watching China as they innovated a new type of ultra-thin gpu board and calling it 5090 "Turbos." Still waiting for Shenzhen listings to post a 5090 official verified with VBIOS crack...
- neals 4mo agoI tried implementing qwen through openrouter and deepinfra. Even without thinking, I had to wait 60s+ for the full result, where haiku or flash would be done in 5 or 6 seconds.
- verdyshd 4mo ago[flagged]
- irishcoffee 4mo agoIt is absolutely mind blowing to see some of the responses here. Open source, run-your-own, pay for nothing, we’re-all-nerds-that-buy-the-hardware-anyways ethos seems basically dead. I guess I’m getting old. I own two 16gb cards and I use them for models, for gpu-pasthru for gaming, 3d model rendering, etc. 14 year old me is mortified at this community.
- nullbio 4mo agoTimes are changing. The open-weight models have needed time to catch up, but they're finally at a point now where we can get almost frontier level capabilities for coding. I just wish we had a way to actually benchmark them properly though. Still seems no one has solved the problem of software architecture, brittleness and bloat as the codebase grows. Models love to add stuff, but they rarely clean up as they go. In a perfect world they'd do both near equally as they're developing. It would be nice if there was an "architecture quality" benchmark that distilled the essence of what it means to have a good architecture, but I suppose that's an open research question with a lot of variables? Like how is good architecture actually quantified and measured? Is there a mechanism that can be re-used across all codebases to clearly denote one that is good and one that is bad, or is it highly subjective and depend on the lens you're looking at it from? Is there a lot more to it than just "how much refactoring effort is required to extend this in the future?". Surely this is something that has been well researched - yet I never really hear anything about it. Makes me wonder why.
- irishcoffee 4mo ago> Surely this is something that has been well researched - yet I never really hear anything about it. Makes me wonder why. Occam’s razor rings true here: where’s the money in it?
- CamperBob2 4mo ago14 year old me is mortified at this community. Same here. There has to be someplace like this that's managed to cultivate a better crowd, but I'll be darned if I can find it.
- hanzeweiasa 4mo ago[flagged]
- mirekrusin 4mo agoon 2x 4090: 90 t/s for 27B Q8 256k context 260 t/s for 35B-A3B Q8 256k context
- tomekowal 4mo agoWith qwen3.6-35b-a3b-mtp using lm-studio on RTX 3090, I was getting 120tokens/s. The mtp (multi token prediction) is the key. I tired coding with Pi and it was much faster than Claude, but for any not-straightforward tasks, it did so so. Either looping itself or not realising easy to spot constraints. But for exploring codebases and asking questions about big stuff I find it better due to sheer speed.
- Havoc 4mo agoVery nice! Though if you're buying a X570 board I'd do crosshair viii dark hero - no buzzy chipset fan and can also do 2x8
- Dollarland 4mo ago[flagged]