4 ms·
Was that cheaper than a Blackwell 6000? But yeah, 4x Blackwell 6000s are ~32-36k, not sure where the other $30k is going.
by ericd 7mo ago
Was that cheaper than a Blackwell 6000?
But yeah, 4x Blackwell 6000s are ~32-36k, not sure where the other $30k is going.
- segmondy 7mo agofolks have too much money than sense, gpt-oss-120b full quant runs on my quad 3090 at 100tk/sec and that's with llama.cpp, with vllm it will probably run at 150tk/sec and that's without batching.
- amarshall 7mo agoYou're almost certainly (definitely, in fact) confusing the 120b and 20b models.
- segmondy 7mo agoI'm most certainly not doing so. seg@seg-epyc:~/models$ du -sh * /llmzoo/models/* | sort -n 4.0K metrics.txt 4.0K opus 4.0K start_llama 8.2G nvidia_Orchestrator-8B-Q8_0.gguf 12K config.ini 34G Qwen3.5-27B 47G Qwen3.5-35B 51G Qwen3.5-27B-BF16 61G gpt-oss-120b-F16.gguf 65G Qwen3.5-35B-BF16 106G Qwen3.5-122B-Q6 117G GLM4.6V 175G MiniMax-M2.5 232G /llmzoo/models/small_models 240G Ernie4.5-300B 377G DeepSeekv3.2-nolight 380G /llmzoo/models/DeepSeek-V3.2-UD 400G /llmzoo/models/Qwen3.5-397B-Q8 424G /llmzoo/models/KimiK2Thinking 443G DeepSeek-Math-v2 443G DeepSeek-V3-0324-Q5 500G /llmzoo/models/GLM5-Q5 546G /llmzoo/models/KimiK2.5
- amarshall 7mo agoOh I missed the "quad" before 3090.
- ericd 7mo agoHow're you fitting a model made for 80 gig cards onto a GPU with 24 gigs at full quant?
- zozbot234 7mo agoMoE layers offload to CPU inference is the easiest way, though a bit of a drag on performance
- ericd 7mo agoYeah, I'd just be pretty surprised if they were getting 100 tokens/sec that way. EDIT: Either they edited that to say "quad 3090s", or I just missed it the first time.
- segmondy 7mo agoyou are correct, I did forget to add quad. you should join us in r/localllama check out what other people are getting. you're welcome. https://www.reddit.com/r/LocalLLaMA/comments/1nunq7s/gptoss120b_performance_on_4_x_3090/ https://www.reddit.com/r/LocalLLaMA/comments/1nunq7s/gptoss1... https://www.reddit.com/r/LocalLLaMA/comments/1p4evyr/most_economical_way_to_run_gptoss120b_for_10_users/ https://www.reddit.com/r/LocalLLaMA/comments/1p4evyr/most_ec...
- ericd 7mo agoThanks for the confirmation, wasn't sure if I was just going a bit senile heh. Yeah, I love /r/localllama, some of the best actual practitioners of this stuff on the internet. Also, crazy awesome frankenrigs to try and get that many huge cards working together. I was considering picking up a couple of the 48 gig 4090/3090s on an upcoming trip to China, but I just ended up getting one of the Max-Q's. But maybe the token throughput would still be higher with the 4090 route? Impressive numbers with those 3090s! What's the rig look like that's hosting all that?
- Havoc 7mo agoHe said quad 3090 not single
- ericd 7mo ago
- Aurornis 7mo ago> gpt-oss-120b full quant runs on my quad 3090 A 120B model cannot fit on 4 x 24GB GPUs at full quantization. Either you're confusing this with the 20B model, or you have 48GB modded 3090s.
- segmondy 7mo agoSome of you folks on here love to argue, gpt-oss-120b was trained in 4 bits, so it pretty much takes up 60gb.
- Aurornis 7mo agoGood point, but you still need KV cache and more. Fitting the model alone to RAM doesn’t get the job done.
- segmondy 7mo agoYeah, it doesn't take much. I'm looking at it right now, KV cache is about 4gb of vram, compute buffer =~ 1.5gb at full 128k context.
- ColonelPhantom 7mo agoGPT-OSS is tailored to be extremely memory efficient. Not only is it natively using the 4.25 bit per token MXFP4 format, but it also uses sliding window attention for half of its layers. It also doesn't have that many layers, only 36 for the 120B version and 24 for the 120B version. (The 120B is also much much sparser than the 20B.) I found a Reddit comment claiming only 36 KiB per token. With that, half a million tokens fits in 18 GB, which is less than one GPU. And three GPUs fit the parameters with room to spare (64 out of 72 GB).
- integralid 7mo agoThanks for chiming in. I'm looking for a reasonably cheap local LLM machine, and multiple 3090s is exactly what I planned to buy. Do you have any recommendations or recommend any reading material before I decide to spend money on that? edit: Found your comment about /r/localllama, but if you have anything more to add I'm still very interested.
- bastawhiz 7mo agoI bought the A100s used for a little over $6k each.
- ericd 7mo agoOh, why'd you go that route? Considering going beyond 80 gigs with nvlink or something?
- bastawhiz 7mo agoWhen the costs come down, I'll add two H100s. Until I have more work to saturate the GPUs, they're really at the limit of what I can make time to use them for. Give me a year of writing code and I'll have the need!