13 ms·
Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
- modeless 1y agoWhat's the best speed people have gotten on 4090s?
- ActorNightly 1y agoYou can't fit the model into 4090 without quantization, its like 64 gigs. For home use, Gemma27B QAT is king. Its almost as good as Deepseek R1
- modeless 1y agoThe 20B one fits.
- steinvakt2 1y agoDoes it fit on a 5080 (16gb)?
- jwitthuhn 1y agoHaven't tried myself but it looks like it probably does. The weight files total 13.8 GB which gives you a little left over to hold your context.
- northern-lights 1y agoIt fits on a 5070TI, so should fit on a 5080 as well.
- SirMaster 1y agoYou don't really need it to fit all in VRAM due to the efficient MoE architecture and with llama.cpp The 120B is running at 20 tokens/sec on my 5060Ti 16GB with 64GB of system ram. Now personally I find 20 tokens/sec quite usable, but for some maybe it's not enough.
- dexterlagan 1y agoI have a similar setup but with 32 GB of RAM. Do you partly offload the model to RAM? Do you use LMStudio or other to achieve this? Thanks!
- SirMaster 1y agoYes, LMStudio and it automatically does this.
- asabla 1y agoI'm on a 5090 so it's not apples to apples comparison. But I'm getting ~150t/s for the 20B version using ~16000 context size.
- modeless 1y agoCool, what software?
- asabla 1y agoInitial testing has only been done with ollama. Plan on testing out llama.cpp and vllm when there is enough time
- steinvakt2 1y agoAnd flash attention doesn't work on 5090 yet, right? So currently 4090 is probably faster, or?
- PeterStuer 1y agoI don't think the 4090 has native 4bit support, which will probably have a significant impact.
- diggan 1y ago> And flash attention doesn't work on 5090 yet, right? Flash attention works with GPT-OSS + llama.cpp (tested on 1d72c8418) and other Blackwell card (RTX Pro 6000) so I think it should work on 5090 as well, it's the same architecture after all.
- deleted 1y ago[deleted]
- tmshapland 1y agoSuch a fascinating read. I didn't realize how much massaging needed to be done to get the models to perform well. I just sort of assumed they worked out of the box.
- acters 1y agoPersonally, I think bigger companies should be more proactive and work with some of the popular inference engine software devs with getting their special snowflake LLM to work before it gets released. I guess it is all very much experimental at the end of the day. Those devs are putting in God's work for us to use on our budget friendly hardware choices.
- eric-burel 1y agoSMEs are starting to want local LLMs and it's a nightmare to figure what hardware would work for what models. I am asking devs in my hometown to literally visit their installs to figure combos that work.
- CMCDragonkai 1y agoAre you installing them onsite?
- eric-burel 1y agoSome are asking that yeah but I haven't run an install yet, I am documenting the process. This is a last resort, hosting on European cloud is more efficient but some companies don't even want to hear about cloud hosting.
- mutkach 1y agoThis is a good take, actually. GPT-OSS is not much of a snowflake (judging by the model's architecture card at least) but TRT-LLM treats every model like that - there is too much hardcode - which makes it very difficult to just use it out-of-the-box for the hottest SotA thing.
- deleted 1y ago[deleted]
- magicalhippo 1y agoMaybe I'm especially daft this morning but I don't get the point of the speculative decoding. How does the target model validate the draft tokens without running the inference as normal? Because if it is doing just that, I don't get the point as you can't trust the draft tokens before they are validated, so you're still stuck waiting for the target model.
- joliu 1y agoIt does run inference, but on the batch of tokens that were drafted, akin to the prefill phase. So your draft model can decode N new tokens, then the real model does one inference pass to score the N new drafted tokens. Prefill is computation bound whereas decode is bandwidth bound, so in practice doing one prefill over N tokens is cheaper than doing N decode passes.
- furyofantares 1y agoNot an expert, but here's how I understand it. You know how input tokens are cheaper than output tokens? It's related to that. Say the model so far has "The capital of France". The small model generates "is Paris.", which let's say is 5 tokens. You feed the large model "The capital of France is Paris." to validate all 5 of those tokens in a single forward pass.
- ahmedfromtunis 1y agoBut what would happen if the small model's prediction was "is Rome."? Wouldn't that result in costlier inference if the small model is "wrong" more than it is correct. Also, if the small model would be sufficiently more "correct" than "wrong", wouldn't be more efficient to get rid of the large model at this point?
- cwyers 1y agoSo, the way speculative decoding works, the model begins predicting at the first wrong token, so you still get 'is' for free.
- deleted 1y ago
- littlestymaar 1y agoVery fast “Sorry I can't help with that” generator.
- jeffhuys 1y agoJust "liberate" it
- sarthaksoni 1y agoReading this made me realize how easy it is to set up GPT-OSS 20B in comparison. I had it running on my Mac in five minutes, thanks to Llama.
- DrPhish 1y agoIts also easy to do 120b on CPU if you have the resources. I had 120b running on my home LLM CPU inference box in just as long as it took to download the GGUFs, git pull and rebuild llama-server. I had it running at 40t/s with zero effort and 50t/s with a brief tweaking. Its just too bad that even the 120b isn't really worth running compared to the other models that are out there. It really is amazing what ggerganov and the llama.cpp team have done to democratize LLMs for individuals that can't afford a massive GPU farm worth more than the average annual salary.
- wkat4242 1y agoWhat hardware do you have? 50tk/s is really impressive for cpu.
- DrPhish 1y ago2xEPYC Genoa w/768GB of DDR5-4800 and an A5000 24GB card. I built it in January 2024 for about $6k and have thoroughly enjoyed running every new model as it gets released. Some of the best money I’ve ever spent.
- wkat4242 1y agoWow nice!! That's a really good deal for that much hardware. How many tokens/s do you get for DeepSeek-R1?
- DrPhish 1y agoThanks, it was a bit of a gamble at the time (lots of dodgy ebay parts), but it paid off. R1 starts at about 10t/s on an empty context but quickly falls off. I'd say the majority of my tokens are generating around 6t/s. Some of the other big MoE models can be quite a bit faster. I'm mostly using QwenCoder 480b at Q8 these days for 9t/s average. I've found I get better real-world results out of it than K2, R1 or GLM4.5.
- eric-burel 1y ago"Encourage Open-Source and Open-Weight AI" is the part just after "Ensure that Frontier AI Protects Free Speech and American Values" in America's AI Action Plan. I know this is not rational but OpenAI OSS models kinda give me chills as I am reading the Plan in parallel. Anyway I like seeing oss model providers talking about hardware, because that's a limiting point for most developers that are not familiar with this layer.
- geertj 1y ago> Ensure that Frontier AI Protects Free Speech and American Values I am in the early phases of collecting my thoughts on this topic so bear with me, but it this a bad thing? AI models will have a world view. I think I prefer them having a western world view, as that has built our modern society and has proven to be most successful in making the lives of people better. At the very minimum I would want a model to document its world view, and be aligned to it so that it does not try to socially engineer me to surreptitiously change mine.
- petesergeant 1y ago> but it this a bad thing? I think the worry is that there’s no fixed definitions here, so the executive can use this to exert partisan or ideological pressure on model providers. Every four years the models get RLHF’d to switch between thinking guns are amazing vs thinking guns are terrible.
- geertj 1y ago> Every four years the models get RLHF’d to switch between thinking guns are amazing vs thinking guns are terrible. I may be naive, but on this specific case, I am hoping that an AI could lead us to a somewhat objective truth. There seems to be enough data points to make some conclusion here. For example, most/all counties in Europe have less gun violence than the US, but there are at least two EU counties with high gun ownership (Finland and Austria) that also have low gun violence. The gun ownership issue is so polarized these days, I don’t think we can trust most people to make reason based arguments about it. Maybe an AI could help us synthesize and interpret the data dispassionately.
- hsaliak 1y agoTLDR: tensorrt
- mutkach 1y ago> Inspired by GPUs, we parallelized this effort across multiple engineers. One engineer tried vLLM, another SGLang, and a third worked on TensorRT-LLM. We were able to quickly get TensorRT-LLM working, which was fortunate as it is usually the most performant inference framework for LLMs. > TensorRT-LLM It is usually the hardest to setup correctly and is often out of the date regarding the relevant architectures. It also requires compiling the model on the exact same hardware-drivers-libraries stack as your production environment which is a great pain in the rear end to say the least. Multimodal setups also been a disaster - at least for a while - when it was near-impossible to make it work even for mainstream models - like Multimodal Llamas. The big question is whether it's worth it, since when running the GPT-OSS-120B on H100 using vLLM is flawless in comparison - and the throughput stays at 130-140 t/s for a single H100. (It's also somewhat a clickbait of a title - I was expecting to see 500t/s for a single GPU, when in fact it's just a tensor-parallel setup) It's also funny that they went for a separate release of TRT-LLM just to make sure that gpt-oss will work correctly, TRT-LLM is a mess
- philipkiely 1y agoTRT-LLM has its challenges from a DX perspective and yeah for Multi-modal we still use vLLM pretty often. But for the kind of traffic we are trying to serve -- high volume and latency sensitive -- it consistently wins head-to-head in our benchmarking and we have invested a ton of dev work in the tooling around it.
- wcallahan 1y agoI just used GPT-OSS-120B on a cross Atlantic flight on my MacBook Pro (M4, 128GB RAM). A few things I noticed: - it’s only fast with with small context windows and small total token context; once more than ~10k tokens you’re basically queueing everything for a long time - MCPs/web search/url fetch have already become a very important part of interacting with LLMs; when they’re not available the LLM utility is greatly diminished - a lot of CLI/TUI coding tools (e.g., opencode) were not working reliably offline at this time with the model, despite being setup prior to being offline That’s in addition to the other quirks others have noted with the OSS models.
- MoonObserver 1y agoM2 Max processor. I saw 60+ tok/s on short conversations, but it degraded to 30 tok/s as the conversation got longer. Do you know what actually accounts for this slowdown? I don’t believe it was thermal throttling.
- summarity 1y agoPhysics: You always have the same memory bandwidth. The longer the context, the more bits will need to pass through the same pipe. Context is cumulative.
- VierScar 1y agoNo I don't think it's the bits. I would say it's the computation. Inference requires performing a lot of matmul, and with more tokens the number of computation operations increases exponentially - O(n^2) at least. So increasing your context/conversation will quickly degrade performance I seriously doubt it's the throughput of memory during inference that's the bottleneck here.
- zozbot234 1y agoTypically, the token generation phase is memory-bound for LLM inference in general, and this becomes especially clear as context length increases (since the model's parameters are a fixed quantity.) If it was pure compute bound there would be huge gains to be had by shifting some of the load to the NPU (ANE) but AIUI it's just not so.
- blitzar 1y ago> widely-available H100 GPUs Just looked in the parts drawer at home and dont seem to have a $25,000 GPU for some inexplicable reason.
- KolmogorovComp 1y agoavailable != cheap
- blitzar 1y agoavailable /əˈveɪləbl/ adjective: available able to be used or obtained; at someone's disposal
- swexbe 1y agoYou can rent one from most cloud providers for a few bucks an hour.
- koakuma-chan 1y agoMight as well just use openai api
- deleted 1y ago[deleted]
- Kurtz79 1y agoDoes it even make sense calling them 'GPUs' (I just checked NVIDIA product page for the H100 and it is indeed so)? There should be a quicker way to differentiate between 'consumer-grade hardware that is mainly meant to be used for gaming and can also run LLMs inference in a limited way' and 'business-grade hardware whose main purpose is AI training or running inference for LLMs".
- smcleod 1y agoTensorRT-LLM is a right nightmare to setup and maintain. Good on them for getting it to work for them - but it's not for everyone.
- philipkiely 1y agoWe have built a ton of tooling on top of TRT-LLM and use it not just for LLMs but also for TTS models (Orpheus), STT models (Whisper), and embedding models.
- nektro 1y ago> we were the clear leader running on NVIDIA GPUs for both latency and throughput per public data from real-world use on OpenRouter. Baseten: 592.6 tps Groq: 784.6 tps Cerebras: 4,245 tps still impressive work
- philipkiely 1y agoYeah the custom hardware providers are super good at TPS. Kudos to their teams for sure, and the demos of instant reasoning are incredibly impressive. That said, we are serving the model at its full 131K context window, and they are serving 33K max, which could matter for some edge case prompts. Additionally, NVIDIA hardware is much more widely available if you are scaling a high-traffic application.
- lagrange77 1y agoWhile you're here.. Do you guys know a website that clearly shows which OS LLM models run on / fit into a specific GPU(setup)? The best heuristic i could find for the necessary VRAM is Number of Parameters × (Precision / 8) × 1.2 from here [0]. [0] https://medium.com/@lmpo/a-guide-to-estimating-vram-for-llms-637a7568d0ea https://medium.com/@lmpo/a-guide-to-estimating-vram-for-llms...
- diggan 1y agoMaybe I'm spoiled by having great internet connection, but I usually download the weights and try to run them via various tools (llama.cpp, LM Studio, vLLM and SGLang typically) and see what works. There seems to be so many variables involved (runners, architectures, implementations, hardware and so on) that none of the calculators I've tried so far been accurate, both in the way that they've over-estimated and under-estimated what I could run. So in the end, trying to actually run them seems to be the only fool-proof way of knowing for sure :)
- reactordev 1y agohuggingface has this built in if you care to fill out your software and hardware profile here: https://huggingface.co/settings/local-apps https://huggingface.co/settings/local-apps Then on the model pages, it will show you whether you can use it.
- diggan 1y agoInteresting, never knew about that! I filled out my details, then went to https://huggingface.co/openai/gpt-oss-120b https://huggingface.co/openai/gpt-oss-120b but I'm not sure if I see any difference? Where is it supposed to show if I can run it or not?
- reactordev 1y agoYou’ll see green check next to models you can use on the model card. https://huggingface.co/unsloth/gpt-oss-20b-GGUF https://huggingface.co/unsloth/gpt-oss-20b-GGUF
- philipkiely 1y agoWent to bed with 2 votes, woke up to this. Thank you so much HN!
- radarsat1 1y agoWould love to try fully local agentic coding. Is it feasible yet? I have a laptop with a 3050 but that's not nearly enough VRAM, I guess. Still, would be interested to know what's possible today on reasonable consumer hardware.
- zackangelo 1y agoGPT-OSS will run even faster on Blackwell chips because of its hardware support for fp4. If anyone is working on training or inference in Rust, I'm currently working on adding fp8 and fp4 support to cudarc[0] and candle[1]. This is being done so I can support these models in our inference engine for Mixlayer[2]. [0] https://github.com/coreylowman/cudarc/pull/449 https://github.com/coreylowman/cudarc/pull/449 [1] https://github.com/huggingface/candle/pull/2989 https://github.com/huggingface/candle/pull/2989 [2] https://mixlayer.com https://mixlayer.com
- diggan 1y agoAh, interesting. As someone with a RTX Pro 6000, is it ready today to be able to run gpt-oss-120b inference, or are there still missing pieces? Both linked PRs seems merged already, so unsure if it's ready to be played around with or not.
- mikewarot 1y agoYou know what's actually hard to find in all this? The actual dimensions of the arrays in the model GPT-OSS-120B. At least with statically typed languages, you know how big your arrays are at a glance. I'm trying to find it in the GitHub repo[1], and I'm not seeing it. I'm just trying to figure out how wide the datastream through this is, in particular, the actual data (not the weights) that flow through all of it. The width of the output stream. Just how big is a token at the output, prior to reducing it with "temperature" to a few bytes? Assume infinitely fast compute in a magic black box, but you have to send the output through gigabit ethernet... what's the maximum number of tokens per second? [1] https://github.com/openai/gpt-oss/tree/main/gpt_oss https://github.com/openai/gpt-oss/tree/main/gpt_oss
- amluto 1y agoWhat’s the application where you want to stream out the logits for each consecutive token while still sampling each token according to the usual rule? Keep in mind that, if you are doing the usual clever tricks like restricting the next token sampled to something that satisfies a grammar, you need to process the logits and sample them and return a token before running the next round of inference.
- mikewarot 1y agoI know the actual output of the model is wider than a token.... but I can't find it (the actual width, or number of bytes) in the source. Perhaps it's my very casual familiarity with Python that's limiting me, but I don't see any actual declarations of array sizes anywhere in the code. I'm just trying to calculate the actual bandwidth required for the full output of the model, not just a token to be handed off to the user. I need this so I can compute just what bandwidth a fully FPGA (later ASIC) based implementation of the model would result in. Edit/Append: I asked GPT-5, and it estimated: Total bytes = 50,000 tokens × 4 bytes/token = 200,000 bytes Which sounds about right to me. This yields a maximum of about 500 logits/second on Gigabit ethernet. The actual compute of the model is peanuts compared to just shuffling the data around.
- steeve 1y agoAccording to https://huggingface.co/openai/gpt-oss-120b/blob/main/config.json https://huggingface.co/openai/gpt-oss-120b/blob/main/config.... That’s 2880 values (so multiply by dtype)
- OldfieldFund 1y agolaughs in Cerebras
- adsharma 1y agoWhat's the best number on vLLM and SGlang so far on H100? It's sad that MLPerf takes a long time to catch up to SOTA models.
- Davidzheng 1y agoif I have a mac with 128Gb of integrated ram and I want to try this model, should I be using llama.cpp, mlx, or vllm, or something else? Sorry but I literally don't understand how I'm supposed to decide. Is it just compare inference speeds?