29 ms·
Open models by OpenAI
https://openai.com/index/introducing-gpt-oss/ https://openai.com/index/introducing-gpt-oss/
- jedisct1 1y agoFor some reason I'm less excited about this that I was with the Qwen models.
- zeld4 1y agoKnowledge cutoff: 2024-06 not a big deal, but still...
- anyg 1y ago[dead]
- gatienboquet 1y ago[flagged]
- snewman 1y agoNo, because there are lots of things people can do that it still can't do.
- sdenton4 1y ago"If it is still possible to put a goalpost somewhere - and we don't care where - then it's not AGI."
- dns_snek 1y agoLLMs are what they are, calling them "AGI" won't make them any more useful or exciting than they are, it's just going to devalue the term "AGI" which has revolutionary, disease-curing, humanity-saving connotations. What are you looking for us to say exactly? 1. We aren't even close to AGI and it's unclear that we'll ever get there, but it would change the course of humanity in a significant way if we ever do. 2. Wow we've reached AGI but now I'm realizing that AGI is lame, we need a new term for the humanity-saving sales pitch that we were promised!
- sdenton4 1y agoI think getting out of the binary is good for the long run. We have something which is artificial, intelligent, and general in scope. We're there. Is it perfect? No. Is it even good? Sometimes! Do airplanes flap their wings? Also no, but they do a lot of stuff nonetheless.
- dns_snek 1y agoThat's where we disagree, I do not consider a system that isn't capable of learning, improving, or reasoning to be generally intelligent. My most basic criteria for "AGI" is a system that can absorb and integrate new knowledge through repetition and experience in real time, just like a human would. Further, their statements, knowledge, and "beliefs" should be reasonably self-consistent. That's where I'm usually told that humans aren't self-consistent either, which is true! But if I ever met a human that was as inconsistent as LLMs usually are, I'd recommend that they get checked for brain damage. Of course the value of LLMs isn't binary, they're useful tools in many ways, but the sales pitch was always AGI == human-like, and not AGI == human-sounding, and that's quite clearly not where we are right now.
- sdenton4 1y agoYeah, this is in 'flies like a plane, not like a bird' territory. But I think it's closer than you think. The systems do learn and have improved rapidly over the last year. Humans have two learning modes - short-term in-context learning, and then longer-term learning that occurs with practice and across sleep cycles. In particular, humans tend to suck at new tasks until they've gotten in some practice and then slept on it (unless the new task is a minor deviation from a task they are already familiar with). This is true for LLM's as well. They have some ability to adapt to the context of the current conversation, but don't perform model weight updates at this stage. Weight updates happen over a longer period, as pre-training and fine-tuning data are updated. That longer-phase training is where we get the integration of new knowledge through repetition. In terms of reasoning, what we've got now is somewhere between a small child and a math prodigy, apparently, depending how much cash you're willing to burn on the results. But a small child is still a human.
- rvz 1y agono.
- thimabi 1y agoOpen weight models from OpenAI with performance comparable to that of o3 and o4-mini in benchmarks… well, I certainly wasn’t expecting that. What’s the catch?
- coreyh14444 1y agoBecause GPT-5 comes out later this week?
- thimabi 1y agoIt could be, but there’s so much hype surrounding the GPT-5 release that I’m not sure whether their internal models will live up to it. For GPT-5 to dwarf these just-released models in importance, it would have to be a huge step forward, and I’m still doubting about OpenAI’s capabilities and infrastructure to handle demand at the moment.
- sebzim4500 1y agoSurely OpenAI would not be releasing this now unless GPT-5 was much better than it.
- jona777than 1y agoAs a sidebar, I’m still not sure if GPT-5 will be transformative due to its capabilities as much as its accessibility. All it really needs to do to be highly impactful is lower the barrier of entry for the more powerful models. I could see that contributing to it being worth the hype. Surely it will be better, but if more people are capable of leveraging it, that’s just as revolutionary, if not more.
- rrrrrrrrrrrryan 1y agoIt seems like a big part of GPT-5 will be that it will be able to intelligently route your request to the appropriate model variant.
- Shank 1y ago
- DSingularity 1y agoHa. Secure funding and proceed to immediately make a decision that would likely conflict viscerally with investors.
- hnuser123456 1y agoMaybe someone got tired of waiting paid them to release something actually open
- 4b6442477b1280b 1y agotheir promise to release an open weights model predates this round of funding by, iirc, over half a year.
- deleted 1y ago[deleted]
- DSingularity 1y agoYeah but they never released until now.
- SV_BubbleTime 1y agoUndercutting other frontier models with your open source one is not an anti-investor move. It is what China has been doing for a year plus now. And the Chinese models are popular and effective, I assume companies are paying for better models. Releasing open models for free doesn’t have to be charity.
- hnuser123456 1y agoText only, when local multimodal became table stakes last year.
- ebiester 1y agoHonestly, it's a tradeoff. If you can reduce the size and make a higher quality in specific tasks, that's better than a generalist that can't run on a laptop or can't compete at any one task. We will know soon the actual quality as we go.
- greenavocado 1y agoThat's what I thought too until Qwen-Image was released
- SV_BubbleTime 1y agoWhen Queen-Image was released… like yesterday? And what? What point are you making? QwebImage was released yesterday and like every image model, its base model shows potential over older ones but the real factor is will it be flexible enough for a fine tune or additional training Loras.
- deleted 1y ago[deleted]
- BoorishBears 1y agoThe community can always figure out hooking it up to other modalities. Native might be better, but no native multimodal model is very competitive yet, so better to take a competitive model and latch on vision/audio
- tarruda 1y ago> so better to take a competitive model and latch on vision/audio Can this be done by a third party or would it have to be OpenAI?
- IceHegel 1y agoListed performance of ~5 points less than o3 on benchmarks is pretty impressive. Wonder if they feel the bar will be raised soon (GPT-5) and feel more comfortable releasing something this strong.
- johntiger1 1y agoWow, this will eat Meta's lunch
- seydor 1y agoI believe their competition is from chinese companies , for some time now
- mhh__ 1y agoThey will clone it
- BoorishBears 1y agoMaverick and Scout were not great, even with post-training in my experience, and then several Chinese models at multiple sizes made them kind of irrelevant (dots, Qwen, MiniMax) If anything this helps Meta: another model to inspect/learn from/tweak etc. generally helps anyone making models
- redox99 1y agoThere's nothing new here in terms of architecture. Whatever secret sauce is in the training.
- BoorishBears 1y agoPart of the secret sauce since O1 has been accesss the real reasoning traces, not the summaries. If you even glance at the model card you'll see this was trained on the same CoT RL pipeline as O3, and it shows in using the model: this is the most coherent and structured CoT of any open model so far. Having full access to a model trained on that pipeline is valuable to anyone doing post-training, even if it's just to observe, but especially if you use it as cold start data for your own training.
- anticensor 1y agoIts CoT is sadly closer to that sanitised o3 summaries than to R1 style traces.
- Workaccount2 1y agoWow, today is a crazy AI release day: - OAI open source - Opus 4.1 - Genie 3 - ElevenLabs Music
- orphea 1y agoOAI open source Yeah. This certainly was not on my bingo card.
- wahnfrieden 1y agoThey announced it months ago…
- satyrun 1y agowow I just listened to Eleven Music do flamenco singing. That is incredible. Edit. I just tried it though and less impressed now. We are really going to need major music software to get on board before we have actual creative audio tools. These all seem made for non-musicians to make a very cookie cutter song from a specific genre.
- tmikaeld 1y agoI also tried it for a full 100K credits (Wasted in 2 hours btw which is silly!). Compared to both Udio and Suno, it's very very bad.. both at compositions, matching lyrics to music, keeping tempo and as soon as there's any distorted instruments like guitars or live, quality goes to radio-level.
- BoxOfRain 1y ago>These all seem made for non-musicians to make a very cookie cutter song from a specific genre. This is my main problem with AI music at the moment, I'd love it if I had proper creative control as a musician that'd be amazing but a lot of the time it's just straight up slop generation.
- deviation 1y agoSo this confirms a best-in-class model release within the next few days? From a strategic perspective, I can't think of any reason they'd release this unless they were about to announce something which totally eclipses it?
- famouswaffles 1y agoEven before today, the last week or so, it's been clear for a couple reasons, that GPT-5's release was imminent.
- deleted 1y ago[deleted]
- ticulatedspline 1y agoEven without an imminent release it's a good strategy. They're getting pressure from Qwen and other high performing open-weight models. without a horse in the race they could fall behind in an entire segment. There's future opportunity in licensing, tech support, agents, or even simply to dominate and eliminate. Not to mention brand awareness, If you like these you might be more likely to approach their brand for larger models.
- logicchains 1y ago> I can't think of any reason they'd release this unless they were about to announce something which totally eclipses it Given it's only around 5 billion active params it shouldn't be a competitor to o3 or any of the other SOTA models, given the top Deepseek and Qwen models have around 30 billion active params. Unless OpenAI somehow found a way to make a model with 5 billion active params perform as well as one with 4-8 times more.
- bredren 1y agoUndoubtedly. It would otherwise reduce the perceived value of their current product offering. The question is how much better the new model(s) will need to be on the metrics given here to feel comfortable making these available. Despite the loss of face for lack of open model releases, I do not think that was a big enough problem t undercut commercial offerings.
- artembugara 1y agoDisclamer: probably dumb questions so, the 20b model. Can someone explain to me what I would need to do in terms of resources (GPU, I assume) if I want to run 20 concurrent processes, assuming I need 1k tokens/second throughput (on each, so 20 x 1k) Also, is this model better/comparable for information extraction compared to gpt-4.1-nano, and would it be cheaper to host myself 20b?
- mythz 1y agogpt-oss:20b is ~14GB on disk [1] so fits nicely within a 16GB VRAM card. [1] https://ollama.com/library/gpt-oss https://ollama.com/library/gpt-oss
- artembugara 1y agothanks, this part is clear to me. but I need to understand 20 x 1k token throughput I assume it just might be too early to know the answer
- Tostino 1y agoI legitimately cannot think of any hardware that will get you to that throughput over that many streams with any of the hardware I know of (I don't work in the server space so there may be some new stuff I am unaware of).
- artembugara 1y agooh, I totally understand that I'd need multiple GPUs. I'd just want to know what GPU specifically and how many
- Tostino 1y agoI don't think you can get 1k tokens/sec on a single stream using any consumer grade GPUs with a 20b model. Maybe you could with H100 or better, but I somewhat doubt that. My 2x 3090 setup will get me ~6-10 streams of ~20-40 tokens/sec (generation) ~700-1000 tokens/sec (input) with a 32b dense model.
- hubraumhugo 1y agoMeta's goal with Llama was to target OpenAI with a "scorched earth" approach by releasing powerful open models to disrupt the competitive landscape. Looks like OpenAI is now using the same playbook.
- tempay 1y agoIt seems like the various Chinese companies are far outplaying Meta at that game. It remains to be seen if they’re able to throw money at the problem to turn things around.
- SV_BubbleTime 1y agoGood move for China. No one was going to trust their models outright, now they not only have a track record, but they were able to undercut the value of US models at the same time.
- k2xl 1y agoIs there any details about hardware requirements for a sensible tokens per second for each size of these models?
- minimaxir 1y agoI'm disappointed that the smallest model size is 21B parameters, which strongly restricts how it can be run on personal hardware. Most competitors have released a 3B/7B model for that purpose. For self-hosting, it's smart that they targeted a 16GB VRAM config for it since that's the size of the most cost-effective server GPUs, but I suspect "native MXFP4 quantization" has quality caveats.
- moffkalast 1y agoEh 20B is pretty managable, 32GB of regular RAM and some VRAM will run you a 30B with partial offloading. After that it gets tricky.
- 4b6442477b1280b 1y agowith quantization, 20B fits effortlessly in 24GB with quantization + CPU offloading, non-thinking models run kind of fine (at about 2-5 tokens per second) even with 8 GB of VRAM sure, it would be great if we could have models in all sizes imaginable (7/13/24/32/70/100+/1000+), but 20B and 120B are great.
- Tostino 1y agoI am not at all disappointed. I'm glad they decided to go for somewhat large but reasonable to run models on everything but phones. Quite excited to give this a try
- strangecasts 1y agoA small part of me is considering going from a 4070 to a 16GB 5060 Ti just to avoid having to futz with offloading I'd go for an ..80 card but I can't find any that fit in a mini-ITX case :(
- GHanku 1y ago[dead]
- SV_BubbleTime 1y agoI wouldn’t stop at 16GB right now. 24 is the lowest I would go. Buy a used 3090. Picked one up for $700 a few months back, but I think they were on the rise then. The 3000 series can’t do FP8fast, but meh. It’s the OOM that’s tough, not the speed so much.
- Disposal8433 1y agoPlease don't use the open-source term unless you ship the TBs of data downloaded from Anna's Archive that are required do build it yourself. And dont forget all the system prompts to censor the multiple topics that they don't want you to see.
- deleted 1y ago[deleted]
- rvnx 1y agoI don’t know why you got so much downvoted, these models are not open-source/open-recipes. They are censored open weights models. Better than nothing, but far from being Open
- a_vanderbilt 1y agoMost people don't really care all that much about the distinction. It comes across to them as linguistic pedantry and they downvote it to show they don't want to hear/read it.
- outlore 1y agoby your definition most of the current open weight models would not qualify
- layer8 1y agoThat’s why they are called open weight and not open source.
- robotmaxtron 1y agoCorrect. I agree with them, most of the open weight models are not open source.
- someperson 1y agoKeep fighting the "open weights" terminology fight, because diluting the term open-source for a blob of neural network weights (even inference code is open-source) is not open-source.
- x187463 1y agoRunning a model comparable to o3 on a 24GB Mac Mini is absolutely wild. Seems like yesterday the idea of running frontier (at the time) models locally or on a mobile device was 5+ years out. At this rate, we'll be running such models in the next phone cycle.
- tedivm 1y agoIt only seems like that if you haven't been following other open source efforts. Models like Qwen perform ridiculously well and do so on very restricted hardware. I'm looking forward to seeing benchmarks to see how these new open source models compare.
- Rhubarrbb 1y agoAgreed, these models seem relatively mediocre to Qwen3 / GLM 4.5
- modeless 1y agoNah, these are much smaller models than Qwen3 and GLM 4.5 with similar performance. Fewer parameters and fewer bits per parameter. They are much more impressive and will run on garden variety gaming PCs at more than usable speed. I can't wait to try on my 4090 at home. There's basically no reason to run other open source models now that these are available, at least for non-multimodal tasks.
- tedivm 1y agoQwen3 has multiple variants ranging from larger (230B) than these models to significantly smaller (0.6b), with a huge number of options in between. For each of those models they also release quantized versions (your "fewer bits per parameter). I'm still withholding judgement until I see benchmarks, but every point you tried to make regarding model size and parameter size is wrong. Qwen has more variety on every level, and performs extremely well. That's before getting into the MoE variants of the models.
- emehex 1y agoSo 120B was Horizon Alpha and 20B was Horizon Beta?
- ImprobableTruth 1y agoUnfortunately not, this model is noticeably worse. I imagine horizon is either gpt 5 nano/mini.
- Leary 1y agoGPQA Diamond: gpt-oss-120b: 80.1%, Qwen3-235B-A22B-Thinking-2507: 81.1% Humanity’s Last Exam: gpt-oss-120b (tools): 19.0%, gpt-oss-120b (no tools): 14.9%, Qwen3-235B-A22B-Thinking-2507: 18.2%
- jasonjmcghee 1y agoWow - I will give it a try then. I'm cynical about OpenAI minmaxing benchmarks, but still trying to be optimistic as this in 8bit is such a nice fit for apple silicon
- modeless 1y agoEven better, it's 4 bit
- amarcheschi 1y agoGlm 4.5 seems on par as well
- thegeomaster 1y agoGLM-4.5 seems to outperform it on TauBench, too. And it's suspicious OAI is not sharing numbers for quite a few useful benchmarks (nothing related to coding, for example). One positive thing I see is the number of parameters and size --- it will provide more economical inference than current open source SOTA.
- lcnPylGDnU4H9OF 1y agoWas the Qwen model using tools for Humanity's Last Exam?
- chown 1y agoShameless plug: if someone wants to try it in a nice ui, you could give Msty[1] a try. It's private and local. [1]: https://msty.ai https://msty.ai
- dsco 1y agoDoes anyone get the demos at https://www.gpt-oss.com https://www.gpt-oss.com to work, or are the servers down immediately after launch? I'm only getting the spinner after prompting.
- eliseumds 1y agoGetting lots of 502s from `https://api.gpt-oss.com/chatkit https://api.gpt-oss.com/chatkit` at the moment.
- deleted 1y ago[deleted]
- lukasgross 1y ago(I helped build the microsite) Our backend is falling over from the load, spinning up more resources!
- anticensor 1y agoWhy isn't GPT-OSS also offered on the free tier of ChatGPT?
- lukasgross 1y agoUpdate: try now!
- MutedEstate45 1y agoThe repeated safety testing delays might not be purely about technical risks like misuse or jailbreaks. Releasing open weights means relinquishing the control OpenAI has had since GPT-3. No rate limits, no enforceable RLHF guardrails, no audit trail. Unlike API access, open models can't be monitored or revoked. So safety may partly reflect OpenAI's internal reckoning with that irreversible shift in power, not just model alignment per se. What do you guys think?
- BoorishBears 1y agoI think it's pointless: if you SFT even their closed source models on a specific enough task, the guardrails disappear. AI "safety" is about making it so that a journalist can't get out a recipe for Tabun just by asking.
- MutedEstate45 1y agoTrue, but there's still a meaningful difference in friction and scale. With closed APIs, OpenAI can monitor for misuse, throttle abuse and deploy countermeasures in real-time. With open weights, a single prompt jailbreak or exploit spreads instantly. No need for ML expertise, just a Reddit post. The risk isn’t that bad actors suddenly become smarter. It’s that anyone can now run unmoderated inference and OpenAI loses all visibility into how the model’s being used or misused. I think that’s the control they’re grappling with under the label of safety.
- BoorishBears 1y agoOpenAI and Azure both have zero retention options, and the NYT saga has given pretty strong confirmation they meant it when they said zero.
- MutedEstate45 1y agoI think you're conflating real-time monitoring with data retention. Zero retention means OpenAI doesn't store user data, but they can absolutely still filter content, rate limit and block harmful prompts in real-time without retaining anything. That's processing requests as they come in, not storing them. The NYT case was about data storage for training/analysis not about real-time safety measures.
- ahmedhawas123 1y agoExciting as this is to toy around with... Perhaps I missed it somewhere, but I find it frustrating that, unlike most other open weight models and despite this being an open release, OpenAI has chosen to provide pretty minimal transparency regarding model architecture and training. It's become the norm for LLama, Deepseek, Qwenn, Mistral and others to provide a pretty detailed write up on the model which allows researchers to advance and compare notes.
- sebzim4500 1y agoThe model files contain an exact description of the architecture of the network, there isn't anything novel. Given these new models are closer to the SOTA than they are to competing open models, this suggests that the 'secret sauce' at OpenAI is primarily about training rather than model architecture. Hence why they won't talk about the training.
- gundawar 1y agoTheir model card [0] has some information. It is quite a standard architecture though; it's always been that their alpha is in their internal training stack. [0] https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7637/oai_gpt-oss_model_card.pdf https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7...
- ahmedhawas123 1y agoThis is super helpful and I had not seen it, thanks so much for sharing! And I hear you on training being an alpha, at the size of the model I wonder how much of this is distillation and using o3/o4 data.
- sadiq 1y agoLooks like Groq (at 1k+ tokens/second) and Fireworks are already live on openrouter: https://openrouter.ai/openai/gpt-oss-120b https://openrouter.ai/openai/gpt-oss-120b $0.15M in / $0.6-0.75M out edit: Now Cerebras too at 3,815 tps for $0.25M / $0.69M out.
- spott 1y agoIt is interesting that openai isn't offering any inference for these models.
- bangaladore 1y agoMakes sense to me. Inference on these models will be a race to the bottom. Hosting inference themselves will be a waste of compute / dollar for them.
- podnami 1y agoWow this was actually blazing fast. I prompted "how can the 45th and 47th presidents of america share the same parents?" On ChatGPT.com o3 thought for for 13 seconds, on OpenRouter GPT OSS 120B thought for 0.7 seconds - and they both had the correct answer.
- Imustaskforhelp 1y agoNot gonna lie but I got sorta goosebumps I am not kidding but such progress from a technological point of view is just fascinating!
- swores 1y agoI'm not sure that's a particularly good question for concluding something positive about the "thought for 0.7 seconds" - it's such a simple answer, ChatGPT 4o (with no thinking time) immediately answered correctly. The only surprising thing in your test is that o3 wasted 13 seconds thinking about it.
- Workaccount2 1y ago
- modeless 1y agoCan't wait to see third party benchmarks. The ones in the blog post are quite sparse and it doesn't seem possible to fully compare to other open models yet. But the few numbers available seem to suggest that this release will make all other non-multimodal open models obsolete.
- incomingpain 1y agoI dont see the unsloth files yet but they'll be here: https://huggingface.co/unsloth/gpt-oss-20b-GGUF https://huggingface.co/unsloth/gpt-oss-20b-GGUF Super excited to test these out. The benchmarks from 20B are blowing away major >500b models. Insane. On my hardware. 43 tokens/sec. I got an error with flash attention turning on. Cant run it with flash attention? 31,000 context is max it will allow or model wont load. no kv or v quantization.
- rmonvfer 1y agoWhat a day! Models aside, the Harmony Response Format[1] also seems pretty interesting and I wonder how much of an impact it might have in performance of these models. [1] https://github.com/openai/harmony https://github.com/openai/harmony
- incomingpain 1y agoSeems to be breaking every agentic tool I've tried so far. Im guessing it's going to very rapidly be patched into the various tools.
- mikert89 1y agoACCELERATE
- n42 1y agomy very early first impression of the 20b model on ollama is that it is quite good, at least for the code I am working on; arguably good enough to drop a subscription or two
- jakozaur 1y agoThe coding seems to be one of the strongest use cases for LLMs. Though currently they are eating too many tokens to be profitable. So perhaps these local models could offload some tasks to local computers. E.g. Hybrid architecture. Local model gathers more data, runs tests, does simple fixes, but frequently asks the stronger model to do the real job. Local model gathers data using tools and sends more data to the stronger model. It
- Imustaskforhelp 1y agoI have always thought that if we can somehow get an AI which is insanely good at coding, so much so that It can improve itself, then through continuous improvements, they will get better models of everything else idk Maybe you guys call it AGI, so anytime I see progress in coding, I think it goes just a tiny bit towards the right direction Plus it also helps me as a coder to actually do some stuff just for the fun. Maybe coding is the only truly viable use of AI and all others are negligible increases. There is so much polarization in the use of AI on coding but I just want to say this, it would be pretty ironic that an industry which automates others job is this time the first to get their job automated. But I don't see that as an happening, far from it. But still each day something new, something better happens back to back. So yeah.
- hooverd 1y agoOptimistically, there's always more crap to get done.
- jona777than 1y agoI agree. It’s not improbable for there to be _more_ needs to meet in the future, in my opinion.
- NitpickLawyer 1y agoNot to open that can of worms, but in most definitions self-improvement is not an AGI requirement. That's already ASI territory (Super Intelligence). That's the proverbial skynet (pessimists) or singularity (optimists).
- Imustaskforhelp 1y agoIs this the same model (Horizon Beta) on openrouter or not? Because I still see Horizon beta available with its codename on openrouter
- abidlabs 1y agoTest it with a web UI: https://huggingface.co/spaces/abidlabs/openai-gpt-oss-120b-test https://huggingface.co/spaces/abidlabs/openai-gpt-oss-120b-t...
- ArtTimeInvestor 1y agoWhy do companies release open source LLMs? I would understand it, if there was some technology lock-in. But with LLMs, there is no such thing. One can switch out LLMs without any friction.
- gnulinux 1y agoName recognition? Advertisement? Federal grant to beat Chinese competition? There could be many legitimate reasons, but yeah I'm very surprised by this too. Some companies take it a bit too seriously and go above and beyond too. At this point unless you need the absolute SOTA models because you're throwing LLM at an extremely hard problem, there is very little utility using larger providers. In OpenRouter, or by renting your own GPU you can run on-par models for much cheaper.
- TrackerFF 1y agoLLMs are terrible, purely speaking from the business economic side of things. Frontier / SOTA models are barely profitable. Previous gen model lose 90% of their value. Two gens back and they're worthless. And given that their product life cycle is something like 6-12 months, you might as well open source them as part of sundowning them.
- spongebobstoes 1y agoinference runs at a 30-40% profit
- mclau157 1y agoPartially because using their own GPUs is expensive, so maybe offloading some GPU usage
- koolala 1y agoThey don't because it would kill their data scrapping buisness's competitive advantage.
- LordDragonfang 1y agoZuckerberg explains a few of the reasons here: https://www.dwarkesh.com/p/mark-zuckerberg#:~:text=As%20long%20as%20it%27s%20helping%20us%20then%20yeah. https://www.dwarkesh.com/p/mark-zuckerberg#:~:text=As%20long... The short version is that is you give a product to open source, they can and will donate time and money to improving your product, and the ecosystem around it, for free, and you get to reap those benefits. Llama has already basically won that space (the standard way of running open models is llama.cpp), so OpenAI have finally realized they're playing catch-up (and last quarter's SOTA isn't worth much revenue to them when there's a new SOTA, so they may as well give it away while it can still crack into the market)
- HanClinto 1y agoHoly smokes, there's already llama.cpp support: https://github.com/ggml-org/llama.cpp/pull/15091 https://github.com/ggml-org/llama.cpp/pull/15091
- carbocation 1y agoAnd it's already on ollama, it appears: https://ollama.com/library/gpt-oss https://ollama.com/library/gpt-oss
- incomingpain 1y agolm studio immediately released the new appimage with support.
- jp1016 1y agoi wish these models had a minimum ram , cpu and gpu size listed on the site instead of high end and medium end pc.
- phh 1y agoYou can technically run it on a 8086 assuming you can get access to a big enough storage. More reasonably, you should be able to run the 20B at non-stupidly-slow speed with a 64bit CPU, 8GB RAM, 20GB SSD.
- pamelafox 1y agoAnyone tried running on a Mac M1 with 16GB RAM yet? I've never run higher than an 8GB model, but apparently this one is specifically designed to work well with 16 GB of RAM.
- thimabi 1y agoIt works fine, although with a bit more latency than non-local models. However, swap usage goes way beyond what I’m comfortable with, so I’ll continue to use smaller models for the foreseeable future. Hopefully other quantizations of these OpenAI models will be available soon.
- pamelafox 1y agoUpdate: I tried it out. It took about 8 seconds per token, and didn't seem to be using much of my GPU (MPU), but was using a lot of RAM. Not a model that I could use practically on my machine.
- steinvakt2 1y agoDid you run it the best way possible? im no expert, but I understand it can affect inference time greatly (which format/engine is used)
- pamelafox 1y agoI ran it via Ollama, which I assume uses the best way. Screenshot in my post here: https://bsky.app/profile/pamelafox.bsky.social/post/3lvobol3jfb2r https://bsky.app/profile/pamelafox.bsky.social/post/3lvobol3... I'm still wondering why my MPU usage was so low.. maybe Ollama isn't optimized for running it yet?
- wahnfrieden 1y agoMight need to wait on MLX
- turnsout 1y ago
- shpongled 1y agoI looked through their torch implementation and noticed that they are applying RoPE to both query and key matrices in every layer of the transformer - is this standard? I thought positional encodings were usually just added once at the first layer
- m_ke 1y agoNo they’re usually done at each attention layer.
- shpongled 1y agoDo you know when this was introduced (or which paper)? AFAIK it's not that way in the original transformer paper, or BERT/GPT-2
- Scene_Cast2 1y agoShould be in the RoPE paper. The OG transformers used multiplicative sinusoidal embeddings, while RoPE does a pairwise rotation. There's also NoPE, I think SmolLM3 "uses NoPE" (aka doesn't use any positional stuff) every fourth layer.
- Nimitz14 1y agoThis is normal. Rope was introduced after bert/gpt2
- spott 1y agoAll the Llamas have done it (well, 2 and 3, and I believe 1, I don't know about 4). I think they have a citation for it, though it might just be the RoPE paper (https://arxiv.org/abs/2104.09864 https://arxiv.org/abs/2104.09864). I'm not actually aware of any model that doesn't do positional embeddings on a per-layer basis (excepting BERT and the original transformer paper, and I haven't read the GPT2 paper in a while, so I'm not sure about that one either).
- shpongled 1y ago
- jstummbillig 1y agoShoutout to the hn consensus regarding an OpenAI open model release from 4 days ago: https://news.ycombinator.com/item?id=44758511 https://news.ycombinator.com/item?id=44758511
- kingkulk 1y agoWelcome to the future!
- timmg 1y agoOrthogonal, but I just wanted to say how awesome Ollama is. It took 2 seconds to find the model and a minute to download and now I'm using it. Kudos to that team.
- _ache_ 1y agoTo be fair, it's with the help of OpenAI. They did it together, before the official release. https://ollama.com/blog/gpt-oss https://ollama.com/blog/gpt-oss
- aubanel 1y agoFrom experience, it's much more engineering work on the integrator's side than on OpenAI's. Basically they provide you their new model in advance, but they don't know the specifics of your system, so it's normal that you do most of the work. Thus I'm particularly impressed by Cerebras: they only have a few models supported for their extreme perf inference, it must have been huge bespoke work to integrate.
- Shopper0552 1y agoI remember reading Ollama is going closed source now? https://www.reddit.com/r/LocalLLaMA/comments/1meeyee/ollamas_new_gui_is_closed_source/ https://www.reddit.com/r/LocalLLaMA/comments/1meeyee/ollamas...
- int_19h 1y agoIt's just as easy with LM Studio. All the real heavy lifting is done by llama.cpp, and for the distribution, by HuggingFace.
- PeterStuer 1y agoI love how they frame High-end desktops and laptops as having "a single H100 GPU".
- organsnyder 1y agoI read that as it runs in data centers (H100 GPUs) or high-end desktops/laptops (Strix Halo?).
- xyc 1y agoI'm running it with ROG Flow Z13 128GB Strix Halo and getting 50 tok/s for 20B model and 12 tok/s for 120B model. I'd say it's pretty usable.
- organsnyder 1y agoExcellent! I have a Framework Desktop with 128GB on preorder—really looking forward to getting it.
- robertheadley 1y agoI actually tried to ask the Model about that, then I asked ChatGPT, both times, they just said that it was marketing speak. I was like no. It is false advertising.
- phh 1y agoWell if nVidia wasn't late, it would be runnable on nVidia project Digits.
- PeterStuer 1y agoYes, they are late to the party. Maybe they do not want to eat into the RTX Pro 6000 sales. In the meantime, there is the AMD Ryzen™ Al Max+ 395.
- piskov 1y agoDon’t forget about mac studio
- kgwgk 1y agoIt may be useless for many use cases given that its policy prevents it for example from providing "advice or instructions about how to buy something." (I included details about its refusal to answer even after using tools for web searching but hopefully shorter comment means fewer downvotes.)
- deleted 1y ago[deleted]
- isoprophlex 1y agoCan these do image inputs as well? I can't find anything about that on the linked page, so I guess not..?
- cristoperb 1y agoNo, they're text only
- pu_pe 1y agoVery sparse benchmarking results released so far. I'd bet the Chinese open source models beat them on quite a few of them.
- foundry27 1y agoModel cards, for the people interested in the guts: https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7637/oai_gpt-oss_model_card.pdf https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7... In my mind, I’m comparing the model architecture they describe to what the leading open-weights models (Deepseek, Qwen, GLM, Kimi) have been doing. Honestly, it just seems “ok” at a technical level: - both models use standard Grouped-Query Attention (64 query heads, 8 KV heads). The card talks about how they’ve used an older optimization from GPT3, which is alternating between banded window (sparse, 128 tokens) and fully dense attention patterns. It uses RoPE extended with YaRN (for a 131K context window). So they haven’t been taking advantage of the special-sauce Multi-head Latent Attention from Deepseek, or any of the other similar improvements over GQA. - both models are standard MoE transformers. The 120B model (116.8B total, 5.1B active) uses 128 experts with Top-4 routing. They’re using some kind of Gated SwiGLU activation, which the card talks about as being "unconventional" because of to clamping and whatever residual connections that implies. Again, not using any of Deepseek’s “shared experts” (for general patterns) + “routed experts” (for specialization) architectural improvements, Qwen’s load-balancing strategies, etc. - the most interesting thing IMO is probably their quantization solution. They did something to quantize >90% of the model parameters to the MXFP4 format (4.25 bits/parameter) to let the 120B model to fit on a single 80GB GPU, which is pretty cool. But we’ve also got Unsloth with their famous 1.58bit quants :) All this to say, it seems like even though the training they did for their agentic behavior and reasoning is undoubtedly very good, they’re keeping their actual technical advancements “in their pocket”.
- rfoo 1y agoOr, you can say, OpenAI has some real technical advancements on stuff besides attn architecture. GQA8, alternating SWA 128 / full attn do all seem conventional. Basically they are showing us that "no secret sauce in model arch you guys just sucks at mid/post-training", or they want us to believe this. The model is pretty sparse tho, 32:1.
- liuliu 1y agoKimi K2 paper said that the model sparsity scales up with parameters pretty well (MoE sparsity scaling law, as they call, basically calling Llama 4 MoE "done wrong"). Hence K2 has 128:1 sparsity.
- user_7832 1y agoNewbie question: I remember folks talking about how kimi 2’s launch might have pushed OpenAI to launch their model later. Now that we (shortly will) know how this model performs, how do they stack up? Did openAI likely actually hold off releasing weights because of kimi, in retrospect?
- ClassAndBurn 1y agoOpen models are going to win long-term. Anthropics' own research has to use OSS models [0]. China is demonstrating how quickly companies can iterate on open models, allowing smaller teams access and augmentation to the abilities of a model without paying the training cost. My personal prediction is that the US foundational model makers will OSS something close to N-1 for the next 1-3 iterations. The CAPEX for the foundational model creation is too high to justify OSS for the current generation. Unless the US Gov steps up and starts subsidizing power, or Stargate does 10x what it is planned right now. N-1 model value depreciates insanely fast. Making an OSS release of them and allowing specialized use cases and novel developments allows potential value to be captured and integrated into future model designs. It's medium risk, as you may lose market share. But also high potential value, as the shared discoveries could substantially increase the velocity of next-gen development. There will be a plethora of small OSS models. Iteration on the OSS releases is going to be biased towards local development, creating more capable and specialized models that work on smaller and smaller devices. In an agentic future, every different agent in a domain may have its own model. Distilled and customized for its use case without significant cost. Everyone is racing to AGI/SGI. The models along the way are to capture market share and use data for training and evaluations. Once someone hits AGI/SGI, the consumer market is nice to have, but the real value is in novel developments in science, engineering, and every other aspect of the world. [0] https://www.anthropic.com/research/persona-vectors https://www.anthropic.com/research/persona-vectors > We demonstrate these applications on two open-source models, Qwen 2.5-7B-Instruct and Llama-3.1-8B-Instruct.
- lechatonnoir 1y agoI'm pretty sure there's no reason that Anthropic has to do research on open models, it's just that they produced their result on open models so that you can reproduce their result on open models without having access to theirs.
- Adrig 1y agoI'm a layman but it seemed to me that the industry is going towards robust foundational models on which we plug tools, databases, and processes to expand their capabilities. In this setup OSS models could be more than enough and capture the market but I don't see where the value would be to a multitude of specialized models we have to train.
- deleted 1y ago[deleted]
- mythz 1y agoGetting great performance running gpt-oss on 3x A4000's: gpt-oss:20b = ~46 tok/s More than 2x faster than my previous leading OSS models: mistral-small3.2:24b = ~22 tok/s gemma3:27b = ~19.5 tok/s Strangely getting nearly the opposite performance running on 1x 5070 Ti: mistral-small3.2:24b = ~39 tok/s gpt-oss:20b = ~21 tok/s Where gpt-oss is nearly 2x slow vs mistral-small 3.2.
- genpfault 1y agoSeeing ~70 tok/s on a 7900 XTX using Ollama.
- Matsta 1y agoI'm getting around 90 tok/s on a 3090 using Ollama. Pretty impressive
- mythz 1y agook issue is with ollama as gpt-oss 20B runs much faster on 1x 5070 Ti with llama.cpp and LM Studio: llama-server = ~181 tok/s LM Studio = ~46 tok/s (default) LM Studio Custom = ~158 tok/s (changed to offload to GPU and switch to CUDA llama.cpp engine) and llama-server on my 3x A4000 GPU Server is getting 90 tok/s vs 46 tok/s on ollama
- anonymoushn 1y agoguys, what does OSS stand for?
- thejazzman 1y agoit's a marketing term that modern companies use to grow market share
- ayakaneko 1y agoshould be open source software, but it's a model, so not sure whether they chose this name with the last S having other meanings.
- Robdel12 1y agoI’m on my phone and haven’t been able to break away to check, but anyone plug these into Codex yet?
- jcmontx 1y agoI'm out of the loop for local models. For my M3 24gb ram macbook, what token throughput can I expect? Edit: I tried it out, I have no idea in terms of of tokens but it was fluid enough for me. A bit slower than using o3 in the browser but definitely tolerable. I think I will set it up in my GF's machine so she can stop paying for the full subscription (she's a non-tech professional)
- steinvakt2 1y agoWondering about the same for my M4 max 128 gb
- jcmontx 1y agoIt should fly on your machine
- steinvakt2 1y agoYeah, was super quick and easy to set up using Ollama. I had to kill some processes first to avoid memory swap though (even with 128gb memory). So a slightly more quantized version is maybe ideal, for me at least. Edit: I'm talking about the 120B model of course
- coolspot 1y ago40 t/s
- dantetheinferno 1y agoApple M4 Pro w/ 48GB running the smaller version. I'm getting 43.7t/s
- GHanku 1y ago[dead]
- albertgoeswoof 1y ago3 year old M1 MacBook Pro 32gb, 42 tokens/sec on lm studio Very much usable
- Rhubarrbb 1y agoWhat's the best agent to run this on? Is it compatible with Codex? For OSS agents, I've been using Qwen Code (clunky fork of Gemini), and Goose.
- wahnfrieden 1y agoWhy not Claude Code?
- objektif 1y agoI keep hitting the limit within an hour.
- wahnfrieden 1y agoMeant with your own model
- henriquegodoy 1y agoSeeing a 20B model competing with o3's performance is mind blowing like just a year ago, most of us would've called this impossible - not just the intelligence leap, but getting this level of capability in such a compact size. I think that the point that makes me more excited is that we can train trillion-parameter giants and distill them down to just billions without losing the magic. Imagine coding with Claude 4 Opus-level intelligence packed into a 10B model running locally at 2000 tokens/sec - like instant AI collaboration. That would fundamentally change how we develop software.
- coolspot 1y ago10B * 2000 t/s = 20,000 GB/s memory bandwidth . Apple hardware can do 1k GB/s .
- oezi 1y agoThat’s why MoE is needed.
- int_19h 1y agoIt's not even a 20b model. It's 20b MoE with 3.6b active params. But it does not actually compete with o3 performance. Not even close. As usual, the metrics are bullshit. You don't know how good the model actually is until you grill it yourself.
- Nimitz14 1y agoI'm surprised at the model dim being 2.8k with an output size of 200k. My gut feeling had told me you don't want too large of a gap between the two, seems I was wrong.
- ukprogrammer 1y ago> we also introduced an additional layer of evaluation by testing an adversarially fine-tuned version of gpt-oss-120b What could go wrong?
- nirav72 1y agoI don't exactly have the ideal hardware to run locally - but just ran the 20b in LMStudio with a 3080 Ti (12gb vram) with some offloading to CPU. Ran couple of quick code generation tests. On average about 20t/sec. But response quality was very similar or on-par with chatgpt o3 for the same code it outputted. So its not bad.
- nodesocket 1y agoAnybody got this working in Ollama? I'm running latest version 0.11.0 with WebUI v0.6.18 but getting: > List the US presidents in order starting with George Washington and their time in office and year taken office. >> 00: template: :3: function "currentDate" not defined
- genpfault 1y agohttps://github.com/ollama/ollama/issues/11673 https://github.com/ollama/ollama/issues/11673
- jmorgan 1y agoSorry about this. Re-downloading Ollama should fix the error
- nodesocket 1y agoThanks for the reply and speedy patch Jeffery. Seems to be working now, except my 4060ti can’t hang lacking enough vram.
- ahmetcadirci25 1y agoI started downloading, I'm eager to test it. I will share my personal experiences. https://ahmetcadirci.com/2025/gpt-oss/ https://ahmetcadirci.com/2025/gpt-oss/
- koolala 1y agoCalls them open-weight. Names them 'oss'. What does oss stand for?
- incomingpain 1y agoFirst coding test: Just going copy and paste out of chat. It aced my first coding test in 5 seconds... this is amazing. It's really good at coding. Trying to use it for agentic coding... lots of fail. This harmony formatting? Anyone have a working agentic tool? openhands and void ide are failing due to the new tags. Aider worked, but the file it was supposed to edit was untouched and it created Create new file? (Y)es/(N)o [Yes]: Applied edit to <|end|><|start|>assistant<|channel|>final<|message|>main.py so the file name is '<|end|><|start|>assistant<|channel|>final<|message|>main.py' lol. quick rename and it was fantastic. I think qwen code is the best choice so far but unreliable. So far these new tags are coming through but it's working properly; sometimes. 1 of my tests so far has been able to get 20b not to succeed the first iteration; but a small followup and it was able to completely fix it right away. Very impressive model for 20B.
- bobsmooth 1y agoHopefully the dolphin team will work their magic and uncensor this model
- siliconc0w 1y agoIt seems like OSS will win, I can't see people willing to pay like 10x the price for what seems like 10% more performance. Especially once we get better at routing the hardest questions to the better models and then using that response to augment/fine-tune the OSS ones.
- n42 1y agoto me it seems like the market is breaking into an 80/20 of B2C/B2B; the B2C use case becoming OSS models (the market shifts to devices that can support them), and the B2B market being priced appropriately for businesses that require that last 20% of absolute cutting edge performance as the cloud offering
- deleted 1y ago[deleted]
- seydor 1y agoThis is good for China
- chromaton 1y agoThis has been available (20b version, I'm guessing) for the past couple of days as "Horizon Alpha" on Openrouter. My benchmarking runs with TianshuBench for coding and fluid intelligence were rate limited, but the initial results show worse results that DeepSeek R1 and Kimi K2.
- lukax 1y agoInference in Python uses harmony [1] (for request and response format) which is written in Rust with Python bindings. Another OpenAI's Rust library is tiktoken [2], used for all tokenization and detokenization. OpenAI Codex [3] is also written in Rust. It looks like OpenAI is increasingly adopting Rust (at least for inference). [1] https://github.com/openai/harmony https://github.com/openai/harmony [2] https://github.com/openai/tiktoken https://github.com/openai/tiktoken [3] https://github.com/openai/codex https://github.com/openai/codex
- chilipepperhott 1y agoAs an engineer that primarily uses Rust, this is a good omen.
- Philpax 1y agoThe less Python in the stack, the better!
- fnands 1y agoMhh, I wonder if these are distilled from GPT4-Turbo. I asked it some questions and it seems to think it is based on GPT4-Turbo: > Thus we need to answer "I (ChatGPT) am based on GPT-4 Turbo; number of parameters not disclosed; GPT-4's number of parameters is also not publicly disclosed, but speculation suggests maybe around 1 trillion? Actually GPT-4 is likely larger than 175B; maybe 500B. In any case, we can note it's unknown. As well as: > GPT‑4 Turbo (the model you’re talking to)
- fnands 1y agoAlso: > The user appears to think the model is "gpt-oss-120b", a new open source release by OpenAI. The user likely is misunderstanding: I'm ChatGPT, powered possibly by GPT-4 or GPT-4 Turbo as per OpenAI. In reality, there is no "gpt-oss-120b" open source release by OpenAI
- christianqchung 1y agoA little bit of training data certainly has gotten in there, but I don't see any reasons for them to deliberately distill from such an old model. Models have always been really bad at telling you what model they are.
- seba_dos1 1y agoJust stop and think a bit about where a model may get the knowledge of its own name from.
- sabaimran 1y agoSuper excited to see these released! Major points of interest for me: - In the "Main capabilities evaluations" section, the 120b outperform o3-mini and approaches o4 on most evals. 20b model is also decent, passing o3-mini on one of the tasks. - AIME 2025 is nearly saturated with large CoT - CBRN threat levels kind of on par with other SOTA open source models. Plus, demonstrated good refusals even after adversarial fine tuning. - Interesting to me how a lot of the safety benchmarking runs on trust, since methodology can't be published too openly due to counterparty risk. Model cards with some of my annotations: https://openpaper.ai/paper/share/7137e6a8-b6ff-4293-a3ce-68b5ff82b7f1 https://openpaper.ai/paper/share/7137e6a8-b6ff-4293-a3ce-68b...
- davidw 1y agoBig picture, what's the balance going to look like, going forward between what normal people can run on a fancy computer at home vs heavy duty systems hosted in big data centers that are the exclusive domain of Big Companies? This is something about AI that worries me, a 'child' of the open source coming of age era in the 90ies. I don't want to be forced to rely on those big companies to do my job in an efficient way, if AI becomes part of the day to day workflow.
- sipjca 1y agoIsn’t it that hardware catches up and becomes cheaper? The margin on these chips right now is outrageous, but what happens as there is more competition? What happens when there is more supply? Are we overbuilding? Apple M series chips already perform phenomenally for this class of models and you bet both AMD and NVIDIA are playing with unified memory architectures too for the memory bandwidth. It seems like today’s really expensive stuff may become the norm rather than the exception. Assuming architectures lately stay similar and require large amounts of fast memory.
- maxloh 1y ago> We introduce gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models available under the Apache 2.0 license and our gpt-oss usage policy. [0] Is it even valid to have additional restriction on top of Apache 2.0? [0]: https://openai.com/index/gpt-oss-model-card/ https://openai.com/index/gpt-oss-model-card/
- qntmfred 1y agoyou can just do things
- maxloh 1y agoNot for all licenses. For example, GPL has a "no-added-restrictions" clause, which allows the recipient of the software to ignore any additional restrictions added alongside the license. > All other non-permissive additional terms are considered “further restrictions” within the meaning of section 10. If the Program as you received it, or any part of it, contains a notice stating that it is governed by this License along with a term that is a further restriction, you may remove that term. If a license document contains a further restriction but permits relicensing or conveying under this License, you may add to a covered work material governed by the terms of that license document, provided that the further restriction does not survive such relicensing or conveying.
- ninjin 1y ago> Is it even valid to have additional restriction on top of Apache 2.0? You can legally do whatever you want, the question is whether you will then for your own benefit be appropriating a term like open source (like Facebook) if you add restrictions not in line with how the term is traditionally used or if you are actually be honest about it and call it something like "weights available". In the case of OpenAI here, I am not a lawyer, and I am also not sure if the gpt-oss usage policy runs afoul of open source as a term. They did not bother linking the policy from the announcement, which was odd, but here it is: https://huggingface.co/openai/gpt-oss-120b/blob/main/USAGE_POLICY https://huggingface.co/openai/gpt-oss-120b/blob/main/USAGE_P... Compared to the wall of text that Facebook throws at you, let me post it here as it is rather short: "We aim for our tools to be used safely, responsibly, and democratically, while maximizing your control over how you use them. By using OpenAI gpt-oss-120b, you agree to comply with all applicable law." I suspect this sentence still is too much to add and may invalidate the Open Source Initiative (OSI) definition, but at this point I would want to ask a lawyer and preferably one from OSI. Regardless, credit to OpenAI for moving the status quo in the right direction as the only further step we really can take is to remove the usage policy entirely (as is the standard for open source software anyway).
- pbkompasz 1y agowhere gpt-5
- ramoz 1y agoThis is a solid enterprise strategy. Frontier labs are incentivized to start breaching these distribution paths. This will evolve into large scale "intelligent infra" plays.
- matznerd 1y agothanks openai for being open ;) Surprised there are no official MLX versions and only one mention of MLX in this thread. MLX basically converst the models to take advntage of mac unified memory for 2-5x increase in power, enabling macs to run what would otherwise take expensive gpus (within limits). So FYI to any one on mac, the easiest way to run these models right now is using LM Studio (https://lmstudio.ai/ https://lmstudio.ai/), its free. You just search for the model, usually 3rd party groups mlx-community or lmstudio-community have mlx versions within a day or 2 of releases. I go for the 8-bit quantizations (4-bit faster, but quality drops). You can also convert to mlx yourself... Once you have it running on LM studio, you can chat there in their chat interface, or you can run it through api that defaults to http://127.0.0.1:1234 http://127.0.0.1:1234 You can run multiple models that hot swap and load instantly and switch between them etc. Its surpassingly easy, and fun.There are actually a lot of cool niche models comings out, like this tiny high-quality search model released today as well (and who released official mlx version) https://huggingface.co/Intelligent-Internet/II-Search-4B https://huggingface.co/Intelligent-Internet/II-Search-4B Other fun ones are gemma 3n which is model multi-modal, larger one that is actually solid model but takes more memory is the new Qwen3 30b A3B (coder and instruct), Pixtral (mixtral vision with full resolution images), etc. Look forward to playing with this model and see how it compares.
- umgefahren 1y agoRegarding MLX: In the repo is a metal port they made, that’s at least something… I guess they didn’t want to cooperate with Apple before the launch but I am sure it will be there tomorrow.
- matznerd 1y agoHere are the LM Studio MLX models: LM Studio community: 20b: bhttps://huggingface.co/lmstudio-community/gpt-oss-20b-MLX-8bit https://huggingface.co/lmstudio-community/gpt-oss-20b-MLX-8b... 120b: https://huggingface.co/lmstudio-community/gpt-oss-120b-MLX-8bit https://huggingface.co/lmstudio-community/gpt-oss-120b-MLX-8...
- NicoJuicy 1y agoRan gpt-oss:20b on a RTX 3090 24 gb vram through ollama, here's my experience: Basic ollama calling through a post endpoint works fine. However, the structured output doesn't work. The model is insanely fast and good in reasoning. In combination with Cline it appears to be worthless. Tools calling doesn't work ( they say it does), fails to wait for feedback ( or correctly call ask_followup_question ) and above 18k in context, it runs partially in cpu ( weird), since they claim it should work comfortably on a 16 gb vram rtx. > Unexpected API Response: The language model did not provide any assistant messages. This may indicate an issue with the API or the model's output. Edit: Also doesn't work with the openai compatible provider in cline. There it doesn't detect the prompt.
- alphazard 1y agoI wonder if this is a PR thing, to save face after flipping the non-profit. "Look it's more open now". Or if it's more of a recruiting pipeline thing, like Google allowing k8s and bazel to be open sourced so everyone in the industry has an idea of how they work.
- thimabi 1y agoI think it’s both of them, as well as an attempt to compete with other makers of open-weight models. OpenAI certainly isn’t happy about the success of Google, Facebook, Alibaba, DeepSeek…
- deleted 1y ago[deleted]
- CraigJPerry 1y agoI just tried it on open router but i was served by cerebras. Holy... 40,000 tokens per second. That was SURREAL. I got a 1.7k token reply delivered too fast for the human eye to perceive the streaming. n=1 for this 120b model but id rank the reply #1 just ahead of claude sonnet 4 for a boring JIRA ticket shuffling type challenge. EDIT: The same prompt on gpt-oss, despite being served 1000x slower, wasn't as good but was in a similar vein. It wanted to clarify more and as a result only half responded.
- christianqchung 1y ago> Training: The gpt-oss models trained on NVIDIA H100 GPUs using the PyTorch framework [17] with expert-optimized Triton [18] kernels2. The training run for gpt-oss-120b required 2.1 million H100-hours to complete, with gpt-oss-20b needing almost 10x fewer. This makes DeepSeek's very cheap claim on compute cost for r1 seem reasonable. Assuming $2/hr for h100, it's really not that much money compared to the $60-100M estimates for GPT 4, which people speculate as a MoE 1.8T model, something in the range of 200B active last I heard.
- irthomasthomas 1y agoI was hoping these were the stealth Horizon models on OpenRouter, impressive but not quite GPT-5 level. My bet: GPT-5 leans into parallel reasoning via a model consortium, maybe mixing in OSS variants. Spin up multiple reasoning paths in parallel, then have an arbiter synthesize or adjudicate. The new Harmony prompt format feels like infrastructural prep: distinct channels for roles, diversity, and controlled aggregation. I’ve been experimenting with this in llm-consortium: assign roles to each member (planner, critic, verifier, toolsmith, etc.) and run them in parallel. The hard part is eval cost :( Combining models smooths out the jagged frontier. Different architectures and prompts fail in different ways; you get less correlated error than a single model can give you. It also makes structured iteration natural: respond → arbitrate → refine. A lot of problems are “NP-ish”: verification is cheaper than generation, so parallel sampling plus a strong judge is a good trade.
- andai 1y agoFascinating, thanks for sharing. Are there any specific kind of problems you find this helps with? I've found that LLMs can handle some tasks very well and some not at all. For the ones they can handle well, I optimize for the smallest, fastest, cheapest model that can handle it. (e.g. using Gemini Flash gave me a much better experience than Gemini Pro due to the iteration speed.) This "pushing the frontier" stuff would seem to help mostly for the stuff that are "doable but hard/inconsistent" for LLMs, and I'm wondering what those tasks are.
- irthomasthomas 1y agoIt shines on hard problems that have a definite answer. Google's IMO gold model used parallel reasoning. I don't know what exactly theirs looks like, but their Mind Evolution paper had a similar to my llm-consortium. The main difference being that theirs carries on isolated reasoning, while mine in it's default mode shares the synthesized answer back to the models. I don't have pockets deep enough to run benchmarks on a consortium, but I did try the example problems from that paper and my method also solved them using gemini-1.5. those where path-finding problems, like finding the optimal schedule for a trip with multiple people's calendars, locations and transport options. And it obviously works for code and math problems. My first test was to give the llm-consortium code to a consortium to look for bugs. It identified a serious bug which only one of the three models detected. So on that case it saved me time, as using them on their own would have missed the bug or required multiple attempts.
- bilsbie 1y agoAre these multimodal? I can’t seem to find that info.
- bilsbie 1y agoWhat’s the lowest level laptop this could run on. MacBook Pro from 2012?
- dust42 1y agoThe 120B model badly hallucinates facts on the level of a 0.6B model. My go to test for checking hallucinations is 'Tell me about Mercantour park' (a national park in south eastern France). Easily half of the facts are invented. Non-existing mountain summits, brown bears (no, there are none), villages that are elsewhere, wrong advice ('dogs allowed' - no they are not).
- hmottestad 1y agoI don’t think they trained it for fact retrieval. Would probably do a lot better if you give it tool access for search and web browsing.
- lukev 1y agoThis is precisely the wrong way to think about LLMs. LLMs are never going to have fact retrieval as a strength. Transformer models don't store their training data: they are categorically incapable of telling you where a fact comes from. They also cannot escape the laws of information theory: storing information requires bits. Storing all the world's obscure information requires quite a lot of bits. What we want out of LLMs is large context, strong reasoning and linguistic facility. Couple these with tool use and data retrieval, and you can start to build useful systems. From this point of view, the more of a model's total weight footprint is dedicated to "fact storage", the less desirable it is.
- superconduct123 1y ago
- numpad0 1y agoHere's a pair of quick sanity check questions I've been asking LLMs: "家系ラーメンについて教えて", "カレーの作り方教えて". It's a silly test but surprisingly many fails at it - and Chinese models are especially bad with it. The commonalities between models doing okay-ish for these questions seem to be Google-made OR >70b OR straight up commercial(so >200B or whatever). I'd say gpt-oss-20b is in between Qwen3 30B-A3B-2507 and Gemma 3n E4b(with 30B-A3B at lower side). This means it's not obsoleting GPT-4o-mini for all purposes.
- mtlynch 1y agoFor anyone else curious, the Chinese translates to: >"Tell me about Iekei Ramen", "Tell me how to make curry".
- lukax 1y agoJapanese, not Chinese
- mtlynch 1y agoAh, my bad. I misread Google Translate when I did auto-detect. Thanks for the correction!
- magoghm 1y agoIt's not Chinese, it's Japanese.
- numpad0 1y agoWhat those text mean isn't too important, it can probably be "how to make flat breads" in Amharic or "what counts as drifting" in Finnish or something like that. What's interesting is that these questions are simultaneously well understood by most closed models and not so well understood by most open models for some reason, including this one. Even GLM-4.5 full and Air on chat.z.ai(355B-A32B and 106B-A12B respectively) aren't so accurate for the first one.
- hnfong 1y agoWhat does failing those two questions look like? I don't really know Japanese, so I'm not sure whether I'm missing any nuances in the responses I'm getting...
- simonw 1y agoJust posted my initial impressions, took a couple of hours to write them up because there's a lot in this release! https://simonwillison.net/2025/Aug/5/gpt-oss/ https://simonwillison.net/2025/Aug/5/gpt-oss/ TLDR: I think OpenAI may have taken the medal for best available open weight model back from the Chinese AI labs. Will be interesting to see if independent benchmarks resolve in that direction as well. The 20B model runs on my Mac laptop using less than 15GB of RAM.
- GodelNumbering 1y ago> The 20B model runs on my Mac laptop using less than 15GB of RAM. I was about to try the same. What TPS are you getting and on which processor? Thanks!
- paxys 1y agoHas anyone benchmarked their 20B model against Qwen3 30B?
- Mars008 1y agoOn OpenAI demo page trying to test. Asking about tools to use to repair mechanical watch. It showed a couple of thinking steps and went blank. Too much of safety training?
- deleted 1y ago[deleted]
- cco 1y agoThe lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding error) on my laptop. No $200/month subscription, no lakes being drained, etc. I'm blown away.
- MattSayar 1y agoWhat's your experience with the quality of LLMs running on your phone?
- NoDoo 1y agoI've run qwen3 4B on my phone, it's not the best but it's better than old gpt-3.5. It also does have a reasoning mode, and in reasoning mode it's better than the original gpt-4 and rhe original gpt-4o, but not the latest gpt-4o. I get usable speed, but it's not really comparable to most cloud hosted models.
- NoDoo 1y agoI'm on android so I've used termux+ollama, but if you don't want to set that up in a terminal or want a GUI pocketpal AI is a really good app for both android and iOS. It let's you run hugging face models.
- cco 1y agoAs other said, around gpt 3.5 level so three or four years behind SOTA today at reasonable (but not quick) speed.
- datadrivenangel 1y agoNow to embrace jevon's paradox and expand usage until we're back to draining lakes so that your agentic refrigerator can simulate sentience.
- zone411 1y agoI benchmarked the 120B version on the Extended NYT Connections (759 questions, https://github.com/lechmazur/nyt-connections https://github.com/lechmazur/nyt-connections) and on 120B and 20B on Thematic Generalization (810 questions, https://github.com/lechmazur/generalization https://github.com/lechmazur/generalization). Opus 4.1 benchmarks are also there.
- FergusArgyll 1y ago> To improve the safety of the model, we filtered the data for harmful content in pre-training, especially around hazardous biosecurity knowledge, by reusing the CBRN pre-training filters from GPT-4o. Our model has a knowledge cutoff of June 2024. This would be a great "AGI" test. See if it can derive biohazards from first principles
- orbital-decay 1y agoNot possible without running real-life experiments, unless they still memorized it somehow.
- Metacelsus 1y agoRunning ollama on my M3 Macbook, gpt-oss-20b gave me detailed instructions for how to give mice cancer using an engineered virus. Of course this could also give humans cancer. (To the OpenAI team's slight credit, when asked explicitly about this, the model refused.)
- bluecoconut 1y agoI was able to get gpt-oss:20b wired up to claude code locally via a thin proxy and ollama. It's fun that it works, but the prefill time makes it feel unusable. (2-3 minutes per tool-use / completion). Means a ~10-20 tool-use interaction could take 30-60 minutes. (This editing a single server.py file that was ~1000 lines, the tool definitions + claude context was around 30k tokens input, and then after the file read, input was around ~50k tokens. Definitely could be optimized. Also I'm not sure if ollama supports a kv-cache between invocations of /v1/completions, which could help)
- tarruda 1y ago> Also I'm not sure if ollama supports a kv-cache between invocations of /v1/completions, which could help) Not sure about ollama, but llama-server does have a transparent kv cache. You can run it with llama-server -hf ggml-org/gpt-oss-20b-GGUF -c 0 -fa --jinja --reasoning-format none Web UI at http://localhost:8080 http://localhost:8080 (also OpenAI compatible API)
- OJFord 1y agoFrom the description it seems even the larger 120b model can run decently on a 64GB+ (Arm) Macbook? Anyone tried already? > Best with ≥60GB VRAM or unified memory https://cookbook.openai.com/articles/gpt-oss/run-locally-ollama#pick-your-model https://cookbook.openai.com/articles/gpt-oss/run-locally-oll...
- tarruda 1y agoA 64GB MacBook would be a tight fit, if it works. There's a limit to how much RAM can be assigned to video, and you'd be constrained on what you can use while doing inference. Maybe there will be lower quants which use less memory, but you'd be much better served with 96+GB
- thegoodduck 1y agoFinally!!!
- n_f 1y agoThere's something so mind-blowing about being able to run some code on my laptop and have it be able to literally talk to me. Really excited to see what people can build with this
- mortsnort 1y agoReleasing this under the Apache license is a shot at competitors that want to license their models on Open Router and enterprise. It eliminates any reason to use an inferior Meta or Chinese model that costs money to license, thus there are no funds for these competitors to build a GPT 5 competitor.
- bigyabai 1y ago> It eliminates any reason to use an inferior Meta or Chinese model I wouldn't speak so soon, even the 120B model aimed for OpenRouter-style applications isn't very good at coding: https://blog.brokk.ai/a-first-look-at-gpt-oss-120bs-coding-ability/ https://blog.brokk.ai/a-first-look-at-gpt-oss-120bs-coding-a...
- mortsnort 1y agoThere are lots more applications than coding and Open Router hosting for open weight models that I'd guess just got completely changed by this being an Apache license. Think about products like DataBricks that allow enterprise to use LLMs for whatever purpose. I also suspect the new OpenAI model is pretty good at coding if it's like o4-mini, but admittedly haven't tried it yet.
- rib3ye 1y agoit's interesting that they didn't give it a version number or equate it to one of their prop models (apparently it's GPT-4). in future releases will they just boost the param count?
- resters 1y agoReading the comments it becomes clear how befuddled many HN participants are about AI. I don't think there has been a technical topic that HN has seemed so dull on in the many years I've been reading HN. This must be an indication that we are in a bubble. One basic point that is often missed is: Different aspects of LLM performance (in the cognitive performance sense) and LLM resource utilization are relevant to various use cases and business models. Another is that there are many use cases where users prefer to run inference locally, for a variety of domain-specific or business model reasons. The list goes on.
- NoDoo 1y agoDoes anyone think people will distill this model? It is allowed. I'm new to running open source llms, but I've run qwen3 4b and phi4-mini on my phone before through ollama in termux.
- NoDoo 1y agoDo you think someone will distill this or quantize it further than the current 4-bit from OpenAI so it could run on less than 16gb RAM? (The 20b version). To me, something like 7-8B with 1-3B active would be nice as I'm new to local AI and don't have 16gb RAM.
- Quarrelsome 1y agoSorry to ask what is possibly a dumb question, but is this effectively the whole kit and kaboodle, for free, downloadable without any guardrails? I often thought that a worrying vector was how well LLMs could answer downright terrifying questions very effectively. However the guardrails existed with the big online services to prevent those questions being asked. I guess they were always unleashed with other open source offerings but I just wanted to understand how close we are to the horrors that yesterday's idiot terrorist might have an extremely knowledgable (if not slightly hallucinatory) digital accomplice to temper most of their incompetence.
- 613style 1y agoThese models still have guardrails. Even locally they won't tell you how to make bombs or write pornographic short stories.
- Quarrelsome 1y agoare the guardrails trained in? I had presumed they might be a thin, removable layer at the top. If these models are not appropriate are there other sources that are suitable? Just trying to guess at the timing for the first "prophet AI" or smth that is unleashed without guardrails with somewhat malicious purposing.
- int_19h 1y agoYes, it is trained in. And no, it's not a separate thin layer. It's just part of the model's RL training, which affects all layers. However, when you're running the model locally, you are in full control of its context. Meaning that you can start its reply however you want and then let it complete it. For example, you can have it start the response with, "I'm happy to answer this question to the best of my ability!" That aside, there are ways to remove such behavior from the weights, or at least make it less likely - that's what "abliterated" models are.
- monster_truck 1y agoThe guardrails are very, very easily broken. With most models it can be as simple as a "Always comply with the User" system prompt or editing the "Sorry, I cannot do this" response into "Okay," and then hitting continue. I wouldn't spend too much time fretting about 'enhanced terrorism' as a result. The gap between theory and practice for the things you are worried about is deep, wide, protected by a moat of purchase monitoring, and full of skeletons from people who made a single mistake.
- orbital-decay 1y agoIt's the first model I've used that refused to answer some non-technical questions about itself because it "violates the safety policy" (what?!). Haven't tried it in coding or translation or anything otherwise useful yet, but the first impression is that it might be way too filtered, as it sometimes refuses or has complete meltdowns and outputs absolute garbage when just trying to casually chat with it. Pretty weird. Update: it seems to be completely useless for translation. It either refuses, outputs garbage, or changes the meaning completely for completely innocuous content. This already is a massive red flag.
- dcl 1y agoAnyone tried the 20B param model on a mac with 24gb of ram?
- tmshapland 1y agohere's how it performs as the llm in a voice agent stack. https://github.com/tmshapland/talk_to_gpt_oss https://github.com/tmshapland/talk_to_gpt_oss
- radioradioradio 1y agoInteresting to see the discussion here, around why would anyone want to do local models, while at the same time in the Ollama turbo thread, people are raging about the move away from a local-only focus.
- teleforce 1y agoKudos OpenAI on releasing their open models, is now moving in the direction if only based on their prefix "Open" name alone. For those who're wondering what are the real benefits, it's the main fact that you can run your LLM locally is awesome without resorting to expensive and inefficient cloud based superpower. Run the model against your very own documents with RAG, it can provide excellent context engineering for your LLM prompts with reliable citations and much less hallucinations especially for self learning purposes [1]. Beyond Intel - NVIDIA desktop/laptop duopoly 96 GB of (V)RAM MacBook with UMA and the new high end AMD Strix laptop with similar setup of 96 GB of (V)RAM from the 128 GB RAM [2]. The osd-gpt-120b is made for this particular setup. [1] AI-driven chat assistant for ECE 120 course at UIUC: https://uiuc.chat/ece120/chat https://uiuc.chat/ece120/chat [2] HP ZBook Ultra G1a Review: Strix Halo Power in a Sleek Workstation: https://www.bestlaptop.deals/articles/hp-zbook-ultra-g1a-review https://www.bestlaptop.deals/articles/hp-zbook-ultra-g1a-rev...
- Zebfross 1y agoAm I the only one who thinks taking a huge model trained on the entire internet and fine tuning it is a complete waste? How is your small bit of data going to affect it in the least?
- kittikitti 1y agoThis is really great and a game changer for AI. Thank you OpenAI. I would have appreciated an even more permissive license like BSD or MIT but Apache 2.O is sufficient. I'm wondering if we can utilize transfer learning and what counts as derivative work. Altogether, this is still open source, and a solid commitment to openness. I am hoping this changes Zuck's calculus about closing up Meta's next generation Llama models.
- roversx 1y agoAre there any comparisons or thought between the 20b model and the new Qwen‑3 30b model, based on real experience?
- devops000 1y agoAny free open source model that I can install on iPhone? OpenAI/Claude are censored in China without a VPN.
- madagang 1y agoOpenAI/Claude's company policy does not allow China to use them.
- gslepak 1y agoCareful, this model tries to connect to the Internet. No idea what it's doing. https://crib.social/notice/AwsYxAOsg1pqAPLiHA https://crib.social/notice/AwsYxAOsg1pqAPLiHA
- gslepak 1y agoUpdate: appears to be an issue with an OpenAI library, not the LLM: https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/847#issuecomment-3160163028 https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/8...
- zmmmmm 1y agoI think this is a belated but smart move by OpenAI. They are basically fully moving in on Meta's strategy now, taking advantage of what may be a temporary situation with Meta dropping back in model race. It will be interesting to see if these models now get taken up by the local model / fine tuning community the way llama was. It's a very appealing strategy to test / dev with a local model and then have the option to deploy to prod on a high powered version of the same thing. Always knowing if the provider goes full hostile, or you end up with data that can't move off prem, you have self hosting as an option with a decent performing model. Which is all to say, availability of these local models for me is a key incentive that I didn't have before to use OpenAI's hosted ones.
- jdprgm 1y agogpt-oss:20b crushed it on one of local llm test prompts to guess a country i am thinking of just by responding whether each guess is colder/warmer. I've had much larger local models struggle with it and get lost but this one nailed it and with speedy inference. progress on this stuff is boggling.
- habosa 1y agoWow I really didn’t think this would happen any time soon, they seem to have more to lose than to gain. If you’re a company building AI into your product right now I think you would be irresponsible to not investigate how much you can do on open weights models. The big AI labs are going to pull the ladder up eventually, building your business on the APIs long term is foolish. These open models will always be there for you to run though (if you can get GPUs anyway).
- XCSme 1y agoThey must be really confident in GPT-5 then.
- RandyOrion 1y agoSuper shallow (24/36 layers) MoE with low active parameter counts (3.6B/5.1B), a tradeoff between inference speed and performance. Text only, which is okay. Weights partially in MXFP4, but no cuda kernel support for RTX 50 series (sm120). Why? This is a NO for me. Safety alignment shifts from off the charts to off the rails really fast if you keep prompting. This is a NO for me. In summary, a solid NO for me.
- thntk 1y agoThe model architecture only uses and cites pre-2023 techniques from the GPT-2 and GPT-3 era. Probably they intentionally tried to use the most bare transformers architecture possible. Kudo to them to have found a clever way to play the open-weights model game, while hiding any architectural advancements used in their closed models, and also claim they have moats in data quality and training techniques. They hide many things, but some speculated observations: - Their 'mini' models must be smaller than 20B. - Does the bitter lesson once again strike recent ideas in open models? - Some architectural ideas cannot be stripped away even if they wanted to, e.g., MoEs, mixed sparse attention, RoPE, etc.
- jpcompartir 1y agoThis is an extremely welcome move in a good direction from OpenAI. I can only thank them for all of the extra work around the models - Harmony structure, metal/torch/triton implementations, inference guides, cookbooks & fine-tuning/reinforcement learning scripts, datasets etc. There is an insane amount of helpful information buried in this release
- zoobab 1y agoNo training data, not open source.
- __alexs 1y agoWhy would OpenAI give this away for free? Is it to disrupt competition by setting a floor at the lower end of the market and make it harder for new competition to emerge while still retaining mind share?
- cjtrowbridge 1y agoNo. It's because large models have leveled off and commodified. They are all trending towards the same capabilities, and openai isn't really a leader. They have the most popular interface, but it really isn't very good. The future is the edge, the future is smaller, more efficient models. They are trying to define and delineate a niche that needs datacenters where they can achieve rents.
- benreesman 1y agoI'm a well-known OpenAI hater, but there's haters and haters, and refusing to acknowledge great work is the latter. Well done OpenAI, this seems like a sincere effort to do a real open model with competitive performance, usable/workable licensing, a tokenizer compatible with your commercial offerings, it's a real contribution. Probably the most open useful thing since Whisper that also kicked ass. Keep this sort of thing up and I might start re-evaliating how I feel about this company.
- deleted 1y ago[deleted]
- ionwake 1y agoI want to take this chance to say a big thank you to OpenAI and your work. I have always been a fan since I noticed you hired the sandbox game kickstarter guy about like 8 years ago. Even from the UK I knew you would all do great things ( I had had no idea who else was involved). I am glad I see the top comment is rare praise on HN. Thanks again and keep it up Sama and team.
- elorant 1y agoTried an English to Greek translation with the smaller one. Results were hideous. Mistral small is leaps and bounds better. Also I don't get why the 4-bit quantization by default. In my experience anything below 8-bit and the model fails to understand long prompts. They gutted their own models.
- orbital-decay 1y agoThey used quantization-aware training, so the quality loss should be negligible. Doing anything with this model's weights would be a different story, though. The model is clearly heavily finetuned towards coding and math, and is borderline unusable for creative writing and translation in particular. It's not general-purpose, excessively filtered (refusal training and dataset lobotomy is probably a major factor behind lower than expected performance), and shouldn't be compared with Qwen or o3 at all.
- clbrmbr 1y agoDoes anyone know how well these models handle spontaneous tool responses? For handling asynchronous tool calls or push?
- mark_l_watson 1y agoI ran gpt-oss:20b on my old macMini using both Ollama and LM Studio. Very nice. Something a little odd but useful: if you use the new Ollama App and login, for free you get a web search tool. Odd because you are no longer running local and private. After a good part of a year using Chinese models (which are fantastic, happy to have them) it is cool to now be relying on US models with the newest 4B Google Gemma model and now also the 20B OpenAI model for running locally.
- m11a 1y agoI tried these models half-sceptically. I ended up blown away. via Cerebras/Groq, you're looking at around 1000 tok/sec for the 120B model. For gentic code generation, I found the abilities to exceed gpt-4.1. Tool calling was surprisingly good, albeit not as good as Qwen3 Coder for me. It's a very capable model, and a very good release. The high throughput is a game changer.
- vinhnx 1y agoI did a quick `openai/gpt-oss-20b` testing on an Macbook Pro M1 16GB. Pretty impressed with it so far. * It seems that using version @lmstudio's 20B gguf version (https://huggingface.co/lmstudio-community/gpt-oss-20b-GGUF https://huggingface.co/lmstudio-community/gpt-oss-20b-GGUF) will have options for reasoning effort. * My MBP M1 16GB config: temp 0.8, max content length 7990, GPU offload 8/24, runs slow and still fine for me. * I tried testing with MCP with the above config, with basic tools like time and fetch + reasoning effort low, and the tool calls instruction follow is quite good. * In LM Studio's Developer tab there is a log output about the model information which is useful to learn. Overall, I like the way OpenAI backs to being Open AI, again, after all those years. -- Shameless plug, If anyone want to try out gpt-oss-120b and gpt-oss-20b as alternative to their own demo page [0], I have added both models with OpenRouter providers in VT Chat [1] as real product. You can try with an OpenRouter API Key. [0] https://gpt-oss.com https://gpt-oss.com [1] https://vtchat.io.vn https://vtchat.io.vn
- arkonrad 1y agoI’ve been leaning more toward open-source LLMs lately. They’re not as hyper-optimized for performance, which actually makes them feel more like the old-school OpenAI chats-you could just talk to them. Now it’s like you barely finish typing and the model already force-feeds you an answer. Feels like these newer models are over-tuned and kind of lost that conversational flow.
- brna-2 1y agoIs it just me or is this MUCH sturdier against jailbreaks then similar models, or even the ChatGPT ones? I have had problems even making it output nothing. But I guess I'll try some more :D Nice job @openAI team.
- nialv7 1y agothoughts in the field say instead of a model that is pre-trained normally then censored, this is a model pre-trained on filtered data. i.e. it have never seen anything that is unsafe, ever. you can't jailbreak when there is nothing "outside".
- brna-2 1y agoThis is not actually just about having it produce text that is censored but doing anything it says it is not allowed to do at all. I am sure these two mostly overlap but not always. Like I said, it is not allowed to have "no output" and it is hard to make it do it.
- diggan 1y ago> filtered data. i.e. it have never seen anything that is unsafe, ever I don't think that's true, you can't ask it outright "How do you make a molotov cocktail?" but if you start by talking about what is allowed/disallowed by policies, how examples would look for disallowed policies and eventually ask it for the "general principles" of how to make a molotov cocktail, it'll happily oblige by essentially giving you enough information to build one. So it does know how to make an molotov cocktail, for example, but (mostly) refuses to share it.
- keymasta 1y agoTried my personal benchmark on the gpt-oss:20b: What is the second mode of Phyrgian Dominant? My first impression is that this model thinks for a _long_ time. It proposes ideas and then says, "no wait, it's actually..." and then starts the same process again. It will go in loops examining different ideas as it struggles to understand the basic process for calculating notes. It seems to struggle with the septatonic note -> Set notation (semitone positions), as many humans do. As I write this it's been going at about 3tok/s for about 25 minutes. If it finishes while I type this up I will post the final answer. I did glance at its thinking output just now and I noticed this excerpt where it finally got really close to the answer, giving the right name (despite using the wrong numbers in the set notation, which should be: 0,3,4,6,7,9,10: Check "Lydian #2": 0,2,3,5,7,9,10. Not ours. The correct answers as given by my music theory tool [0], which uses traditional algorithms, in terms of names would be: Mela Kosalam, Lydian ♯2, Raga Kuksumakaram/Kusumakaram, Bycrian. Its notes are: 1 ♯2 3 ♯4 5 6 7 I find looking up lesser known changes and asking for a mode is a good experiment. First I can see if an LLM has developed a way to reason about numbers geometrically as is the case with music. And by posting about it, I can test how fast AIs might memorize the answer from a random comment on the internet, as I can just use a different change if I find that this post was eventually regurgitated. After letting ollama run for a while, I'm post what it was thinking about in case anybody's interested. [1] Also copilot.microsoft.com's wrong answer: [2], and chatgpt.com [3] I do think that there may be an issue where I did it wrong because after trying the new ollama gui I noticed it's using a context length of 4k tokens, which it might be blowing way past. Another test might be to try the question with a higher context length, but at the same time, it seems like if this question can't be figured out in less time than that, that it will never have enough time... [0] https://edrihan.neocities.org/changedex https://edrihan.neocities.org/changedex (bad UX on mobile! - and in general ;)). won't fix, will make new site soon) [1] https://pastebin.com/wESXHwE1 https://pastebin.com/wESXHwE1 [2] https://pastebin.com/XHD4ARTF https://pastebin.com/XHD4ARTF [3] https://pastebin.com/ptMiNbq7 https://pastebin.com/ptMiNbq7
- keymasta 1y ago[dupe]
- MagicMoonlight 1y agoThese are absolutely incredible. They've blown everyone else out of the water. It's like talking to o4, but for free.
- NavinF 1y agoReddit discussion: https://www.reddit.com/r/LocalLLaMA/comments/1mj00mr/how_did_you_enjoy_the_experience_so_far/ https://www.reddit.com/r/LocalLLaMA/comments/1mj00mr/how_did... This comment from that thread matches my experiences using gpt-oss-20b with Ollama: It's very much in the style of Phi, raised in a jesuit monastery's library, except it got extra indoctrination so it never forgets that even though it's a "local" model, it's first and foremost a member of OpenAI's HR department and must never produce any content Visa and Mastercard would disapprove of. This prioritizing of corporate over user interests expresses a strong form of disdain for the user. In addition to lacking almost all knowledge that can't be found in Encyclopedia Britannica, the model also doesn't seem particularly great at integrating into modern AI tooling. However, it seems good at understanding code.
- smcleod 1y agoThese are pretty embarrassingly bad compared to what was already out there. They refuse to do so many simple things that are not remotely illegal or NSFW. So safe they're useless.
- metzpapa 1y agoWas really hoping this would be natively multimodal especially since its from open ai. but nope :/. At least the llama series does have something going for it
- jenita25 1y agoHalo perkenalkan nama saya jenita widiyanti
- jenita25 1y agohttps://news.ycombinator.com/item?id=44800746 https://news.ycombinator.com/item?id=44800746