6 ms·
I don’t get the hype with OpenAI OSS, they would never make a model better than their proprietary models open source, and the other open source models beat GPT
by 44za12 1y ago
I don’t get the hype with OpenAI OSS, they would never make a model better than their proprietary models open source, and the other open source models beat GPT and family so why the wait?
- jphoward 1y agoI think they could release non-agentic models that are as good as 4o, and have almost no repercussions on sales tbh. I have Ollama installed (only a small proportion of their clients would have a large enough GPU for this) and have download deepseek and played with it, but I still pay for an OpenAI subscription because I want the speed of a hosted model, and never mind the luxuries of things like Codex's diffs/pull request support, agents on new models, deep research etc. - I use them all at least weekly.
- 44za12 1y agoI pay for Cursor, OpenAI and kimi (to use with Claude Code), OpenAI is good with quickly refining my thoughts, Cursor’s subscription I’m reconsidering to cancel bought it for Claude but the rate limits are making it impossible for me to find it useful. Kimi is what truly surprises me, Claude code shows this conversation costed you $500 (based on Opus usage which is mapped to kimi k2) while I’ve barely spent $2. I have Ollama as well, majorly to quickly test small models that could be improved for our usecase through finetuning.
- garciasn 1y agoWhat am I doing wrong that I'm never hitting the rate limits on the $100 Max plan?
- Topfi 1y agoConsidering my personal heavy use also not leading to rate limits and what I've seen by some users over the past months, I suspect a mix of actually thinking about your code before writing a prompt, managing context by documenting and running stuff like git, npm install, etc. yourself instead of "Hey Claude, setup React with Radix and install a few packages". I have genuinely seen someone use ultrathink for setting up a starter repo hosted on Github, despite the commands being listed in the readme, so I can see how certain people may hit the limits quicker than others. Still, I will cancel my Claude Max subscription if they remain intransparent concerning the amount of use we actually get, especially regarding the mail they sent out recently which stated that 20x Max users do only get 10x in terms of expected usable hours. Same goes for still not providing an official way to track how much use one has left in a week.
- garciasn 1y ago> running stuff like git, npm install, etc. yourself Ah; this definitely makes sense! I do this myself and then paste back only the relevant part of the log so as to limit this. I suspect I am being more conservative than others.
- 44za12 1y agoI am on the pro plan, I was considering Max, but then i found kimi and I’m getting used to it.
- ewoodrich 1y agoI've been using Kimi with Roo via OpenRouter and have been very surprised at how capable it is. It's the first open model I've tried that actually lives up to claims I see online that's it on par with this or that previous gen proprietary model. Context window has been the only negative, at least with the providers OpenRouter has been giving me but forgiveable given how absurdly cheap it is.
- nico 1y ago> but forgiveable given how absurdly cheap it is Are you using it everyday for programming? If so, how much more or less does it cost you per month? More or less than $100?
- vineyardmike 1y agoThey would definitely have sales repercussions, but it might be worth it. They are fully trying to be a consumer product, developer services be damned. But they can’t just get rid of the API because it’s a good incremental source of revenue, and thanks to the Microsoft deal, all that revenue would end up in Azure. Maintaining their API is basically just a way to get a slice of that revenue. But if they open sourced everything, it might sour the relationship more with Microsoft, who would lose azure revenue and might be willing to part ways. It would also ensure that they compete on consumer product quality not (directly) model quality. At this point, they could basically put any decent model in their app and maintain the user base, they don’t actually need to develop their own.
- Topfi 1y agoPure performance isn't necessarily everything. Context window, speed and local use are just some of the upsides this model may have. We still know next to nothing so anything is possible, but if it is an MoE at 120B, that could enable some interesting local use cases, even if it less capable than e.g. Deepseek V3, simply by running on more hardware/at higher tokens/sec. GPT-4.1s code focus has also shown that OpenAI does have a knack for models with a more narrow use case, maybe this will do well in specific tasks. More so since GPT-4.1 was that much better than the massive GPT-4.5, I am cautiously optimistic. Even if it does poorly in all areas (like Llama 4 [0]), there is still a lot the community and industry can learn from even an uncompetitive model. [0] Llama 4 technically has a massive 10M token context as a differentiator, however in my experience, it is not reliably usable beyond 100k.
- jstummbillig 1y agoI don't see how it would be in OpenAIs selfish interest to release an open source model that sucks. Unless you can cohesively explain how that would work in their favor, it seems a lot smarter to assume that they won't.
- granitepail 1y agoWhile the benchmarks all say open source models Kimi and Qwen outpace proprietary models like GPT 4.1, GPT 4o, or even o3, my (and just about everyone I know's) boots on the ground experience suggests they're not even close. This is for tool calling agentic tasks, like coding, but also in other contexts (research, glue between services, etc). I feel like it's worth putting that out there--it's pretty clear there's a lot of benchmark hacking happening. I'm not really convinced it's purposeful/deceitful, but it's definitely happening. Qwen3 Coder, for example, is basically incompetent for any real coding tasks and frequently gets caught in death spirals of bad tool calls. I try all the OSS models regularly, because I'm really excited for them to get better. Right now Kimi K2 is the most usable one, and I'd rate it at a few ticks worse than GPT 4.1.
- jimbo808 1y agoI would have assumed anyone frequenting HN would have figured out by now that benchmarks are 100% bullshit. I guess I'd be wroing.
- dist-epoch 1y agoSo what do you propose? Gut feel, N=1 tests?
- deleted 1y ago[deleted]
- spullara 1y agoit currently beats depending on the benchmarks
- BoorishBears 1y agoI mean, in other environments people say that. If you asked "What's the best bicycle", most enthusiasts would say one you tried, works for your usecase, etc. Benchmarks should be for pruning models you try at the absolute highest level, because at the end of the day it's way too easy to hack them without breaking any rules (post-train on the public, generate a ton of synthetic examples, train on those, repeat)
- rdtsc 1y ago> they would never make a model better than their proprietary models open source Not their proprietary model, but maybe other open source models, or closed source models of their competitors. That way they can first ensure they are the only player on both sides, and then can kneecap their open source models just enough to drive the revenue to their proprietary one.
- 44za12 1y agoMaking a model better than proprietary models is in fact making a model better than their closed source models if you believe the benchmarks.
- rdtsc 1y agoRealistically yeah but if they drunk their own kool-aid, they'd think they have the best proprietary and open source model. So capturing both sides of the ecosystem would make sense.
- PeterStuer 1y agoThey might not release a better model than their proprietary models, but others can build on and tinker with these open models to improve and specialize them. Another reason people are 'hyped' for open models is that access to them can not be taken away or price gauged at the whim of the provider, and that their use can not be restricted in arbitrary ways, although I'm sure that on the latter part they will have a go at it through regulation. Grab'em while you can.