5 ms·
While the benchmarks all say open source models Kimi and Qwen outpace proprietary models like GPT 4.1, GPT 4o, or even o3, my (and just about everyone I know's)
by granitepail 1y ago
While the benchmarks all say open source models Kimi and Qwen outpace proprietary models like GPT 4.1, GPT 4o, or even o3, my (and just about everyone I know's) boots on the ground experience suggests they're not even close. This is for tool calling agentic tasks, like coding, but also in other contexts (research, glue between services, etc). I feel like it's worth putting that out there--it's pretty clear there's a lot of benchmark hacking happening. I'm not really convinced it's purposeful/deceitful, but it's definitely happening. Qwen3 Coder, for example, is basically incompetent for any real coding tasks and frequently gets caught in death spirals of bad tool calls. I try all the OSS models regularly, because I'm really excited for them to get better. Right now Kimi K2 is the most usable one, and I'd rate it at a few ticks worse than GPT 4.1.
- jimbo808 1y agoI would have assumed anyone frequenting HN would have figured out by now that benchmarks are 100% bullshit. I guess I'd be wroing.
- dist-epoch 1y agoSo what do you propose? Gut feel, N=1 tests?
- deleted 1y ago[deleted]
- spullara 1y agoit currently beats depending on the benchmarks
- BoorishBears 1y agoI mean, in other environments people say that. If you asked "What's the best bicycle", most enthusiasts would say one you tried, works for your usecase, etc. Benchmarks should be for pruning models you try at the absolute highest level, because at the end of the day it's way too easy to hack them without breaking any rules (post-train on the public, generate a ton of synthetic examples, train on those, repeat)
- int_19h 1y agoAt the moment, the only way you can tell if the model is good for a particular task is by trying it at that task. Gut feel is how you pick the models to test first, and that is also based largely on past experience and educated guesses as to what strengths translate between tasks. You should also remember that there's no free lunch. If you see models below a certain size fail consistently, don't expect a model that is even smaller to somehow magically succeed, no matter how much pixie dust the developer advertises.
- sebzim4500 1y agoTo some extent there must be a free lunch, because today's 30B models are enormously better than the 30B models that existed a year ago. I suppose it's an open question whether there is another free lunch or whether the 30B models in a year will be not much better than our current ones.
- andrewmcwatters 1y agoI think anyone frequenting HN and actually using these tools absolutely knows these benchmarks are 100% bullshit and the only real way to test these things is to just use them yourself. Many small models are supposedly good for controlled tasks, but given a detailed prompt, I can't get any of them to follow simple instructions. They usually just regurgitate the examples in the system prompt. Useless.
- daft_pink 1y agoisn’t the problem with the benchmarks that most people running ai locally are running way lower weights? i have an m4 studio with a lot of unified memory and i’m still no where near running a 120b model. i’m at like 30b apple or nvidia’s going to have to sell 1.5 tb ram machines before benchmark performance is going to be comparable Plus when you use claude or openai, these days it’s performing google searches etc that my local model isn’t doing.
- BoorishBears 1y agoNo, I've deployed a lot of open weight models and the gap between closed source is there even at larger sizes. I'm running a 400B parameter model at FP8 and it still took a lot of post-training to get an even somewhat comparable performance - I think a lot of people implicitly bake in some grace because the models are open weights, and that's not unreasonable because of the flexibility... but in terms of raw performance it's not even close. GPT-3.5 has better world knowledge than some 70B models, and a few even larger.
- daft_pink 1y agoyou're killing my dream of blowing $50-100k on a desktop supercomputer next year and being able to do everything locally ;) "the hacker news dream" - a house, 2 kids, and a desktop supercomputer that can run a 700B model.
- n_kr 1y agoIt may be the way I use it, but qwen3-coder (30b with ollama) is actually helping me with real world tasks. Its a bit worse than big models for the way I use it, but absolutely useful. I do use ai tools with very specific instructions though, like file paths, line numbers if I can, and specific direction about what to do, my own tools, etc. so that may be why I don't see such a huge difference from big models. I should try Kimi K2 too.
- refulgentis 1y agoYou'll see good results, Kimi is basically a micro dosing Sonnet lol. V v v reliable tool calls, but, because it's micro dosing, you don't wanna use it for implementing OAuth, maybe adding comments or strict direction (i.e. a series of text mutations)
- Art9681 1y agoIt has everything to do with the way you use it. And the biggest difference is how fast the model/service can process context. Everything is context. It's the difference between you iterating on an LLM boosted goal for an hour vs 5 minutes. If your workflow involves chatting with an LLM and manually passing chunks, and manually retrieving that response, and manually inserting it, and manually testing.... You get the picture. Sure, even last year's local LLM will do well in capable hands in that scenario. Now try pushing over 100,000 tokens in a single call, every call, in an automated process. I'm talking the type of workflows where you push over a million tokens in a few minutes, over several steps. That's where the moat, no, the chasm, between local setups and a public API lies. No one who does serious work "chats" with an LLM. They trigger workflows where "agents" chew on a complex problem for several minutes. That's where local models fold.
- torginus 1y agoNot sure about benchmarks, but I did use Deepseek when it was novel and cool for a variety of tasks before going back to Claude, and in my experience it was OK, not significantly worse for what I use these models for (writing code small functions at a time, learning about libraries etc.), tham closed stuff at the time.
- lossolo 1y agoWhile that's true for some open source models, I find DeepSeek R1 685B 0528 to be competitive with O3 in my production tests, I've been using it interchangeably for tasks I used to handle with Opus or O3.