6 ms·
For coding you always want to go with the best model in the category, not something that would be the best model if we went 1 year back which GLM 5.1 is, and I'
by mesmertech 4mo ago
For coding you always want to go with the best model in the category, not something that would be the best model if we went 1 year back which GLM 5.1 is, and I'm saying that as a big fan of GLM cause I run a translation site where GLM is good enough for the price.
Most of the money right now is in coding. Openai and Anthropic just have to be 6 months ahead of SOTA open source models and they'll capture most of the enterprise and dev market
- EGreg 4mo agoMost work is not coding. And also, people have it wrong… their models are not the main problem anymore. It’s the RAG
- obsidianbases1 4mo agoDepending on RAG is a workflow problem, not an AI problem
- tomrod 4mo agoWould love to hear more about your thought about the RAG.
- simonw 4mo agoI think RAG is a mostly outdated concept now, it's been subsumed by the idea of a "agent harness" which is exactly what Claude Code and Claude Cowork and OpenAI Codex and Claude.ai and ChatGPT themselves have now become. An agent harness with access to a good search tool is a much more interesting thing than 2024-era RAG systems.
- EGreg 4mo agoAnd how exactly does the agent harness surface ALL the right places that need to be updated, and reason about functions and APIs?
- tomrod 4mo agoI appreciate where you are coming from, as you have surfed the front of the wave of GenAI for years. From my point of view, there is interesting because something is SOTA, and there is interesting because there is still more to build. I definitely understand state of RAG tech. I also view it as barely utilized versus what we can do with it, hence my question. Agent harnesses integrated into good search tools are definitely interesting. Knowledgebasing with partitions and similar structure also remains fruitful for applications, above and beyond standard ElasticSearch on a cache.
- Traveler42 4mo agoI generally agree with this, but would note that it assumes that the data is accessible from a web search. Some data sources will be private.
- simonw 4mo agoYou can configure extra search tools that search private data.
- kgwgk 4mo agoFor coding like for everything else in life cost is a factor.
- mesmertech 4mo agoCost for the value delivered. Like if you offered the current SOTA open source models at $0.1/M, I still think I'd be using Opus or 5.5 at $30/M. Or say GPT 5 which was released Aug 25, I don't think I'd use it for coding for even $0.1. I'd def find other uses for it(translations, agentic workflows, prompt guards etc), but for coding I don't think I'd ever completely switch to a SOTA open model Unless ofc there was an actual speed difference, only reason I'd be willing to go with a worse model couple of percent worse than current best model is if the speed was at least 5x higher. Looking forward to kimi k2.6 offered publicly by Cerebras
- kgwgk 4mo ago> I still think I'd be using That's fine. Other people may not want to pay 300x more and will rather make do with last year's SOTA. > For coding you always want to go with the best model Maybe you meant "For coding I always want to go with the best model"?
- mesmertech 4mo agoBased on current market for LLMs I'd say my use of "you" in the general is fine. Even openrouter which doesn't capture all of the SOTA closed models but nearly all of opensource model usage has Opus as 1st(on last week) on "Programming" category and 3rd in overall rankings https://openrouter.ai/rankings https://openrouter.ai/rankings
- simonw 4mo agoI'd trust the OpenRouter rankings a lot more if they exposed the number of unique users for each model, as opposed to just a token count. Currently I have no way of telling if big changes in their rankings are caused by a single "whale" switching providers, or if it's a more meaningful trend.
- binary0010 4mo agoYes I'm an engineer (20 years most in games/graphics industry) and only use it for code. I've been using glm 5.1 this week a lot. I went in expecting another "decent" but not really "up to standard" open source model. I highly doubt I'll ever use Claude again. I think you are wrong about Claude being any significant level better
- cassianoleal 4mo agoI've been mostly coding with GLM-5.1 as well and I agree with you. DeepSeek V4 Flash is another very good surprise. Incredibly cheap, fast and effective.
- MaKey 4mo agoI've been using DeepSeek v4 Flash with OpenCode for the whole week to refactor a Terraform code base I inherited and it worked surprisingly well.
- aspenmartin 4mo agoWell I think there are a multitude of harder measurements that would disagree with you, but ultimately there is absolutely a use case for cheaper open models (or even cheaper tiers of proprietary models) and in fact the unsolved optimization everyone is trying to get to is how much spend to use for a given task. But there will always be a market, especially in enterprise, for the best performance there is to offer
- ggttk 4mo agoWhy are you boosting so hard? Lmao either you’re a paid poster or you own stock in a frontier firm. Which one?
- aspenmartin 4mo agoSadly neither :(
- blackjack_ 4mo agoThis is a silly take. There is a line of "good enough" for most coding (most CRUD apps and APIs are nothing special), and once we are past that, nobody will care about having the "newest, best" model except extreme outliers. And this base "good enough" model will become an ultra cheap commodity as we already see with GLM, deepseek, etc.
- mesmertech 4mo agoAs long as closed models are 6 months ahead I won't be switching from them to prev. 6 month SOTA open source models. Maybe its just a different calculation if you're in a job, but as an indiehacker I'll take any edge I can get Ofc again, can be convinced to switch if there's however a clear speed difference, like 5x+ for a open source sota even if it was SOTA for 6 months ago
- Andrex 4mo ago> For coding you always want to go with the best model in the category Will this always be true? There will never be an event horizon/point of diminishing returns where something not-bleeding-edge is "good enough" for 51%+ of users?
- mesmertech 4mo agoAs long as closed source is 6 months ahead in terms of current difference. Although this is hard to figure out using simple percent based coding benchmarks, you def. notice it when you're actually trying to do a long task. Even simple things like UI "taste" is enough for me to use opus instead of 5.5 though even though 5.5 is strictly better for anything that doesn't have a UI, ie backend, scripts, making agent workflows etc
- dogleash 4mo ago> For XXX you always want to go with XXX, not XXX Oh, hey, I recognize you. Thank you for the very forward and thorough orbital sander recommendation at Home Depot. That's exactly what I wanted to deal with on my holiday weekend. You just know so much about this and the rest of us are simple passersbys.
- mesmertech 4mo agoYep sorry was just pulling it out my rear, not like a market trend that nearly every enterprise uses Anthropic or Openai models for coding or that Anthropic has had such ridiculous growth that they're 10x-ing year over year
- dogleash 4mo agoI'm ribbing you for writing like a condescending guru that invalidates the evaluatory capability of your peers. Not the meat of your evaluation (not to say that it's any good either, just that it's irrelevant).
- odie5533 4mo agoIf I generate code with Claude, ChatGPT, and GLM 5.1, I can't say which model is which reliably. I exclusively use Claude more out of superstition than reason.
- eikenberry 4mo ago> For coding you always want to go with the best model in the category [..] And this is why many companies go out of business. You always want the best bang for your buck, sometimes this is the "best model" and sometimes it is not.
- lunar_mycroft 4mo ago> For coding you always want to go with the best model in the category This is transparently false, because the best "model" is still competent human developers. They're just more expensive. If you're willing to use current LLMs at all, it means you're willing to sacrifice quality for a better price, and your disagreement with the comment you were replying to is entirely about what the optimum tradeoff is.
- aspenmartin 4mo agoWell it may be false that you always want the best model, but the point is performance of you+<agent> is far more cost effective than you+someone else
- lunar_mycroft 4mo agoMaybe, but that's a different claim than the one I was responding to. And also raises the question of "if the lower quality but cheaper output of frontier models is more cost effective than humans, is the even lower quality but even cheaper output of OSS models is more cost effective still?" With an absolute rule like GP suggested ("no, you always want the best code generator") the answer is clear, but it get much murkier if you reject such rules (as you have to to be an LLM coding proponent)
- aspenmartin 4mo agoI think that’s a fair and good q and point.
- noname120 4mo agoIt was true 6 months ago, not anymore. Frontier models now outperform developers on many tasks, be it on quality/readability/maintainability, and let’s not talk about speed…
- suddenlybananas 4mo agoWhy is anthropic hiring software developers then?
- solomatov 4mo ago>For coding you always want to go with the best model in the category, not something that would be the best model if we went 1 year back which GLM 5.1 is, and I'm saying that as a big fan of GLM cause I run a translation site where GLM is good enough for the price. Currently, the difference is substantial, but what happens if capabilities saturate?
- aspenmartin 4mo agoThen the house of cards comes crumbling down, but there is so much evidence to point to this not happening that it requires a bit of a theory for how that may happen
- solomatov 4mo ago> but there is so much evidence to point to this not happening Could you explain this?
- aspenmartin 4mo agoWell I think there are several fairly stable trends that paint a pretty compelling picture: - performance scales with compute very very reliably. We have “scaling laws” (and have for years) and they are almost miraculously stable and show no sign of being invalidated at all even at the very largest scales. There are some theoretical bases for this though I’m not as familiar with the details - these scaling laws are on an unintuitive quantity (validation loss on pretraining datasets), so we can look at downstream performance. Benchmarks are a minefield of junk but there are many decent ones and enough variety of techniques and data sources and scoring methods etc that in aggregate they are useful. The single number that I think is the best summary statistic across the crazy (O(100k)) number of benchmarks is the “epoch capability index” (just some branding over a reasonably standard statistical model that was really well thought out and a great idea). The trends in this are extremely stable. Eyeballing the trend over time on their graph we’re getting basically a GPT-4 to GPT-5 level capability improvement every ~18 months - coding agents are not limited by the quality of the human training data they’re trained on, this is such a massive misconception: human data is only a bootstrap to a reinforcement learning phase. This combined with the fact that we have verifiable rewards means it’s just a matter of when not if for any given level of reliability. - the massive compute investment implies that the compute that we’re building over the next 2-3 years will 10x the effective compute for training models. That combined with various R&D contributions (historically which have been very significant and there is no shortage of wins here), better data curation and flywheels, richer data (wait until conversation capability gets good) means we have several orders of magnitude of runway that we know of, today. In short I don’t see any compelling evidence to suggest all of the trends we observe in many different ways will end any time soon.
- RevEng 4mo agoI strongly disagree. I'm an engineer - I'm all about the fastest, cheapest thing that meets the requirements. I don't need Opus 4.7, even for my complex programming tasks. It costs over 10x other models available that still give good enough answers. Those smaller models are also a lot faster to output tokens, which saves me time. Once the model gets good enough, the returns on bigger models diminishes quickly. I don't want to spend 10x the money and wait 5x the time to get answers that are equivalent.
- yokoprime 4mo agoSame here, i can't say i've seen any difference in 4.6 vs 4.7 other than price
- danny_codes 4mo agoWhy? If it's good enough, it's good enough. Though I read the code that gets vibed so maybe my use-case is different.
- yokoprime 4mo agoIt's driven a lot by the harness too. If you're using claude code, you're actively being pushed towards newer models, even though older ones work perfectly fine for your use cases
- danny_codes 4mo agoYeah wouldn't touch ClaudeCode when there are so many better harnesses that are free and portable. Seems like a waste of time to learn a proprietary tool when the FOSS ones are better.
- Perz1val 4mo agoAnd you propose the same companies that have been cost cutting and avoiding buying you a chair for ever won't start objecting to a $200/dev/month subscription? The finance department won't have a say?
- deleted 4mo ago[deleted]
- vidarh 4mo agoI have stats from a harness that tells me glm5.1 is far more cost effective for us than Opus with the rate of defects and rework taken into account. In fact, with a decent harness I'm now increasingly favouring eHaiku over Opus for execution too. Opus is still worth it for planning, though, and far better at one-shotting things.
- r0b05 4mo agoWhy do need to go with the best model for coding?