4 ms·
I can't notice any difference to 4.6 from 3 weeks ago, except that this model burns way more tokens, and produces much longer plans. To me it seem like this mod
by EmanuelB 6mo ago
I can't notice any difference to 4.6 from 3 weeks ago, except that this model burns way more tokens, and produces much longer plans. To me it seem like this model is just the same as 4.6 but with a bigger token budget on all effort levels. I guess this is one way how Anthropic plans to make their business profitable.
During the past weeks of lobotomized opus, I tried a few different open weight models side by side with "opus 4.6" on the same issue. The open weights outperformed opus 4.6, and did it way faster and cheaper. I tried the same problem against Opus 4.7 today and it did manage to find one additional edge case that is not critical, but should be logged. So based on my experience, the open weight models managed to solve the exact problem I needed fixed, while Opus 4.7 seem to think a bit more freely at the bigger picture. However Opus 4.7 also consumed way more tokens at a higher price, so the price difference was 10-20x higher on Opus compared to the open weights models. I will use Opus for code review and minor final fixes, and let the open weights models do the heavy lifting from now on. I need a coding setup I can rely on, and clearly Anthropic is not reliable enough to rely on.
Why pay 200$ to randomly get rug-pulled with no warning, when I can pay 20$ for 90% of the intelligence with reliable and higher performance?
- sanderjd 6mo agoWhich open weights models did you use for this comparison, and how are you running them?
- parasti 6mo agoWhich open weights model?
- InvisGhost 6mo agoIt goes to a different school, you wouldn't know if
- Scrounger 6mo ago> Which open weights model? Yes, I'm also wondering! Currently I'm testing out gemma4:26b and qwen3.6:35b-a3b-q4_K_M locally on my M2 Max Macbook Pro. Not the fastest, but reasonable. However, I am also interested in getting as close as possible in performance to Opus 4.6 while minimizing my costs.
- hk__2 6mo ago> I am also interested in getting as close as possible in performance to Opus 4.6 while minimizing my costs. Aren’t we all? ;)
- taffydavid 6mo agoGemma4 on an m2? That sounds promising. I have an m3 max, going to try that today
- itsdavesanders 6mo agoRemember, Open Weight doesn't necc. mean local. They are probably running on a larger version online, closer to Claude specs. (lol and probably distilled from Claude)
- deleted 6mo ago[deleted]
- misja111 6mo agoI'm actually seeing a similar thing when comparing 4.6 and 4.5. It burns a lot more tokens, does show more how it is thinking along the way, but I don't see a strong difference in the end result. Occasionally 4.6 even seems to get stuck in its 'processing' phase, while 4.5 doesn't on the same task.
- spaceman_2020 6mo agoYeah my rate limits are getting exhausted way faster now. Its also way slower and overplans unless you steer it closely. I can’t rely on this anymore.
- mattmanser 6mo agoI just don't believe you. The vast gulf between open weights and frontier models that existed 6 months ago has suddenly disappeared? It's far more likely you're just bad at assessing model output.
- jamiejquinn 6mo agoOr that gulf doesn't exist for the problems they are trying to solve?
- michaelscott 6mo agoTheir problem space may be just fine with open weight models regardless, but yes the release of gemma 4, GLM 5.1 and qwen 3.5 (and now 3.6!) have all happened in the last 6 months
- elAhmo 6mo agoIts funny to think that with a model release Anthropic can slide in some instructions ("be a bit more detailed" or something similar) that affect the token output by a few percent, 5-10%, which will not be noticeable by most users but over the course of the year would bring solid growth (once the VC craze is over, if ever) and increase income. "Regular companies" would love to have a growth like that without effectively doing anything.
- weird-eye-issue 6mo agoI like how some people are accusing them of reducing the overall token usage to screw over Claude Code users and then there are yet other people that are accusing them of deliberately increasing token usage to screw over API users (or maybe to get subscription users to upgrade, I'm not really sure)
- rrr_oh_man 6mo agoIt's almost as if there are different people with different motivations and ideas about how the world should work
- doix 6mo agoI suspect the real issue is that they just change stuff "randomly" and the experience gets worse/better cheaper/more expensive. Since you have no way of knowing when they change stuff, you can't really know if they did change something or it's just bias. I've experienced that so many times in the last month that I switched to codex. The worst part is, it could be entirely in my head. It's so hard to quantify these changes, and the effort it takes isn't worth it to me. I just go by "feeling".
- 1dom 6mo agoThe issue is business and transparency. Transparency is often in the customer's interest at the individual business's expense. There are very, very few things that can be completely transparent without giving competitors an advantage. The nice solution solution to this is to be better and faster than your competitors, but sometimes it's easier just to remove transparency.
- paulluuk 6mo agoIf open weight models are sufficient for your engineering problems, then you should absolutely use them. But I haven't seen a single open weight model that can get even close to the complexity in my projects. They sometimes work for small toy examples or leetcode puzzles, but not very any real project. Really curious what models you've found that could replace current state of the art.
- fragmede 6mo agoqwen3.6-35b-a3b, released today. https://qwen.ai/blog?id=qwen3.6-35b-a3b https://qwen.ai/blog?id=qwen3.6-35b-a3b https://news.ycombinator.com/item?id=47792764 https://news.ycombinator.com/item?id=47792764
- brunooliv 6mo agoAlso my experience
- berkes 6mo agoI've been using devstral2 with great success for a few months now. The hosted version, not running one locally or such. Devstral is open. Devstral is good, Opus better. But not much. For me, "good" is "good enough". The difference, IME lies in context engineering: skills, agents.md, subagents, tools, prompts. A Devstral with good skills performs far better than an "blank" claude code. Claude with good skills performs even better, but hardly noticable, IME. I am convinced I've plateaued. Better performance comes from improving skills and other "memory", prompting smarter, better context management and, above all, from the tooling around it and the stability of the services. I do still run Claude with Opus alongside Mistral with Devstral2. Sometimes to just compare outputs, often to doublecheck, but mostly to doublecheck my statement that the difference between Devstral2 and Opus is marginally and easily covered by better context engineering.
- berkes 6mo agoSomeone just asked my what I dislike most about Mistral and about Claude code. I run both in zed editor. Claude codes' integration is subpar - it's ACP does not report tasks, doesn't give diffs and so on. Mistral has rate limits that I hit just too often. I'm now using Mistral Pro, where this is worse, using pay-as-you-go is better but costs me 10x the pro. The agent then stops with an error.
- weird-eye-issue 6mo ago> Why pay 200$ to randomly get rug-pulled with no warning, when I can pay 20$ for 90% of the intelligence with reliable and higher performance? Then go do that. Good luck!