19 ms·
Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
- DeathArrow 6mo agoWe (my wallet and I) are pretty happy with GLM 5.1 and MiniMax 2.7.
- anabranch 6mo agoI wanted to better understand the potential impact for the tokenizer change from 4.6 and 4.7. I'm surprised that it's 45%. Might go down (?) with longer context answers but still surprising. It can be more than 2x for small prompts.
- pawelduda 6mo agoNot very encouraging for longer use, especially that the longer the conversation, the higher the chance the agent will go off the rails
- someuser54541 6mo agoShould the title here be 4.6 to 4.7 instead of the other way around?
- freak42 6mo agoabsolutely!
- UltraSane 6mo agoWriting Opus 4.6 to 4.7 does make more sense for people who read left to right.
- pixelatedindex 6mo agoI’m impressed with anyone who can read English right to left.
- einpoklum 6mo agoRight to Left English - read can, who? Anyone with [which] impressed am I.
- y1n0 6mo agoYoda, you that is?
- adrian_b 6mo agoEnglish can be read in a different order than the normal order when the sentences contain words for which it is easy to guess whether they are agents or patients, e.g. when the agents are animate nouns and the patients are inanimate nouns, or when pronouns are used for the agents or patients. Otherwise, the non-standard order can be understood incorrectly. While the distinction between agents and patients is the most important that depends on word order in English, there are also other order-dependent distinctions, e.g. between beneficiary and patient, when the beneficiary is not marked by a preposition, or between a noun and its attribute, e.g. "police dog" is not the same as "dog police" and unless there is a detailed context you cannot know what is meant when the word order is wrong. English is one of the languages with the most rigid word order. There are languages, especially among older languages, where almost any word order can be used without causing ambiguities, because all the possible roles of the words are marked by prepositions, postpositions or affixes (or sometimes by accentuation shifts).
- einpoklum 6mo agoIn my example, the RTL reading is indeed a misunderstanding. I even cheated, because it really should have been: > Left to Right English - read can, who? Anyone with [which] impressed am I. and the causation is wrong; instead of the ability being impressive, it's the impressive character than allows reading in the opposite order. So, you're right, and now I'll wait for the dog police to come pick me up.
- jlongman 6mo agoYou might like https://en.wikipedia.org/wiki/Boustrophedon https://en.wikipedia.org/wiki/Boustrophedon
- embedding-shape 6mo agoBut the page is not in a language that should be read right to left, doesn't that make that kind of confusing?
- usrnm 6mo agoDid you mean "right to left"?
- embedding-shape 6mo agoI very much did, it got too confusing even for me. Thanks!
- UltraSane 6mo agoI kept mentally verifying that English is written left to right.
- bee_rider 6mo agoErr, how so?
- therobots927 6mo agoWow this is pretty spectacular. And with the losses anthro and OAI are running, don’t expect this trend to change. You will get incremental output improvements for a dramatically more expensive subscription plan.
- deleted 6mo ago[deleted]
- falcor84 6mo agoIndeed, and if we accept the argument of this tech approaching AGI, we should expect that within x years, the subscription cost may exceed the salary cost of a junior dev. To be clear, I'm not saying that it's a good thing, but it does seem to be going in this direction.
- dgellow 6mo agoIf LLMs do reach AGI (assuming we have an actual agreed upon definition), it would make sense to pay way more than a junior salary. But also, LLMs won’t give us AGI (again, assuming we have an actual, meaningful definition)
- therobots927 6mo agoI absolutely do not accept that argument. It’s clear models hit a plateau roughly a year ago and all incremental improvements come at an increasingly higher cost. And junior devs have never added much value. The first two years of any engineer’s career is essentially an apprenticeship. There’s no value add from have a perpetually junior “employee”.
- deleted 6mo ago[deleted]
- justindotdev 6mo agoi think it is quite clear that staying with opus 4.6 is the way to go, on top of the inflation, 4.7 is quite... dumb. i think they have lobotomized this model while they were prioritizing cybersecurity and blocking people from performing potentially harmful security related tasks.
- vessenes 6mo ago4.7 is super variable in my one day experience - it occasionally just nails a task. Then I'm back to arguing with it like it's 2023.
- aenis 6mo agoMy experience as well, unfortunately. I am really looking forward to reading, in a few years, a proper history of the wild west years of AI scaling. What is happening in those companies at the moment must be truly fascinating. How is it possible, for instance, that I never, ever, had an instance of not being able to use Claude despite the runaway success it had, and - i'd guess - expotential increase in infra needs. When I run production workloads on vertex or bedrock i am routinely confronted with quotas, here - it always works.
- dgellow 6mo agoThat has been my Friday experience as well… very frustrating to go back to the arguing, I forgot how tense that makes me feel
- bcherny 6mo agoHey, Boris from the Claude Code team here. People were getting extra cyber warnings when using old versions of Claude Code with Opus 4.7. To fix it, just run claude update to make sure you're on the latest. Under the hood, what was happening is that older models needed reminders, while 4.7 no longer needs it. When we showed these reminders to 4.7 it tended to over-fixate on them. The fix was to stop adding cyber reminders. More here: https://x.com/ClaudeDevs/status/2045238786339299431 https://x.com/ClaudeDevs/status/2045238786339299431
- coldtea 6mo agoThis, the push towards per-token API charging, and the rest are just a sign of things to come when they finally establish a moat and full monoply/duopoly, which is also what all the specialized tools like Designer and integrations are about. It's going to be a very expensive game, and the masses will be left with subpar local versions. It would be like if we reversed the democratization of compilers and coding tooling, done in the 90s and 00s, and the polished more capable tools are again all proprietary.
- throwaway041207 6mo agoYep, between this and the pricing for the code review tool that was released a couple weeks ago (15-25 a review), and the usage pricing and very expensive cost of Claude Design, I do wonder if Anthropic is making a conscious, incremental effort to raise the baseline for AI engineering tasks, especially for enterprise customers. You could call it a rug pull, but they may just be doing the math and realize this is where pricing needs to shift to before going public.
- zozbot234 6mo agoThere's been speculation that the code review might actually be Mythos. It would seem to explain the cost.
- quux 6mo agoIf only there were an Open AI company who's mandate, built into the structure of the company, were to make frontier models available to everyone for the good of humanity. Oh well
- slowmovintarget 6mo agoThings used to be better... really. OpenAI was built as you say. Google had a corporate motto of "Don't be evil" which they removed so they could, um, do evil stuff without cognitive dissonance, I guess. This is the other kind of enshitification where the businesses turn into power accumulators.
- 6mo ago
- ai_slop_hater 6mo agoDoes anyone know what changed in the tokenizer? Does it output multiple tokens for things that were previously one token?
- quux 6mo agoIt must, if it now outputs more tokens than 4.6's tokenizer for the same input. I think the announcement and model cards provide a little more detail as to what exactly is different
- deleted 6mo ago[deleted]
- l5870uoo9y 6mo agoMy impression the reverse is true when upgrading to GPT-5.4 from GPT-5; it uses fewer tokens(?).
- andai 6mo agoBut with the same tokenizer, right? The difference here is Opus 4.7 has a new tokenizer which converts the same input text to a higher number of tokens. (But it costs the same per token?) > Claude Opus 4.7 uses a new tokenizer, contributing to its improved performance on a wide range of tasks. This new tokenizer may use roughly 1x to 1.35x as many tokens when processing text compared to previous models (up to ~35% more, varying by content), and /v1/messages/count_tokens will return a different number of tokens for Claude Opus 4.7 than it did for Claude Opus 4.6. > Pricing remains the same as Opus 4.6: $5 per million input tokens and $25 per million output tokens. ArtificialAnalysis reports 4.7 significantly reduced output tokens though, and overall ~10% cheaper to run the evals. I don't know how well that translates to Claude Code usage though, which I think is extremely input heavy.
- ausbah 6mo agois it really unthinkable that another oss/local model will be released by deepseek, alibaba, or even meta that once again give these companies a run for their money
- pitched 6mo agoNow that Anthropic have started hiding the chain of thought tokens, it will be a lot harder for them
- zozbot234 6mo agoAnthropic and OpenAI never showed the true chain of thought tokens. Ironically, that's something you only get from local models.
- amelius 6mo agoI'm betting on a company like Taalas making a model that is perhaps less capable but 100x as fast, where you could have dozens of agents looking at your problem from all different angles simultaneously, and so still have better results and faster.
- andai 6mo agoYeah, it's a search problem. When verification is cheap, reducing success rate in exchange for massively reducing cost and runtime is the right approach.
- never_inline 6mo agoYou underestimating the algorithmic complexity of such brute forcing, and the indirect cost of brittle code that's produced by inferior models
- 100ms 6mo agoI'm excited for Taalas, but the worry with that suggestion is that it would blow out energy per net unit of work, which kills a lot of Taalas' buzz. Still, it's inevitable if you make something an order of magnitude faster, folk will just come along and feed it an order of magnitude more work. I hope the middleground with Taalas is a cottage industry of LLM hosts with a small-mid sized budget hosting last gen models for quite cheap. Although if they're packed to max utilisation with all the new workloads they enable, latency might not be much better than what we already have today
- kalkin 6mo agoAFAICT this uses a token-counting API so that it counts how many tokens are in the prompt, in two ways, so it's measuring the tokenizer change in isolation. Smarter models also sometimes produce shorter outputs and therefore fewer output tokens. That doesn't mean Opus 4.7 necessarily nets out cheaper, it might still be more expensive, but this comparison isn't really very useful.
- manmal 6mo agoWhy is it not useful? Input token pricing is the same for 4.7. The same prompt costs roughly 30% more now, for input.
- deleted 6mo ago[deleted]
- kalkin 6mo agoThat's valid, but it's also worth knowing it's only one part of the puzzle. The submission title doesn't say "input".
- deleted 6mo ago[deleted]
- dktp 6mo agoThe idea is that smarter models might use fewer turns to accomplish the same task - reducing the overall token usage Though, from my limited testing, the new model is far more token hungry overall
- fny 6mo agoI'm going to suggest what's going on here is Hanlon's Razor for models: "Never attribute to malice that which is adequately explained by a model's stupidity." In my opinion, we've reached some ceiling where more tokens lead only to incremental improvements. A conspiracy seems unlikely given all providers are still competing for customers and a 50% token drives infra costs up dramatically too.
- tailscaler2026 6mo agoSubsidies don't last forever.
- pitched 6mo agoRunning an open like Kimi constantly for an entire month will cost around 100-200$, being roughly equal to a pro-tier subscription. This is not my estimate so I’m more than open to hearing refutations. Kimi isn’t at all Opus-level intelligent but the models are roughly evenly sized from the guesses I’ve seen. So I don’t think it’s the infra being subsidized as much as it’s the training.
- nothinkjustai 6mo agoKimi costs 0.3/$1.72 on OpenRouter, $200 for that gives you way more than you would get out of a $200 Claude subscription. There are also various subscription plans you can use to spend even less.
- senordevnyc 6mo agoI’m using Composer 2, Cursor’s model they built on top of Kimi, and it’s great. Not Opus level, but I’m finding many things don’t need Opus level.
- RevEng 6mo agoIt's all I use at work and I've yet to find anything it can't handle. Then again, I'm a principal engineer and I already have designs in mind, so I'm giving it careful instruction and checking its work every time.
- varispeed 6mo agoHow do you get anything sensible out of Kimi?
- gadflyinyoureye 6mo agoI've been assuming this for a while. If I have a complex feature, I use Opus 4.6 in copilot to plan (3 units of my monthly limit). Then have Grok or Gemini (.25-.33) of my monthly units to implement and verify the work. 80% of the time it works every time. Leave me plenty of usage over the month.
- matt3210 6mo ago[flagged]
- ant6n 6mo agoI thought it would be to get better, to stay competitive with the competitors and free models.
- operatingthetan 6mo agoThe long-term pitch of these AI companies is that the AI will essentially replace workers for low cost. If the models don't get to a higher level of 'intelligence' and still struggle with certain basic tasks at the SOTA while also getting more expensive, then the pitch is misleading and unlikely to happen. So yes, I expect the price to go down.
- deleted 6mo ago[deleted]
- Shailendra_S 6mo ago45% is brutal if you're building on top of these models as a bootstrapped founder. The unit economics just don't work anymore at that price point for most indie products. What I've been doing is running a dual-model setup — use the cheaper/faster model for the heavy lifting where quality variance doesn't matter much, and only route to the expensive one when the output is customer-facing and quality is non-negotiable. Cuts costs significantly without the user noticing any difference. The real risk is that pricing like this pushes smaller builders toward open models or Chinese labs like Qwen, which I suspect isn't what Anthropic wants long term.
- c0balt 6mo agoOne could reconsider whether building your business on top of a model without owning the core skills to make your product is viable regardless. A smaller builder might reconsider (re)acquiring relevant skills and applying them. We don't suddenly lose the ability to program (or hire someone to do it) just because an inference provider is available.
- OptionOfT 6mo agoThat's the risk you take on. There are 2 things to consider: * Time to market. * Building a house on someone else's land. You're balancing the 2, hoping that you win the time to market, making the second point obsolete from a cost perspective, or you have money to pivot to DIY.
- duped 6mo ago> if you're building on top of these models as a bootstrapped founder This is going to be blunt, but this business model is fundamentally unsustainable and "founders" don't get to complain their prospecting costs went up. These businesses are setting themselves up to get Sherlocked. The only realistic exit for these kinds of businesses is to score a couple gold nuggets, sell them to the highest bidder, and leave.
- deleted 6mo ago[deleted]
- rachel_rig 5mo ago
- dakiol 6mo agoWe dropped Claude. It's pretty clear this is a race to the bottom, and we don't want a hard dependency on another multi-billion dollar company just to write software We'll be keeping an eye on open models (of which we already make good use of). I think that's the way forward. Actually it would be great if everybody would put more focus on open models, perhaps we can come up with something like the "linux/postgres/git/http/etc" of the LLMs: something we all can benefit from while it not being monopolized by a single billionarie company. Wouldn't it be nice if we don't need to pay for tokens? Paying for infra (servers, electricity) is already expensive enough
- nate8bit 6mo agoAny recommendations on good open ones? What are you using primarily?
- blahblaher 6mo agoqwen3.5/3.6 (30B) works well,locally, with opencode
- zozbot234 6mo agoMind you, a 30B model (3B active) is not going to be comparable to Opus. There are open models that are near-SOTA but they are ~750B-1T total params. That's going to require substantial infrastructure if you want to use them agentically, scaled up even further if you expect quick real-time response for at least some fraction of that work. (Your only hope of getting reasonable utilization out of local hardware in single-user or few-users scenarios is to always have something useful cranking in the background during downtime.)
- pitched 6mo agoFor a business with ten or more engineers/people-using-ai, it might still make sense to set this up. For an individual though, I can’t imagine you’d make it through to positive ROI before the hardware ages out.
- nate8bit 6mo agoMakes me think the model could actually not even be smarter necessarily, just more token dependent.
- hirako2000 6mo agoAsking a seller to sell less. That's an incentive difficult to reconcile with the user's benefit. To keep this business running they do need to invest to make the best model, period. It happens to be exactly what Anthropic's strategy is. That and great tooling.
- subscribed 6mo agoBut they're clearly oversubscribed, massively. And they're selling less and less (suddenly 5 hour window lasts 1 hour on the similar tasks it lasted 5 hours a week ago), so IMO they're scamming. I hope many people are making notes and will raise heat soon.
- hirako2000 6mo agoI agree. I'm rather pointing out the whole strategy dictates the outcome. Anthropic has to keep racing ahead and be stamped offering the best frontier models. It isn't optimal, so the models cost them disproportionately too much to sell at a profitable price. So they keep feeding the hype and push the costs higher, hoping there won't be too much heat and get away with it. I wouldn't like to be a leader at such company, but their pay keep them in line.
- micromacrofoot 6mo agoThe latest qwen actually performs a little better for some tasks, in my experience latest claude still fails the car wash test
- reddit_clone 6mo agoNot just _wrong_. It is confused! It is actually right in the second sentence. This was Friday, Opus 4.6. >I want to wash my car. The car wash is 50 meters away. Should I walk or drive? Walk. It's 50 meters — you're going there to clean the car anyway, so drive it over if it needs washing, but if you're just dropping it off or it's a self-service place, walking is fine for that distance.
- zozbot234 6mo agoThis is actually a good diagnostic of whether the model is skimping on the thinking loop. Try raising thinking effort and it should get it right. Of course, if you're running this in a coding harness with a whole lot of extraneous context, the model will be awfully confused as to what it should be thinking about.
- tiffanyh 6mo agoI was using Opus 4.7 just yesterday to help implement best practices on a single page website. After just ~4 prompts I blew past my daily limit. Another ~7 more prompts & I blew past my weekly limit. The entire HTMl/CSS/JS was less than 300 lines of code. I was shocked how fast it exhausted my usage limits.
- hirako2000 6mo agoI haven't used Claude. Because I suspect this sort of things to come. With enterprise subscription, the bill gets bigger but it's not like VP can easily send a memo to all its staff that a migration is coming. Individuals may end their subscription, that would appease the DC usage, and turn profits up.
- fooster 6mo agoSorry you are missing out. I use claude all day every day with max and what people are reporting here has not been my experience. My current usage is 16% and it resets Thursday.
- semcheck 6mo ago[dead]
- hirako2000 6mo agoSome research argue LLMs (up to 2025) were giving a false sense of higher productivity. That and atrophy, I will pass on what Claude is trying to accomplish. I'm not dismissing LLMs entirely, for certain cases the concerns don't apply, at least not as much.
- sync 6mo agoWhich plan are you on? I could see that happening with Pro (which I think defaults to Sonnet?), would be surprised with Max…
- templar_snow 6mo ago
- mvkel 6mo agoThe cope is real with this model. Needing an instruction manual to learn how to prompt it "properly" is a glaring regression. The whole magic of (pre-nerfed) 4.6 was how it magically seemed to understand what I wanted, regardless of how perfectly I articulated it. Now, Anth says that needing to explicitly define instructions are as a "feature"?!
- blahblaher 6mo agoConspiracy time: they released a new version just so hey could increase the price so that people wouldn't complain so much along the lines of "see this is a new version model, so we NEED to increase the price") similar to how SaaS companies tack on some shit to the product so that they can increase prices
- willis936 6mo agoThe result is the same: they lose their brand of producing quality output. However the more clever the maneuver they try to pull off the more clear it is to their customers that they are not earning trust. That's what will matter at the end of this. Poor leadership at Claude.
- operatingthetan 6mo agoThey are trying to pull a rabbit out of a hat. Not surprising that is their SOP given that AI in concept is an attempt to do the very same thing.
- templar_snow 6mo agoBrutal. I've been noticing that 4.7 eats my Max Subscription like crazy even when I do my best to juggle tasks (or tell 4.7 to use subagents with) Sonnet 4.6 Medium and Haiku. Would love to know if anybody's found ideal token-saving approaches.
- copperx 6mo agoI haven't seen a noticeable difference BUT I've been always using the context mode plugin.
- FireBeyond 6mo agoWhat plugin is this?
- vidarh 6mo agoI assume they mean: https://github.com/mksglu/context-mode https://github.com/mksglu/context-mode
- templar_snow 6mo agoYou mean this? https://github.com/mksglu/context-mode https://github.com/mksglu/context-mode Is it actually good or is this an ad?
- copperx 6mo agocorrect. ad? it's not a paid product afaik.
- monkeydust 6mo ago[flagged]
- dackdel 6mo agoreleases 4.8 and deletes everything else. and now 4.8 costs 500% more than 4.7. i wonder what it would take for people to start using kimi or qwen or other such.
- rectang 6mo agoFor now, I'm planning to stick with Opus 4.5 as a driver in VSCode Copilot. My workflow is to give the agent pretty fine-grained instructions, and I'm always fighting agents that insist on doing too much. Opus 4.5 is the best out of all agents I've tried at following the guidance to do only-what-is-needed-and-no-more. Opus 4.6 takes longer, overthinks things and changes too much; the high-powered GPTs are similarly flawed. Other models such as Sonnet aren't nearly as good at discerning my intentions from less-than-perfectly-crafted prompts as Opus. Eventually, I quit experimenting and just started using Opus 4.5 exclusively knowing this would all be different in a few months anyway. Opus cost more, but the value was there. But now I see that 4.7 is going to replace both 4.5 and 4.6 in VSCode Copilot, and with a 7.5x modifier. Based on the description, this is going to be a price hike for slower performance — and if the 4.5 to 4.6 change is any guide, more overthinking targeted at long-running tasks, rather than fine-grained. For me, that seems like a step backwards.
- trueno 6mo ago> 4.7 is going to replace both 4.5 and 4.6 as in 4.5 is no longer going to be avail? F. ive also been sticking with 4.5 that sucks
- rectang 6mo agohttps://github.blog/changelog/2026-04-16-claude-opus-4-7-is-generally-available/ https://github.blog/changelog/2026-04-16-claude-opus-4-7-is-... > Over the coming weeks, Opus 4.7 will replace Opus 4.5 and Opus 4.6 in the model picker for Copilot Pro+[...] > This model is launching with a 7.5× premium request multiplier as part of promotional pricing until April 30th.
- xstas1 6mo agoPromotional pricing? Are they saying that after the promotion, it will cost more than 7.5x??
- freely0085 6mo ago
- hgoel 6mo agoThe bump from 4.6 to 4.7 is not very noticeable to me in improved capabilities so far, but the faster consumption of limits is very noticeable. I hit my 5 hour limit within 2 hours yesterday, initially I was trying the batched mode for a refactor but cancelled after seeing it take 30% of the limit within 5 minutes. Had to cancel and try a serial approach, consumed less (took ~50 minutes, xhigh effort, ~60% of the remaining allocation IIRC), but still very clearly consumed much faster than with 4.6. It feels like every exchange takes ~5% of the 5 hour limit now, when it used to be maybe ~1-2%. For reference I'm on the Max 5x plan. For now I can tolerate it since I still have plenty of headroom in my limits (used ~5% of my weekly, I don't use claude heavily every day so this is OK), but I hope they either offer more clarity on this or improve the situation. The effort setting is still a bit too opaque to really help.
- _blk 6mo agoFrom what I understand you shouldn't wait more than 5min between prompts without compacting or clearing or you'll pay for reinitializing the cache. With compaction you still pay but it's less input tokens. (Is compaction itself free?)
- hgoel 6mo agoAh I can see how my phrasing might be misleading, but these prompts were made within 5 minutes of each other, the timing I mentioned were what Claude spent working.
- krackers 6mo ago>pay for reinitializing the cache Why can't they save the kv cache to disk then later reload it to memory?
- stavros 6mo agoProbably because the costly operation is loading it onto the GPU, doesn't matter if it's from disk or from your request.
- silverwind 6mo agoStill worth it imho for important code, but it shows that they are hitting a ceiling while trying to improve the model which they try to solve by making it more token-inefficient.
- razodactyl 6mo agoIf anyone's had 4.7 update any documents so far - notice how concise it is at getting straight to the point. It rewrote some of my existing documentation (using Windsurf as the harness), not sure I liked the decrease in verbosity (removed columns and combined / compressed concepts) but it makes sense in respect to the model outputting less to save cost. To me this seems more that it's trained to be concise by default which I guess can be countered with preference instructions if required. What's interesting to me is that they're using a new tokeniser. Does it mean they trained a new model from scratch? Used an existing model and further trained it with a swapped out tokeniser? The looped model research / speculation is also quite interesting - if done right there's significant speed up / resource savings.
- KellyCriterion 6mo agoYesterday, I killed my weekly limit with just three prompts and went into extra usage for ~18USD on top
- axeldunkel 6mo agothe better the tokenizer maps text to its internal representation, the better the understanding of the model what you are saying - or coding! But 4.7 is much more verbose in my experience, and this probably drives cost/limits a lot.
- autoconfig 6mo agoMy initial experience with Opus 4.7 has been pretty bad and I'm sticking to Codex. But these results are meaningless without comparing outcome. Wether the extra token burn is bad or not depends on whether it improves some quality / task completion metric. Am I missing something?
- zuzululu 6mo agoSame I was excited about 4.7 but seeing more anecdotes to conclude its not big of a boost to justify the extra tokenflatino Sticking with codex. Also GPT 5.5 is set to come next week.
- chandureddyvari 6mo ago[dead]
- napolux 6mo agoToken consumption is huge compared to 4.6 even for smaller tasks. Just by "reasoning" after my first prompt this morning I went over 50% over the 5 hours quota.
- bobjordan 6mo agoI've spent the past 4+ months building an internal multi-agent orchestrator for coding teams. Agents communicate through a coordination protocol we built, and all inter-agent messages plus runtime metrics are logged to a database. Our default topology is a two-agent pair: one implementer and one reviewer. In practice, that usually means Opus writing code and Codex reviewing it. I just finished a 10-hour run with 5 of these teams in parallel, plus a Codex run manager. Total swarm: 5 Opus 4.7 agents and 6 Codex/GPT-5.4 agents. Opus was launched with: `export CLAUDE_AUTOCOMPACT_PCT_OVERRIDE=35 claude --dangerously-skip-permissions --model 'claude-opus-4-7[1M]' --effort high --thinking-display summarized` Codex was launched with: `codex --dangerously-bypass-approvals-and-sandbox --profile gpt-5-4-high` What surprised me was usage: after 10 hours, both my Claude Code account and my Codex account had consumed 28% of their weekly capacity from that single run. I expected Claude Code usage to be much higher. Instead, on these settings and for this workload, both platforms burned the same share of weekly budget. So from this datapoint alone, I do not see an obvious usage-efficiency advantage in switching from Opus 4.7 to Codex/GPT-5.4.
- pitched 6mo agoI just switched fully into Codex today, off of Claude. The higher usage limits were one factor but I’m also working towards a custom harness that better integrates into the orchestrator. So the Claude TOS was also getting in the way.
- deleted 6mo ago[deleted]
- bparsons 6mo agoHad a pretty heavy workload yesterday, and never hid the limit on claude code. Perhaps they allowed for more tokens for the launch? Claude design on the other hand seemed to eat through (its own separate usage limit) very fast. Hit the limit this morning in about 45 mins on a max plan. I assume they are going to end up spinning that product off as a separate service.
- aray07 6mo agoCame to a similar conclusion after running a bunch of tests on the new tokenizer It was on the higher end of Anthropics range - closer to 30-40% more tokens https://www.claudecodecamp.com/p/i-measured-claude-4-7-s-new-tokenizer-here-s-what-it-costs-you https://www.claudecodecamp.com/p/i-measured-claude-4-7-s-new...
- jimkleiber 6mo agoI wonder if this is like when a restaurant introduces a new menu to increase prices. Is Opus 4.7 that significantly different in quality that it should use that much more in tokens? I like Claude and Anthropic a lot, and hope it's just some weird quirk in their tokenizer or whatnot, just seems like something changed in the last few weeks and may be going in a less-value-for-money direction, with not much being said about it. But again, could just be some technical glitch.
- hopfenspergerj 6mo agoYou can't accidentally retrain a model to use a different tokenizer. It changes the input vectors to the model.
- jimkleiber 6mo agoI appreciate you saying that, i think sometimes with ai conversations i wade into them without knowing the precise definitions of the terms, I'll try to be more careful next time. Thank you.
- gsleblanc 6mo agoIt's increasingly looking naive to assume scaling LLMs is all you need to get to full white-collar worker replacement. The attention mechanism / hopfield network is fundamentally modeling only a small subset of the full human brain, and all the increasing sustained hype around bolted-on solutions for "agentic memory" is, in my opinion, glaring evidence that these SOTA transformers alone aren't sufficient even when you just limit the space to text. Maybe I'm just parroting Yann LeCun.
- aerhardt 6mo ago> you just limit the space to text And even then... why can't they write a novel? Or lowering the bar, let's say a novella like Death in Venice, Candide, The Metamorphosis, Breakfast at Tiffany's...? Every book's in the training corpus... Is it just a matter of someone not having spent a hundred grand in tokens to do it?
- colechristensen 6mo agoWho says they can't? What's your bar that needs to be passed in order for "written a novella" to be achieved? There's a lot of bad writing out there, I can't imagine nobody has used an LLM to write a bad novella.
- aerhardt 6mo ago> What's your bar that needs to be passed I provide four examples in my comment...
- colechristensen 6mo agoYour qualification for if an LLM can write a novella is it has to be as good as The Metamorphosis? Yes, those are examples of novellas, surely you believe an LLM could write a bad novella? I'm not sure what your point is. Either you think it can't string the words together in that length or your standard is it can't write a foundational piece of literature that stays relevant for generations... I'm not sure which.
- glerk 6mo agoI'd be ok with paying more if results were good, but it seems like Anthropic is going for the Tinder/casino intermittent reinforcement strategy: optimized to keep you spending tokens instead of achieving results. And yes, Claude models are generally more fun to use than GPT/Codex. They have a personality. They have an intuition for design/aesthetics. Vibe-coding with them feels like playing a video game. But the result is almost always some version of cutting corners: tests removed to make the suite pass, duplicate code everywhere, wrong abstraction, type safety disabled, hard requirements ignored, etc. These issues are not resolved in 4.7, no matter what the benchmarks say, and I don't think there is any interest in resolving them.
- xpe 6mo ago> ... but it seems like Anthropic is going for the Tinder/casino intermittent reinforcement strategy: optimized to keep you spending tokens instead of achieving results. This part of the above comment strikes me as uncharitable and overconfident. And, to be blunt, presumptuous. To claim to know a company's strategy as an outsider is messy stuff. My prior: it is 10X to 20X more likely Anthropic has done something other than shift to a short-term squeeze their customers strategy (which I think is only around ~5%) What do I mean by "something other"? (1) One possibility is they are having capacity and/or infrastructure problems so the model performance is degraded. (2) Another possibility is that they are not as tuned to to what customers want relative to what their engineers want. (3) It is also possible they have slowed down their models down due to safety concerns. To be more specific, they are erring on the side of caution (which would be consistent with their press releases about safety concerns of Mythos). Also, the above three possibilities are not mutually exclusive. I don't expect us (readers here) to agree on the probabilities down to the ±5% level, but I would think a large chunk of informed and reasonable people can probably converge to something close to ±20%. At the very least, can we agree all of these factors are strong contenders: each covers maybe at least 10% to 30% of the probability space? How short-sighted, dumb, or back-against-the-wall would Anthropic have to be to shift to a "let's make our new models intentionally _worse_ than our previous ones?" strategy? Think on this. I'm not necessarily "pro" Anthropic. They could lose standing with me over time, for sure. I'm willing to think it through. What would the world have to look like for this to be the case. There are other factors that push back against claims of a "short-term greedy strategy" argument. Most importantly, they aren't stupid; they know customers care about quality. They are playing a longer game than that. Yes, I understand that Opus 4.7 is not impressing people or worse. I feel similarly based on my "feels", but I also know I haven't run benchmarks nor have I used it very long. I think most people viewed Opus 4.6 as a big step forward. People are somewhat conditioned to expect a newer model to be better, and Opus 4.7 doesn't match that expectation. I also know that I've been asking Claude to help me with Bayesian probabilistic modeling techniques that are well outside what I was doing a few weeks ago (detailed research and systems / software development), so it is just as likely that I'm pushing it outside its expertise.
- monkpit 6mo agoDoes this have anything to do with the default xhigh effort?
- ivanfioravanti 6mo agoProbably due to the new tokenizer: https://www.claudecodecamp.com/p/i-measured-claude-4-7-s-new-tokenizer-here-s-what-it-costs-you https://www.claudecodecamp.com/p/i-measured-claude-4-7-s-new...
- alphabettsy 6mo agoI’m trying to understand how this is useful information on its own? Maybe I missed it, but it doesn’t tell you if it’s more successful for less overall cost? I can easily make Sonnet 4.6 cost way more than any Opus model because while it’s cheaper per prompt it might take 10x more rounds (or never) solve a problem.
- senordevnyc 6mo agoEverything in AI moves super quickly, including the hivemind. Anthropic was the darling a few weeks ago after the confrontation with the DoD, but now we hate them because they raised their prices a little. Join us!
- couchdb_ouchdb 6mo agoComments here overall do not reflect my experience -- i'm puzzled how the vast majority are using this technology day to day. 4.7 is absolute fire and an upgrade on 4.6.
- Gareth321 6mo agoI suspect the distinction is API vs subscription. The app has some kind of very restrictive system prompt which appears to heavily restrict compute without some creative coaxing. API remains solid. So if you're using OpenCode or some other harness with an API key, that's why you're still having a good time.
- jbrooks84 6mo agoAmen, yes it uses more tokens and thinks longer but it's amazing
- varispeed 6mo agoI spent one day with Opus 4.7 to fix a bug. It just ran in circles despite having the problem "in front of its eyes" with all supporting data, thorough description of the system, test harness that reproduces the bug etc. While I still believe 4.7 is much "smarter" than GPT-5.4 I decided to give it ago. It was giving me dumb answers and going off the rails. After accusing it many times of being a fraud and doing it on purpose so that I spend more money, it fixed the bug in one shot. Having a taste of unnerfed Opus 4.6 I think that they have a conflict of interest - if they let models give the right answer first time, person will spend less time with it, spend less money, but if they make model artificially dumber (progressive reasoning if you will), people get frustrated but will spend more money. It is likely happening because economics doesn't work. Running comparable model at comparable speed for an individual is prohibitively expensive. Now scale that to millions of users - something gotta give.
- fmckdkxkc 6mo agoI enjoy using Claude but I find the vibing stuff starts to cause source-code amnesia. Even if I design something and put forth a thoughtful plan, the more I increase my output the less I feel the “vibes”. It’s funny everyone says “the cost will just go down” with AI but I don’t know. We need to keep the open source models alive and thriving. Oh, but wait the AI companies are buying all the hardware.
- QuadrupleA 6mo agoOne thing I don't see often mentioned - OpenAI API's auto token caching approach results in MASSIVE cost savings on agent stuff. Anthropic's deliberate caching is a pain in comparison. Wish they'd just keep the KV cache hot for 60 seconds or so, so we don't have to pay the input costs over and over again, for every growing conversation turn.
- QuadrupleA 6mo agoDefinitely seems like AI money got tight the last month or two - that the free beer is running out and enshittification has begun.
- jeremie_strand 6mo ago[dead]
- alekseyrozh 6mo agoIs it just me? I don't feel difference between 4.6 and 4.7
- fathermarz 6mo agoI have been seeing this messaging everywhere and I have not noticed this. I have had the inverse with 4.7 over 4.6. I think people aren’t reading the system cards when they come out. They explicitly explain your workflow needs to change. They added more levels of effort and I see no mention of that in this post. Did y’all forget Opus 4? That was not that long ago that Claude was essentially unusable then. We are peak wizardry right now and no one is talking positively. It’s all doom and gloom around here these days.
- RevEng 6mo agoI have used nothing but Sonnet and composer for a year and they work fine. LLMs were certainly not unusable before and Opus is certainly not necessary, especially considering the cost. People get excited by new records on benchmarks but for most day to day work the existing models are sufficient and far more efficient.
- gck1 6mo ago> They explicitly explain your workflow needs to change How about - don't break my workflow unless the change is meaningful? While we're at it, either make y in x.y mean "groundbreaking", or "essentially same, but slightly better under some conditions". The former justifies workflow adjustments, the latter doesn't.
- andai 6mo agoFor a fair comparison you need to look at the total cost, because 4.7 produces significantly fewer output tokens than 4.6, and seems to cost significantly less on the reasoning side as well. Here is a comparison for 4.5, 4.6 and 4.7 (Output Tokens section): https://artificialanalysis.ai/?models=claude-opus-4-7%2Cclaude-opus-4-6-adaptive%2Cclaude-opus-4-5-thinking https://artificialanalysis.ai/?models=claude-opus-4-7%2Cclau... 4.7 comes out slightly cheaper than 4.6. But 4.5 is about half the cost: https://artificialanalysis.ai/?models=claude-opus-4-7%2Cclaude-opus-4-6-adaptive%2Cclaude-opus-4-5-thinking#cost https://artificialanalysis.ai/?models=claude-opus-4-7%2Cclau... Notably the cost of reasoning has been cut almost in half from 4.6 to 4.7. I'm not sure what that looks like for most people's workloads, i.e. what the cost breakdown looks like for Claude Code. I expect it's heavy on both input and reasoning, so I don't know how that balances out, now that input is more expensive and reasoning is cheaper. On reasoning-heavy tasks, it might be cheaper. On tasks which don't require much reasoning, it's probably more expensive. (But for those, I would use Codex anyway ;)
- matheusmoreira 6mo agoIt thinks less and produces less output tokens because it has forced adaptive thinking that even API users can't disable. Same adaptive thinking that was causing quality issues in Opus 4.6 not even two weeks ago. The one bcherny recommended that people disable because it'd sometimes allocate zero thinking tokens to the model. https://news.ycombinator.com/item?id=47668520 https://news.ycombinator.com/item?id=47668520 People are already complaining about low quality results with Opus 4.7. I'm also spotting it making really basic mistakes. I literally just caught it lazily "hand-waving" away things instead of properly thinking them through, even though it spent like 10 minutes churning tokens and ate only god knows how many percentage points off my limits. > What's the difference between this and option 1.(a) presented before? > Honestly? Barely any. Option M is option 1.(a) with the lifecycle actually worked out instead of hand-waved. > Why are you handwaving things away though? I've got you on max effort. I even patched the system prompts to reduce this. > Fair call. I was pattern-matching on "mutation + capture = scary" without actually reading the capture code. Let me do the work properly. > You were right to push back. I was wrong. Let me actually trace it properly this time. > My concern from the first pass was right. The second pass was me talking myself out of it with a bad trace. It's just a constant stream of self-corrections and doubts. Opus simply cannot be trusted when adaptive thinking is enabled. Can provide session feedback IDs if needed.
- nmeofthestate 6mo agoIs this a weird way of saying Opus got "cheaper" somehow from 4.6 to 4.7?
- hereme888 6mo ago> Opus 4.7 (Adaptive Reasoning, Max Effort) cost ~$4,406 to run the Artificial Analysis Intelligence Index, ~11% less than Opus 4.6 (Adaptive Reasoning, Max Effort, ~$4,970) despite scoring 4 points higher. This is driven by lower output token usage, even after accounting for Opus 4.7's new tokenizer. This metric does not account for cached input token discounts, which we will be incorporating into our cost calculations in the near future.
- throwatdem12311 6mo agoPrice is now getting to be more in line with the actual cost. Th models are dumber, slower and more expensive than what we’ve been paying up until now. OpenAI will do it too, maybe a bit less to avoid pissing people off after seeing backlash to Anthropic’s move here. Or maybe they won’t make it dumber but they’ll increase the price while making a dumber mode the baseline so you’re encouraged to pay more. Free ride is over. Hope you have 30k burning a hole in your pocket to buy a beefy machine to run your own model. I hear Mac Studios are good for local inference.
- eezing 6mo agoNot sure if this equates to more spend. Smarter models make fewer mistakes and thus fewer round trips.
- kuzivaai 6mo ago[dead]
- gverrilla 6mo agoYeah I'm seriously considering dropping my Max subscription, unless they do something in the next few days - something like dropping Sonnet 4.7 cheap and powerful.
- bertil 6mo agoMy impression is that the quality of the conversation is unexpectedly better: more self-critical, the suggestions are always critical, the default choices constantly best. I might not have as many harnesses as most people here, so I suspect it’s less obvious but I would expect this to make it far more valuable for people who haven’t invested as much. After a few basic operations (retrospective look at the flow of recent reviews, product discussions) I would expect this to act like a senior member of the team, while 4.6 was good, but far more likely to be a foot-gun.
- cooldk 6mo agoAnthropic may have its biases, but its product is undeniably excellent.
- ianberdin 6mo agoOpus 4.6 is the main model on https://playcode.io https://playcode.io. Not a secret, the model is the best on the world. Yet it is crazy expensive and this 35% is huge for us. $10,000 becomes $13,500. Don’t forget, anthropic tokenizer also shows way more than other providers. We have experimented a lot with GLM 5.1. It is kinda close, but with downsides: no images, max 100K adequate context size and poor text writing. However, a great designer. So there is no replacement. We pray.
- npollock 6mo agoYou can configure the status line to get a feel for token usage: [Opus 4.6] 3% context | last: 5.2k in / 1.1k out add this to .claude/settings.json "statusLine": { "type": "command", "command": "jq -r '\"[\\(.model.display_name)] \\(.context_window.used_percentage // 0)% context | last: \\(((.context_window.current_usage.input_tokens // 0) / 1000 * 10 | floor / 10))k in / \\(((.context_window.current_usage.output_tokens // 0) / 1000 * 10 | floor / 10))k out\"'" }
- atleastoptimal 6mo agoThe whole version naming for models is very misleading. 4 and 4.1 seem to come from a different "line" than 4.5 and 4.6, and likewise 4.7 seems like a new shape of model altogether. They aren't linear stepwise improvements, but I think overall 4.7 is generally "smarter" just based on conversational ability.
- gck1 6mo agoAnthropic is playing a strange game. It's almost like they want you to cancel the subscription if you're an active user and only subscribe if you only use it once per month to ask what the weather in Berlin is. First they introduce a policy to ban third party clients, but the way it's written, it affects claude -p too, and 3 months later, it's still confusing with no clarification. Then they hide model's thinking, introduce a new flag which will still show summaries of thinking, which they break again in the next release, with a new flag. Then they silently cut the usage limits to the point where the exact same usage that you're used to consumes 40% of your weekly quota in 5 hours, but not only they stay silent for entire 2 weeks - they actively gaslight users saying they didn't change anything, only to announce later that they did, indeed change the limits. Then they serve a lobotomized model for an entire week before they drop 4.7, again, gaslighting users that they didn't do that. And then this. Anthropic has lost all credibility at this point and I will not be renewing my subscription. If they can't provide services under a price point, just increase the price or don't provide them. EDIT: forgot "adaptive thinking", so add that too. Which essentially means "we decide when we can allocate resources for thinking tokens based on our capacity, or in other words - never".
- BrianneLee011 6mo agoWe should clarify 'Scaling up' here. Does higher token consumption actually correlate with better accuracy, or are we just increasing overhead?
- erelong 6mo agowas shocked to see phone verification roll out like last month as well... yikes
- kziad 6mo ago[dead]
- nickvec 6mo agoFor all intents and purposes, aren't the "token change" and "cost change" metrics effectively the same thing?
- Frannky 6mo agoMy subscription was up for renewal today. I gave it a shot with OpenCode Go + Xiaomi model. So far, so good—I can get stuff done the same way it seems.
- vicchenai 6mo agoran into this yesterday building a data pipeline that pulls SEC filings. same prompt, same context window, 4.7 chewed through noticeably more of my api budget than 4.6 did. the output wasnt obviously better either, just... more expensive. what bugs me is the tokenizer change feels like a stealth price hike. if you're charging the same $/token but the same text now costs 35% more tokens, thats just a 35% price increase with extra steps. at least be upfront about it.
- lucid-dev 6mo agoUm, I keep getting "invalid" request despite trying my prompt through various formats as provided in the examples. It looks like you don't allow testing of anything beyond a certain token size. Which makes your test kind of pointless, because if you are chatting about AI with something that's only a few hundred tokens, the data your collecting is pretty minimal and specific, not something that's generally applicable or relevant to wider user outside of that specific context.
- Syzygies 6mo agoI'm a retired mathematician hoping to finish a second proof of a major theorem before I die. AI needs to understand my math and help me code. What I spend on AI isn't going to deplete my retirement savings. So far, Opus 4.7 seems a bit smarter than Opus 4.6 for my use case. That's my only concern. Is an $80 bottle of wine a better value than a $20 or $40 bottle of wine? Pretty much never. If there are those of us willing to buy $80 bottles of wine, of course the market will facilitate this. People can use whatever model they want. I'm too worried about worms crawling through my dead body to waste time on any but the smartest model any moment can offer.
- deleted 6mo ago[deleted]
- xvector 6mo agoI too am finding 4.7 a significant upgrade, it's hard to go back to 4.6 for me. I don't understand everyone calling it a disappointment but clowning on Anthropic is the trendy move these days. And what's missing in all these token count complaints is that 4.7 is actually cheaper overall anyways because it produces fewer output tokens.
- WarmWash 6mo agoIf you are doing math I'd stick with Gemini and ChatGPT. Anthropic doesn't seem interested in doing math, whereas google and OAI trade blows over it (read: are doing math specific training).
- Olivia_Pan 6mo ago[flagged]
- EthanFrostHI 6mo ago[flagged]
- macinjosh 6mo agoOpus 4.7 seems smarter not wiser. More knowledge, maybe, but less grit. It often has been asking me to wrap it up or just be happy with current state, instead of working out a problem.
- isodev 6mo agoIt’s really funny how people are surprised or upset about the pricing “anomalies” of these SaaS models. If you’ve been around in tech, you know it’s probably designed to keep you outraged about it to keep engagement up and essentially free ads. The advice, as always, is to not lock yourself into it.
- someguyiguess 6mo agoHonestly… that is a fair point. It’s not what I would assume by default but now that I’ve read what you said, I mean … I don’t disagree.
- maxbeech 6mo ago[dead]
- jiusanzhou 6mo ago[dead]
- liangyunwuxu 6mo agoIf possible, I will continue to use version 4.6 until it is discontinued.
- ManlyBread 6mo agoI've tried the following prompt: "repeat the following 100 times: FFFFFFFFFFFFF AAAAAAAAAAAAAAAAAAAA" This has resulted in +92.9% cost and token difference. Submission bd2457e5, currently at the top of the leaderboard.
- jbrooks84 6mo agoI don't get all this talk on the new model. I see enhanced capabilities and more token usage. Need to use external validation and specs.
- ozgrakkurt 6mo agoThe design of this thing is atrocious. There should be a clear way to see what the +X% thing means. Is 4.7 using more or is 4.6 using more. Also there should be time distribution for the queries and a way to filter by query time. This is because Anthropic is reported to change the model quality arbitrarily in the background. Also there is no unit in table column headers. For example "Request 4.7" is this the amount of tokens 4.7 consumes? Is it output/input/reasoning etc. Really difficult to make sense of this. People get offended if what they are doing is labeled as slop but this is unfortunately the level of quality I expect from AI related content or code.
- TomGarden 6mo agoHad better, faster results by changing to medium effort. Weird, but the xhigh default chugged forever just to come back with poorer solutions than 4.6 on medium
- spencerkw 6mo agothe tokenizer change is the real story here imo. same text, same prompt, but 4.7 maps it to 1.0-1.35x more tokens at the same per-token price. that's a stealth price increase that doesn't show up on any pricing page. what makes it worse is it compounds with two other things: thinking tokens (invisible but counted against limits) and the more verbose output style. so the effective cost delta is closer to 1.5-2x, not just the 1.35x from the tokenizer alone. practically the only mitigation right now is to keep using 4.6 for tasks where you don't need the reasoning improvements and only use 4.7 when you actually need it. but that means maintaining model selection logic per-task, which most people won't bother with.
- agentseal 6mo ago[dead]
- contractlens_hn 6mo ago[dead]
- Futurmix 6mo ago[flagged]
- carlovalenti 5mo agoFor sure Opus 4.7 is more chatty and talkative, I had to explicitly state a "be concise" preference in the settings. Is anyone experiencing some (very very rare) glitch in the output? Broken words, I mean. I'm using the WEB interface extensively, adaptive thinking ON, PRO plan.