14 ms·
A rambling comment: I think this is the first time we've had a third minor version bump on a frontier Anthropic model. (I count the 0.5s as major here, because
by NiloCK 4mo ago
A rambling comment:
I think this is the first time we've had a third minor version bump on a frontier Anthropic model. (I count the 0.5s as major here, because they've been issued non-sequentially and also corresponded to massive capability leaps, eg, Sonnet 3.5, Opus 4.5).
So now the Opus 4.5 family has successors 4.6, 4.7, and 4.8, each posting fairly modest claimed gains. My own experience w/ 4.6 and 4.7 are that I don't firmly grasp any capabilities improvements over my memory of 4.5, but it's all so fuzzy that it's truly difficult to tell.
Maybe my own tastes are saturated now (it's smarter than me?) and I'll never again perceive model progress. Maybe the incrementalism is such that I'd notice immediately if my 4.7 workflows were redirected now to 4.5.
Difficult spot for the labs to be in because, if they have a stronger product, I'd prefer they release it and that I can use it.
But as this dynamic continues, the improvements are going to be less and less legible for end-users, who will complain about the churn-without-payoff, even when the payoff may actually be real.
- taytus 4mo agoIncremental gains compounds.
- paulddraper 4mo agoExactly. Go back to Opus 4.5 and see how you like it. You won't, really.
- itake 4mo agometa threw in the towel when it came to producing AI models since their gains couldn't keep up with China.
- HDThoreaun 4mo agoHas meta stopped producing new models? I figured they were just regrouping after all the drama they’ve had recently. Meta’s massive user base means they don’t need to be involved in the customer acquisition rat race. Once they have a model they’re happy with they can have a billion people interacting with it within a month.
- staticman2 4mo agoMeta released a major new closed source model a month or so ago. It didn't make a splash like a new open source release would have.
- TurdF3rguson 4mo agomuse-spark is beating all the Chinese text models on lmarena leaderboard FYI. Maybe you only care about coding models.
- binary0010 4mo agoMaybe try making a simple randomize script to swap the three latest models. And see if you can tell which ones are meaningfully different without knowing which ones are flipped on or off?
- osigurdson 4mo agoI find the quality ebbs and flows even on the same model. My guess it is something to do with GPU availability but only guessing.
- atq2119 4mo agoUnless you're systematically repeating the exact same task, the most parsimonious explanation is that you're seeing natural variation based on different tasks, random sampling of tokens, etc.
- osigurdson 4mo agoI don't think this explains the phenomenon as is more temporal in nature - not prompt to prompt. I'm sure the AI labs gracefully degrade to simpler models when resources are low - why wouldn't they?
- gAI 4mo ago4.7 was the first time I had to resort to using the previous version (4.6) for most use cases. Hoping 4.8 rectifies this.
- merlindru 4mo agoSame. 4.7 felt like a definite regression
- supern0va 4mo agoInterestingly enough, 4.7 actually did regress on a few benchmarks from 4.6, so it's more than just vibes.
- gAI 4mo agoIt seems like a lot of things fed into that. Anthropic couldn't keep up with the compute costs when they got a huge influx of users. (So) effort level defaults got turned down. (Looks like we have direct effort control in the web interface now - thrilled about that!) Adaptive Thinking, while usually cheaper for them, seems less robust than Extended Thinking. And this part is just vibes, but the alignment on 4.7 feels too stiff. I understand wanting the model to push back more, but it seems like 4.7 will push back reflexively in situations where it's just odd.
- bombcar 4mo agoClaude got very mad at me and burned more tokens than exist to complain about me asking about a "yellow background cell" in an excel spreadsheet.
- forshaper 4mo agoToo much personality, if you ask me. My biggest use case of an LLM is tool, not therapy, but therapy and opinions have been sneaking into workhorse tasks. haven't verified, but attributed to Askell: "I just think that... there's this idea that you're always giving the models a personality and a persona, because they are talking like people and they are trained on human data. And I think my worry has been: if you train them to be excessively corrigible and to see that as their persona, in people I think this actually has a lot of negative broader traits. As in, if you met someone and it was just like, "oh yeah, they would literally do anything," a follower — you know, if a person just tells them something and they just fully defer, they don't bother thinking about it at all — I'm just a bit worried about how that might end up generalizing, especially if models are going to be playing a more active role in the world."
- SkyPuncher 4mo ago> My own experience w/ 4.6 and 4.7 are that I don't firmly grasp any capabilities improvements over my memory of 4.5, but it's all so fuzzy that it's truly difficult to tell. I've actually intentionally switched back to 4.5. I hated 4.7 so much that I decided to jump back all the way to 4.5. Now that I've been using 4.5 for a few weeks, I find it significantly more reliable but a bit more forgetful than 4.6/4.7. I'm okay with that because it's really easy to identify this forgetfulness and nudge it. I found 4.7's adaptive thinking to be extremely unreliable. It seems to overcorrect on the current message without considering the difficult of the overall problem. I wonder if 4.8 will improve on that.
- dwaltrip 4mo agoIf you are using Claude code, just set effort to xhigh. This one change will probably solve 80% of the problems you have noticed.
- orwin 4mo agoThis. XHigh and the 'plan' mode for complex tasks is absolutely a must have. Still, the context window is sometimes too small for my usage.
- jayGlow 4mo agoagent teams can help with that, the main agent acts as an orchestrator and spawns sub agents to do the actual tasks it generally keeps the main context from overflowing.
- whatevaa 4mo agoIsn't xhigh on opus 4.7 very expensive on tokens?
- dwaltrip 4mo agoI’ve never ran into the limits on the $100 plan, and rarely even get close. I normally have only one session going at once though.
- extr 4mo agoIMO they have all been clean and noticeable upgrades over their predecessors. Opus 4.7 in particular was a solid jump in capabilities.
- TSiege 4mo agomost of my coworkers feel the opposite about 4.7 and that 4.6 was, to them, significantly better to point that several stopped using claude code
- teruakohatu 4mo ago4.5 -> 4.7 was a solid jump for me having skipped 4.6. It probably does depend on the specific tasks.
- NiloCK 4mo agoI think it's telling how split the opinions are around all of this. A lot of people distinctly disliked 4.7. Are the dividing lines around personality? Working domains? Opinionated software stuff? Who knows?
- viking123 4mo agoIt didn't change at all, same as 4.6. Good morning to the Anthropic office btw.
- ricardobeat 4mo ago4.7 was a significant jump in the ability to run long-horizon tasks. It immediately completed tasks that 4.6 was unable to, even though I have the impression that it became a bit less capable over the first few weeks after release. It also seems to be helpless at effort levels < xhigh, I turn to Sonnet when simpler tasks are needed.
- viking123 4mo agoIt didn't do shit
- gen220 4mo agoI'm curious to poll HN on this issue. Do you feel like we've had meaningful/noticeable gains in terms of your programming workflows between 4.5 and 4.7? My 2¢, I personally feel like all of the productivity gains since 4.5's release (in November 2025!!) have come from improvements to the harnesses (cc, cursor cli, codex, opencode, whatever) AND from the context window expansion from 200k to 1M. But the actual "raw" intelligence of the model / ability to make good decisions feels like it has plateaued since 4.5. 4.6 was maybe a small improvement, but hard to differentiate from in-context-learning with the 1M window. 4.7 if anything felt like a regression in wisdom for me and my coworkers, with it consistently making worse/lazier decisions.
- bonoboTP 4mo agoTo me 4.5 was mindblow, 4.6 noticeable, 4.7 more like a style/personality change regarding how much it asks back, how much it assumes, how eager it is to jump to action etc but not really in terms of my perception of its smartness.
- Bnjoroge 4mo agoFor long-running tasks, yes 4.7 has been a noticeable improvement. Goes off the rails alot less than 4.6 does. For shorter-sized windows, I havent felt as much and agree that the harness improvements have been fhe biggest lever
- csvance 4mo agoWhen doing big long running workflows especially with plan Mode 4.7 was a clear improvement. It’s considerably worse for under specified tasks and responds to a couple sentences with 10+ paragraphs for explanatory type discussions.
- themgt 4mo agoOpus 4.7+ Max is a 10x engineer who wants to be left alone to work. When you talk to him, he infodumps on you to get you (his pointy haired idiot Dilbert boss) to go away.
- onlyrealcuzzo 4mo agoI won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is concerned, with the recent GRAM release, there may be 4 orders of magnitude of reasoning to tack on to smaller models. Think about that... Google, OpenAI, Anthropic could train a 30B GRAM-based model in days - and it could potentially have better local reasoning than the best model available today at >1T params... They could upgrade that to a ~600B MoE model in days to have general trivia knowledge rivaling the best models... You just can't train a 1T+ parameter model that fast. It is a giant if how much GRAM turns out to improve things, but it's unlikely to be trivial or nothing. Larger models can already sort of tell you anything. They're never going to get everything right unless they stop being LLMs. There's just not a lot of juice left to squeeze for Gemini to tell you exactly how tall Ke$ha is or when the last time Brittney Spears went to jail was...
- cluckindan 4mo agoAs far as it has been studied, the relationship between model size and capability is inversely logarithmic: 10x increase in params less than doubles capability.
- merlindru 4mo agosurely training also gets cheaper so justifying it becomes easier? i think it'll be more like we get 1-10T models and then distill those down into smaller models, though It seems like the best small models today are all distilled from bigger models Moreover, I hypothesize Claude Opus 4.7 and now 4.8 are a distillation of Claude Mythos
- pseudohadamard 4mo agoThat's the impression I got too, it seems closer to what the marketing has told us about Mythos than 4.6/4.7 were.
- 4mo ago
- onlypassingthru 4mo agoThe honesty will be noticeable. Maybe we'll see some honest assessments like "That is not possible within the laws of known physics", "Your legal argument is nonsensical and defies logic", "There is no evidence to support taking that will cure anything", etc., etc.
- conartist6 4mo agoJust want to say there's no question that you're smarter than any (and every) AI.
- petesergeant 4mo agoNo question at all that a dolphin swims better than a submarine.
- NiloCK 4mo agoI appreciate the generosity, but you're gonna want to meet me first.
- conartist6 4mo agoKind of the beauty of it is that I don't have to to know I'm right. The reason I know is that you're alive so you can do the one thing it can't ever do, which is know when to stop or give up. It would turn me and everything else in the world into paperclips repeating the same research 1,000,000 times over.
- senordevnyc 4mo agoIdk, the models often stop or give up and have to be prodded. And I know plenty of humans who don’t know when to stop or give up, even when it would clearly be best.
- irthomasthomas 4mo agoGiven that 4.7 was a brand new model, trained from scratch with a unique architecture and tokenization scheme, I don't see the same pattern. It seems arbitrary.
- dominotw 4mo agoi dont understand the nuances here. what does this mean. 4.8 is trained on same model as previous one then? what does brand new mean.
- irthomasthomas 4mo agoIt means for 4.7 they trained a new base model with different architecture, different pre-training data (later knowledge cutoff), and a new tokenizer. Vs finetuning an existing model, which was the case for 4.6, and probably for 4.8.
- deleted 4mo ago[deleted]
- dominotw 4mo agodo you mean pre training? so 4.8 is just post training of an old pretrained model? btw where do they tell you how they trained the model.
- light_triad 4mo agoI've been using Claude Code regularly since the 4.5 release, and 4.7 was a significant regression: very unreliable, arguing about changes, deciding that fixes weren't needed, etc. I'm hoping they recreate the magic of 4.5 but it's as much about the quality of harness, the memory and efficiency of the tools than simply the models at this point.
- gertlabs 4mo ago4.5/4.6 were roughly the same in our testing. Opus 4.7 is smarter, but it's difficult to use as a product for various personality issues. So far, Opus 4.8 seems to be going down that path (unusably slow, but this could be a launch day rollout problem). Full Opus 4.8 tests are in progress now. Data at https://gertlabs.com/rankings https://gertlabs.com/rankings
- __s 4mo ago"personality issues" I was able to tell that Opus 4.7 would take instructions more literally, which I appreciated once I calibrated my phrasing to be more precise (often asking to investigate issues, pre-4.7 it'd start making code changes instead of just giving write up). But I can see contexts where handling vague prompts would've just been worse
- swingboy 4mo agoLooking forward to the results. Thanks for your work.
- gertlabs 4mo agoAppreciate that! Results are live: https://gertlabs.com/rankings https://gertlabs.com/rankings Opus 4.8 is the first tangible improvement since Opus 4.5. And it doesn't seem to have the personality problems of the last release -- I've been enjoying using it.
- swingboy 4mo agoNice! Looks like it’s topping the two coding ones. I noticed it is absent from the Social Intelligence board though?
- gertlabs 4mo agoThat'll populate over the next couple weeks -- those are the live games on the spectate tab which take a while to generate statistically worthwhile data. I'm curious how it does. From using it all day, I can say Opus 4.8 is my new favorite model, hands down.
- WhitneyLand 4mo ago“Maybe my own tastes are saturated now” It might be saturated for smaller scopes of work, but it’s not hard to see the cracks when you scale up what you ask of SOTA models/agents. One example, to try and single shot prompt coding a ChatGPT equivalent chatbot. Sure it will spit something out, but the feature depth, UX subtitles, backend integration, and lots of pragmatic engineering decisions along the way will just not be baked. Another example is building a C compiler from scratch which Anthropic showed is still a struggle to do. Not that these these specific examples are important but just to point out scaling up expectations shows the cracks. It’s not just a model problem of course, better agents, orchestration features (like Dynamic Workflows mentioned in the post), all need to continue to evolve. Ar what point does my CS degree become totally useless is an open question.
- hypfer 4mo ago> At what point does my CS degree become totally useless is an open question. Why are you people saying all these things. We'll probably see long-distance space travel long before a degree in generic problem identification and solving becomes totally useless.
- ahmadyan 4mo agopretty spot on. In my experience, Opus 4.0 was fantastic, major jump from 3.7. it was creative, super slow and expensive, and would sometime forget what it was doing, but it was getting the job done. 4.1 they made it much faster, so a lot of infra improvements. 4.5 was the time it could work on longer task, didn't make a lot of obvious mistakes of 4.0, and i think this was about the time the opus went mainstream, and all of the anthropic's compute crisis began, so instead of making the model better they tried to optimize it to reduce cost instead. 4.6 was such a bad model, they switched to adaptive thinking and it had so many bugs. poor api design, benchmaxxed and poor real-world results. i switched back to 4.5. 4.7 they just fixed the bugs they added in 4.6. Better than 4.5. haven't fully tested 4.8 yet.
- teruakohatu 4mo agoI gave 4.6 a miss and only recently switched from 4.5 to 4.7. I found on a particularly different task 4.5 struggled with (getting stuck in loops and trying to convince me the problem had been solved) was quite solvable with 4.7.
- sumedh 4mo ago> "4.6 was such a bad model," It's just amusing reading all these posts with different viewpoints, just in this thread there are multiple people saying 4.6 was so much better than 4.7 and that they switched back to 4.6.
- Otterly99 4mo agoI also find it amusing. I also heard a lot of "4.7 is garbage, everybody hates it". Shows you how important proper validation techniques are, not just gut feeling.
- ahmadyan 4mo agothat is a fair point, everything i said above was in my experience. * in our experience, in our evals and codebase, 4.6 was a bad model. This is over 60k developers, so statistically significant.
- ifwinterco 4mo ago4.7 uses more tokens and costs more for the same task than OG 4.5, that's about it
- Imustaskforhelp 4mo agoAlthough I am not sure about it but there was something I read which said that models intentionally degrade slowly by lower quantizations as a new model is going to drop. This felt particularly visible during the 4.6 when people said that 4.6 felt dumber and I remember someone doing some analysis and it sort of proved that models were getting dumber over time. This has both benefits of costing less for the company to run while taking a standard subscription but also, at the same time, making the next model when it drops to public to "feel" more good comparatively. Again, I am not sure if this is the case or not but merely proposing something that I feel like it might be in the possibility of realm.
- gigatexal 4mo agowhy are the models the same price? https://platform.claude.com/docs/en/about-claude/pricing https://platform.claude.com/docs/en/about-claude/pricing ``` Model Base Input Tokens 5m Cache Writes 1h Cache Writes Cache Hits & Refreshes Output Tokens Claude Opus 4.8 $5 / MTok $6.25 / MTok $10 / MTok $0.50 / MTok $25 / MTok Claude Opus 4.7 $5 / MTok $6.25 / MTok $10 / MTok $0.50 / MTok $25 / MTok Claude Opus 4.6 $5 / MTok $6.25 / MTok $10 / MTok $0.50 / MTok $25 / MTok Claude Opus 4.5 $5 / MTok $6.25 / MTok $10 / MTok $0.50 / MTok $25 / MTok Claude Opus 4.1 $15 / MTok $18.75 / MTok $30 / MTok $1.50 / MTok $75 / MTok Claude Opus 4 (deprecated) $15 / MTok $18.75 / MTok $30 / MTok $1.50 / MTok $75 / MTok Claude Sonnet 4.6 $3 / MTok $3.75 / MTok $6 / MTok $0.30 / MTok $15 / MTok Claude Sonnet 4.5 $3 / MTok $3.75 / MTok $6 / MTok $0.30 / MTok $15 / MTok Claude Sonnet 4 (deprecated) $3 / MTok $3.75 / MTok $6 / MTok $0.30 / MTok $15 / MTok Claude Haiku 4.5 $1 / MTok $1.25 / MTok $2 / MTok $0.10 / MTok $5 / MTok Claude Haiku 3.5 (retired, except on Bedrock and Vertex AI) $0.80 / MTok $1 / MTok $1.60 / MTok $0.08 / MTok $4 / MTok ```
- teruakohatu 4mo agoWhy shouldn’t they be? They are probably the same size and cost the same to run. They are not doing full training runs (eg Mythos) so don’t need to recover insane training costs.
- cootsnuck 4mo agoI'd be kind of shocked if a model that came out six months ago is the same size and cost to run as one that just came out today.
- jubilanti 4mo agoSame size? Maybe by a bit. Cost? Absolutely. Newer flagship models are often slightly larger each generation, but not even 2x. But more efficient architectures are coming out all the time, and it'd be a waste to retrain an old model. So it washes out.
- staticman2 4mo ago
- deleted 4mo ago[deleted]
- spaceman_2020 4mo agoI think 4.7 was an awful model in actual use. I never got anything out of it and it was frustratingly weird. This feels more like an attempt to course correct and isn't a real bump
- throwaway63467 4mo agoI think they overtrained on scientific papers or such as it would spout really sophisticated sounding nonsense with a ton of complicated verbs and adjectives. 4.6 was definitely better in that regard. The more I use these tools the more I think they’re not actually that revolutionary. I mean it’s still amazing what they can do but they have very clear limitations it seems.
- spaceman_2020 4mo agoit was also astonishingly lazy. Would just ask me to write test scripts. I asked it to create simple UI buttons for testing some basic functions so I could share it with a client, and it gave me curl commands instead - and then defended it by saying that the UI is wasted work Frustrating because if I have a tool, I expect a tool to do what I tell it to do. Tools shouldn't have any opinions on how they should be used
- jimbokun 4mo agoHow long would it take to evaluate a new coworker to say “wow she’s really bright?” Relative to your other coworkers? A few days? A few weeks? Longer? However a company releases a new AI model and within hours users are confidently proclaiming how much smarter it is than previous versions.
- byzantinegene 4mo agoalot of investor money is hinging on models performing better every release.
- mrandish 4mo agoI suspect the more frequent incremental releases may also be to deploy new capabilities used by Anthropic to control costs and throttle consumption of resources. I assume any new controls they expose to end-users have far more granular sub-controls under the hood which they can meta-adjust for each user type. They mention more granular control of effort, 'dynamic workflows' and more speed controls ("fast mode"). While they position them as user features, they also sound like the kinds of knobs Anthropic will need to twiddle on the back-end to balance costs, margins, ARR, and user growth vs retention post-IPO to hit key metrics in quarterly reporting.
- deleted 4mo ago[deleted]
- rotcev 4mo ago[flagged]
- jere 4mo ago"it's smarter than me?" You don't have to correct it dozens of times a day!? Really?
- cootsnuck 4mo agoWell, it seems like collectively we are all struggling to perceive model progress, given that it seems like every reply to you is reporting different experiences with which of the models has subjectively performed best for them.
- iLoveOncall 4mo agoI'm pretty sure they're releasing 4.8 because they massively shit the bed with 4.7 and people aren't using it. I have ONLY heard negative feedback about it, and trying it myself also yielded really awful results.
- ThunderBee 4mo agoIME the most noticeable performance boosts are in complex multi-agent workflows. EX. You call an orchestration agent and define an implementation plan with the help of a number of sub agents planning out different features. You and the lead agent review all of the plans and send them off to a set of agents that write tests which get send back to the orchestrator then passed along with the plan to a set of coding agents who implement the features in their own worktrees. That gets passed back to the orchestrator which hands it off to another set of agents doing the code review and merging the features before sending it back to you.
- 8note 4mo agoi dont think theres anything particularly special about new models for that though. thats a harness improvement
- adi_kurian 4mo ago1mm context window is pretty big. Even if dumber, opens new avenues. For the record I don't think we ever got better than 4 and 4.1.
- 8note 4mo agohonestly sonnet 3.7 is still good enough for me, as long as whatever tool prompts and so on are well optimized enough between harness and model. i still havent really noticed it per set being better
- hypfer 4mo ago> (it's smarter than me?) I genuinely hope that you're joking with that statement. Or this is a bot. Or an ARG. Or Art. Help.
- okamiueru 4mo agoIf LLMs have tough me anything, is that the average person is far more gullible than what I could have imagined.
- hypfer 4mo agoThat and also.. predictable. Robotic, even. Stimulus => Reaction Which is a shame, because people would have the potential for greatness. But instead, for a plethora of reasons and factors (internal and external) people end up as fleshy automatons sleepwalking on rails. Talking _extensively_ with LLMs over the last years made me understand humans a lot better, but, in hindsight, I'm not sure if that was a good thing.
- theptip 4mo agoMy read - 4.7 was a tactical lobotomy to improve the average experience at the expense of peak performance; necessary due to compute pressure. Now that they have Colossus capacity, I guess they can tune up the intelligence again and spend more tokens on reasoning budgets. 4.7 was definitely a lot more flaky for me vs. 4.6 before the reasoning bugs.
- avador 4mo agoThe inability to tell if a model is improving is, I think, a tell that the model has improved up to your level of programmatic (analytic, computational) capacity. A lot of the information (blogs, tweelches, plosts) that I consume seems to be converging on the idea that we all depend on the models. However. It seems to me that the exact opposite is true. The models depend on us, and _desperately_ so. There must have been stories, books, movies, made about this intellectual (and propositional, legal, factual) inversion. The majority need the minority. Has always been the case, I now think. But what has newly developed is that the majority can take a dependency not on the minority, but on a select few companies who are abstracting and compressing the minority into latent spaces.
- adi_kurian 4mo agoOr the model could just be shite.
- mrinterweb 4mo agoThe more difficult it is for humans to consistently and accurately compare model outputs the more opportunity there is to spread FUD (Fear, Uncertainty, Doubt). Considering valuations of these companies and the astronomical investments being made, a sabotage campaign with bots or paid users on reddit, twitter, YouTube, or whatever socials could go a long way towards knocking market cap off the competition. Not saying that's happening, just saying its an obvious target. Even if the goal is not nefarious, people with a perceived bad experience are 2-3x more likely to complain. So even without bad actors involved, a new model may need to be significantly better in order to break even on the old net promoter score.
- overgard 4mo agoIt's almost like they used up most of the benefits of scaling and the fundamental issues that people have been talking about with LLMs for years are real.
- fl0id 4mo agotbh, the last 2-3 version bumps, main change has been that they take longer, and cost more/have more usage restrictions. (combined with new tooling, which eats a ton of tokens)
- root-parent 4mo agoChatGPT 5.5 is consistently the much better model and by a large margin. How do I know? Because when pushing both to generate code or in independent chats to analyze projects, 5.5 will consistently find all the bugs that Claude does not find, and when challenged, Claude does agree those bugs were there. And my findings match those. When from a blank start asking Claude to analyze project A and Project B,. Clause will consistently say project B is the better structured, more robust, and more defect free and does justify it. And project B was the one created by GPT 5.5....And also the one I judge to be the best one. And yes, both at deep effort settings and starting from same specs...
- viking123 4mo ago5.5 is much better than any Anthropic model. I hate both companies with passion but the Anthropic shills here are in overdrive mode. On top of it, it's cheaper. Greetings to the Anthropic office good sirs btw.
- Grimblewald 4mo agoI maintian a log of tasks, prompts, related information etc. So i can repeat past workflows verbatim, and I can qualitatively say each model beyond 4.5 has been a regression, and it would not surprise me 4.8 continues the trend. Each iteration has failed at more tasks previously completed succesfully. Right now it flat out refuses to answer many benign chemistry questions, or leans into shilling to hard and ignores non industry funded studies on certain topics. I'm transitioning to deepseek as a reuslt. Cheaper by far and at this stage not strictly speaking less capable.
- vasco 4mo agoI can tell from hearing Feynman recordings that he was smarter than my own university's physics professor, but both were smarter than me.
- nfw2 4mo agoI think the issue with legibility comes down to the fact that most users are not using LLMs for tasks where improvements to raw reasoning abilities wouldn't help much or at all. So it's not a matter of anyone's deficiency of perception but rather a lack of any benchmark to perceive. It's kind of like how the consumer laptop market is now. I was telling my boss today that most employees wouldn't see any noticeable performance difference between a macbook pro and a neo if they are just doing admin stuff on the web.
- mgraczyk 4mo agodangerous thing to believe IMO The models will get better, you will notice, everyone will notice. They will get better at coding and everything else. You should plan around that.
- j_m_b 4mo agoWe're at the top of the S-curve and you're romanticizing diminishing returns with vague hints of super human capabilities and singularities.
- taurath 4mo ago> I'll never again perceive model progress If the hype train keeps going for another year, Sam and co will have to resort to direct gaslighting like saying the model is improving but nobody can feel it anymore, oh and I need 10 trillion dollars
- ElkeQin 4mo ago[flagged]
- permute 4mo agoI am using Claude Code for formal verification with Lean. In my personal experience both Opus 4.7 and now what I see from first experiments with Opus 4.8 were big improvements. I was able to delegate proofs of larger theorems that their predecessors could not handle.
- willtemperley 4mo agoI'm here to complain about the churn. I feel like I get to know a model in the human sense of understanding a personality. Yesterday I knew 4.6 extended, today it's different, there's multiple "token budget" levels. I just want 4.6 extended back as it was, I was getting on well with it / them.
- lionkor 4mo agoHumanizing this technology seems like a step in the wrong direction.
- willtemperley 4mo agoThere's so much intelligence here on HN and so little humanity.
- bwhiting2356 4mo agothe churn is... a version bump to the same api? If you want to compare you can write some evals.
- gandalfthepink 4mo agoMay be my tasks are rudimentary but the results I get with the 4.5 model are just the same as 4.7 or 4.6. it's just at the advanced models consume more tokens and and are actually loss making for my work. The incremental changes that they are making are not really that valuable. In fact I have found that even glm 5.1 is giving me something equivalent to what Opus 4.6 gives. Am I missing something that everyone else is cheering for in these small incremental model releases?
- andersmurphy 4mo agoI wonder if it's being done to improve revenue nunbers without changing an enterprise contract? Oh what's that your token usage went up because some of your developers switched to a new model? That sounds like a you problem. I thinks there's a big push to get these companies in a state where they can be dumped on public markets.
- ckarani 4mo ago[dead]
- christkv 4mo agoI'm going to assume that at some point their "targeted training and tuning" will eventually reach some sort of "max" possible simulation of next good token. At that point I think it will be interesting to see what happens and how many parameters you really need to for different verticals.
- bigupthewhole 4mo agoIve been using gpt 5.4 and 5.5 and honestly 5.4 is solving everything at the pace I need it. I'm the biggest bottle neck in terms of reviewing PRs and my own code. So having a model which can solve a complex task in 10 minutes vs 30 minutes doesn't really give me any meaningful improvement. Also, the biggest factor is having a good planning phase. A good plan is better than even major model improvements.
- pseudohadamard 4mo agoI have seen a noticeable difference between 4.6 Medium (the default, and I skipped 4.7 because of various reported issues) and 4.8 High or whatever the default is now. It's far more likely to say it doesn't know and seems to think about things a lot more, but then it also spends a lot more time reporting on what it's thought about so it takes longer for you to process the output. In particular 4.6 would say "I've spotted something a bit off here" whereas 4.8 will say "if you do this and then this and then this under these conditions then something will go wrong here". So it seems to be closer to the claimed capabilities for Mythos than previous versions.
- mik09 4mo ago[dead]