4 ms·
> OpenAI / Anthropic models have largely stopped advancing I'm shocked anyone could conclude this. This year it became common for people to entirely delegate c
by qnleigh 21d ago
> OpenAI / Anthropic models have largely stopped advancing
I'm shocked anyone could conclude this. This year it became common for people to entirely delegate coding to AI (I know many competent programmers/researchers who do this now). Progress in math has just been insane. An internal model at OAI just resolved one of the most celebrated open problems in mathematics. If anything, progress has accelerated.
- glub 21d ago> This year it became common for people to entirely delegate coding to AI This has been the case for around 2 years now, more reliably - a year. We've mostly stayed there since then. Saying that more people started doing it isn't indicative of significant improvement. Some people just started doing it later. I can't speak about math because I haven't used AI for that application, but I know that there hasn't been any significant advancement in coding in this year on base models. There has been more RL work, more harness work, more tools, they all expanded some capabilities like cyber or orchestration or tool use, but raw intelligence of base models is no longer where the main focus is.
- BobbyJo 20d ago> This has been the case for around 2 years now, more reliably - a year. I have to disagree with this pretty strongly. Opus 4.5 needed a lot of handholding not to work itself into a corner pretty quickly. Fable I basically never need to correct, and I've most become a data source.
- itkovian_ 20d agoWhat are you talking about - I feel like we’re living in parallel realities. If I had to go back to opus 4.5 tomorrow I’d be hugely upset and significantly slowed down
- bel8 21d agoI'm not. Yes we normalized 1m context window and models tend to hallucinate less. But models have been somewhat stagnant since Opus 4.6/7. And in some regards there were even regressions like Claudeisms that are load bearing.
- boshalfoshal 20d agoYes these guys are completely delusional. 2 years ago a model could barely solve the AMC, 1 year ago it got IMO gold, and this year models have solved multiple millenium problems. Even 1 year ago ai code was just unusable (claude code only became available 1.5 years ago!) and now basically everyone I know from independent shops all the way to faang and anthropic/openai themselves exclusively use some AI agent to code. Why does HN continue to delude itself that "models are not improving?" Maybe for the simple things they care about its "roughly the same," but they are _clearly_ improving.
- qnleigh 20d agoWaitwaitwait multiple millennium problems? What was the other one??? Increasing a bound on the fraction of Reimann zeros on the critical line doesn't count; even Anthropic says they don't think this line of work will lead to a solution.
- boshalfoshal 20d agoNavier Stokes, and _allegedly_ the Hodge Conjecture and Birch–Swinnerton-Dyer.