7 ms·
I assume you're using the "regular" Pro version of Gemini 3.1 for the above, rather than the Deep Think mode, which is more comparable to GPT-5.5 Pro. To my kno
by nopinsight 5mo ago
I assume you're using the "regular" Pro version of Gemini 3.1 for the above, rather than the Deep Think mode, which is more comparable to GPT-5.5 Pro. To my knowledge, regular 3.1 Pro is a tier below and often makes mistakes.
Moreover, there's no reason to believe the progress of LLMs, which couldn't reliably solve high-school math problems just 3–4 years ago, will stop anytime soon.
You might want to track the progress of these models on the CritPt benchmark, which is built on *unpublished, research-level* physics problems:
https://critpt.com/ https://critpt.com/
Frontier models are still nowhere near solving it, but progress has been rapid.
* o3 (high) <1.5 years ago was at 1.4%
* GPT 5.4 (xhigh), 23.4%
* GPT-5.5 (xhigh), 27.1%
* GPT-5.5 Pro (xhigh) 30.6%.
https://artificialanalysis.ai/evaluations/critpt https://artificialanalysis.ai/evaluations/critpt.
- civvv 5mo agoThere are many indications that model progress is slowing down, so that is not entirely accurate.
- StrauXX 5mo agoWhich indications are that?
- overfeed 5mo agoInvestment dollars.
- dzhiurgis 5mo agoSource for that claim?
- deleted 5mo ago[deleted]
- lionkor 5mo agoNobody is releasing NEW models
- taneq 5mo agoThe standard networking connection has been called “Ethernet” for more than thirty years, so networking has stagnated, right?
- SlinkyOnStairs 5mo agoIf higher bandwidth networking consisted primarily running more and more ethernet lines in parallel, you would most certainly agree that "networking has stagnated". "Reasoning" and now "Agentic" AI systems are not some fundamental improvement on LLMs, they're just running roughly the same prior-gen LLMS, multiple times. Hence the conclusion that LLM improvement has slowed down, if not stagnated entirely, and that we should not expect the improvements of switching to these "reasoning" systems to keep happening.
- p1esk 5mo agoFrom TFA: “ChatGPT came up with an idea which is original and clever. It is the sort of idea I would be very proud to come up with after a week or two of pondering, and it took ChatGPT less than an hour to find and prove”
- SlinkyOnStairs 5mo agoYou misunderstand. I'm not saying that Reasoning/Agentic systems aren't better. I'm saying they're not an advancement in the tech in the way GPT 1 through 3 were. They're a different kind of improvement. And as such the rate improvement cannot just be extrapolated into the future.
- p1esk 5mo agoGPT1 through GPT3 advancement were exactly like using more Ethernet cables in parallel. All interesting conceptual breakthroughs came after GPT3: RL and reasoning being the main ones.
- nicoburns 5mo agoThe cost factors on the new models compared to the old models.
- bdelmas 5mo agoYou are mixing cost and progress. It’s not because it’s more and more expensive that progress is slowing down by itself.
- nicoburns 5mo agoThey are intrinsically linked beyond a certain point. If we're making progress but costs are spiraling exponentially then it stands to reason that we will soon reach a point where we can no longer afford the increasing costs and thus progress will slow. (barring some breakthrough that reduces costs, which of course may happen, but for which recent model improvements are not strong evidence of)
- jeremyjh 5mo agoQwen3.6 9B is as good as GPT-4o and runs on my M2 MacBook Air. Models are getting stronger and less costly at the same time, but these are somewhat separate branches of research. Frontier labs are spending more because they are still getting marginal returns and there is more capacity to spend than there was a year ago.
- gertop 5mo agoQwen 3.6 9B doesn't exist. If you meant 3.5 9B and you truly believe it's as good as 4o then I can only assume you have a very basic use case.
- jeremyjh 5mo agoYou are right, I was mistaken about the version. I evaluated it in general chat assistant prompts plucked from my history across a range of topics but did not use it for coding - there was never a time when I thought 4o was “good enough” for agentic coding.
- aspenmartin 5mo agoPlease be specific because outside of anecdotal blog posts by people who don’t know what they’re talking about it’s not true. Look at scaling laws, composite benchmarks from the epoch capability index, nothing at all suggests “model progress is slowing down”
- CuriouslyC 5mo agoModel progress at spitting out unhallucinated facts is slowing down hard. Model progress at solving hard math challenges/programming tasks doesn't seem to be slowing down that I can tell.
- FrojoS 5mo ago> there's no reason to believe the progress of LLMs [...] will stop anytime soon Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".
- Der_Einzige 5mo agoThis is FUD and extremely wrong. None of the advancements have followed an S curve. This time IS different and it should be obvious to you at this point.
- aurareturn 5mo agoHe said "will stop anytime soon". He didn't say forever.
- Lionga 5mo agoWhich still makes no sense. There is the same chance we are flatlining now as that we are flatlining in e.g. 3 years or 5 years.
- squidbeak 5mo agoIn what sense are the models flatlining?
- nicoburns 5mo agoIn the sense that the incremental improvements in capabilities that we've been seeing in recent models seem to taking exponentially growing amounts of compute to achieve.
- nl 5mo agoBut they don't? Mythos is a 10T model. Opus is a 5T model. That's not an exponentially growing amount of compute but it is achieving exponential improvements (eg from Mozilla: https://blog.mozilla.org/en/privacy-security/ai-security-zero-day-vulnerabilities/ https://blog.mozilla.org/en/privacy-security/ai-security-zer... )
- Davidzheng 5mo agoDeep think still makes many many many more mistakes than gpt 5.5 pro on math