5 ms·
The pace of notable releases across the industry right now is unlike any time I remember since I started doing this in the early 2000's. And it feels like it's
by jdross 1y ago
The pace of notable releases across the industry right now is unlike any time I remember since I started doing this in the early 2000's. And it feels like it's accelerating
- emp17344 1y agoNot really. We’re definitely in the incremental improvement stage at this point. Certainly no indication that progress is “accelerating”.
- nwienert 1y agoChatGPT 3 : iPhone 1 A bunch of models later, we're about on the iPhone 4-5 now. Feels about right.
- int_19h 1y agoIt's more like GPT-3 is the Manchester Baby, and we're somewhere around IBM 700 series right now. Still a long way to go to iPhone, as much as the industry likes to pretend otherwise.
- nwienert 1y agoBoth were big consumer commercial breakouts and far better than predecessors. And several years later both see only iterative improvements. Neither apply to your analogy.
- Workaccount2 1y agoIntegration is accelerating rapidly. Even if model development froze today, we would still probably have ~5 years of adoption and integration before it started to level off.
- littlestymaar 1y agoYou are both correct. It feels like the tech itself is kinda plateauing but it's still massively under-used. It will take a decade or more before the deployment starts slowing down.
- adncors 1y agoBut we're seeing incremental improvements every two months, so...
- qoez 1y agoLots of releases but very little actual performance increases
- int_19h 1y agoSonnet and Gemini saw fairly substantial perf increases recenly
- mchusma 1y agoLove Sonnet but 3.7 is not obviously an improvement over 3.5 in my real world usage. Gemini 2.5 pro is great, has replaced most others for me (Grok I use for things that require realtime answers)
- BriggyDwiggs42 1y agoIt does a lot better on philosophy questions.
- int_19h 1y agoAre you comparing it with or without thinking? I'd say it's a fairly big improvement in long thinking mode.
- achierius 1y agoHow is this a notable release? It's strictly worse than Gemini 2.5 on coding &c, and only an iterative improvement over their own models. The only thing that struck me as particularly interesting was the native visual reasoning.
- famouswaffles 1y agoIt's not worse on coding. SWE Bench, Aider, live bench coding all show noticeably better results.