3 ms·
> late last year/early this year That's an eternity when it comes to coding models. In my personal experience, we've had almost a step change every ~3 months
by ehsankia 17d ago
> late last year/early this year
That's an eternity when it comes to coding models.
In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.
- pelagicAustral 17d agotbf, I the happiest I've been working with claude is late last year/early this year (before March)...
- epolanski 17d agoThat's because Opus 4.6 was the last good assistant model. Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant. Now it's *you* being the assistant, reviewer, etc.
- pelagicAustral 17d agoIs that right? Why the shit would that happen?
- transdev12 17d agoBecause they’ve essentially exhausted pre training scaling and are looking to post training to expand capabilities, which is really just optimization via reinforcement learning against specific tasks aka bench maxing.
- ctolsen 17d agoTheir ambition isn't your work being amplified by their model, they want you running fifty autonomous long-running agents.
- epolanski 17d agoIn one sense you're right. In another one, Opus 4.6 level already solved 90% of my work-day tasks, so while better models have been instrumental into handling a higher % that does not mean that defaulting on cheaper models can't be good. I run DS 4.1 flash daily, and then cross check with gpt-6-astra and I've nuked 90% of my AI monthly bill while having higher limits and better performance/intelligence than I did just at the beginning of this summer.