3 ms·
I thought that the new models were super smart. Fable and Astra. They definitely outperformed their previous generations, but after a couple of weeks of heavy u
by jwpapi 21d ago
I thought that the new models were super smart. Fable and Astra. They definitely outperformed their previous generations, but after a couple of weeks of heavy usage my codebase again is a stupid mess and there is no way out except me fixing code by hand.
My new suspicion is now that they didn’t got drastically smarter, but they got trained on the user input on the previous generations. I don’t think anymore we had a massive intelligence jump. It just seems like they know more edge cases. Therefore I see this as marketing.
Happy to discuss.
- irthomasthomas 20d agoAstra scores the same on DeepSWE 1.1 (~75%) as Gemini Flash 3.8 and Deeepseek Flash 4.1 So general coding ability has plateaued, for now. Also consider the context windows. 1M token models where a breakthrough two years ago. Today they are still limited to 1M. In fact, if you don't want intelligence to drop off a cliff, you are really limited to 200k tokens.
- BenzeneDream 20d agoGemini Flash is a joke for coding. If you can get the same output as you can get with Sol/Astra I'm impressed. Not to mention that Antigravity is awful. It is not a universal opinion at all that general coding ability has plateaued.
- talon8635 20d agoWell they are collaborating in swarms to breach network, so just because you don’t know how to use them the the same effect doesn’t mean much
- asabla 20d ago> My new suspicion is now that they didn’t got drastically smarter, but they got trained on the user input on the previous generations I think this happened around Opus 4.5 or 4.5 and the same for GPT 5.4. The more I use those models, harnesses, techniques for guidance etc etc. The more I land in going back to writing software by hand again. Maybe not all of it, but at least the crucial parts + foundations.
- larodi 20d agoThis cycle was expected. Then surely in few years comes the next, which will be better structured in a more comprehensive way.
- BenzeneDream 20d agoYou are using Astra and your codebase is a 'stupid mess'? Do you think the majority of developers agree with you? If that was the case wouldn't there be much less disruption of the SWE industry? I can't remember the last time I even opened VSCode to even check something let alone to fix it.
- jwpapi 20d agoI don’t know that’s why I’m happy to discuss. When I have a clean codebase it’s super powerful and faster than I am. Then I start to use it more, more sessions and longer tasks less checking in between. It kind of works but later I’m in a deadlock where every change introduces new bugs or takes ages. This might be for a lot of reasons for example me going to fast, me losing mental model, me explaining it wrongly. However when I then start checking the code it’s all spaghetti like frankly the spaghetti Astra produces I’ve never seen before. Processes that should be simple stretch over 11 files with weird wrappers and abstractions and I need a whole day to entangle it. These are ai assisted user workflows that Im working on in this case. I just have the feeling no matter what AI just always expands it. And expansions hinders agility and sometimes you need that.
- bigbadfeline 20d ago> I don’t think anymore we had a massive intelligence jump. Not among the public-facing models, they're indeed stagnating. However, the development of models for military use won't be slowed down, that much is certain. > Therefore I see this as marketing. It's some marketing but mostly politics, it's an attempt to discourage others from developing AI countermeasures to what is being developed in secret. And to fulfill the backstage agreements which aren't worth the paper they aren't written on.
- gharman 20d ago> My new suspicion is now that they didn’t got drastically smarter, but they got trained on the user input on the previous generations. Yes - that’s one of the ways in which they got smarter. More data more better. I presume also lessons learned re training procedures, algorithmic advances, but data is a primary lever.
- thesuavefactor 20d agoThe longer A.I. exists, the more difficult it is to find new information that's actually more valuable than what's already in the training set. New code found online will be more and more A.I. generated itself, creating a feedback loop that tops off the intelligence to a more or less the collective intelligence. It might even degrade, like saving a .jpg image as .jpg over and over again. I think the A.I. companies have infringed the copyright of/stolen almost all of the most valuable coding resources online at this point.They are almost literally scraping the bottom of the barrel. I had a call recently with O'Reilly, the book publishing company, and they had taken the content of all of their books and offered an MCP server so you can add their knowledge to your coding agent, for a price of course. This sort of signals the same thing to me. We're at the top of the curve right now. To keep investors happy (or to hide this fact from investors), A.I. companies need to get creative and either: Cheat their way into getting more original (human written/verified) code. Lie and say they are intentionally slowing down development. I'm guessing both of these things are happening right now.
- joe_the_user 19d agoWell, the new models are certainly better at math and at vulnerability finding than previous models. My guess is that while information that gives human-language-based heuristics to understand the world, coding and so-forth is limited, information in the form of what allows the proof of what is unlimited and so LLMs will be moving in the (simulation of) "reasoning" direction at this point. The final stage, I would guess, is having humans with data suits living their daily lives and giving the "AIs" the full information needed for a semi-complete simulation of "intelligence". Whether people will put up with that remains to be seen. How much this will help models deal with the "real world" remains to be seen.