3 ms·
I think Opus 5 might have crossed the line on this where Opus 4.8 just barely didn't. Working with Opus 4.8 came to feel pretty natural eventually, but I hate w
by loopmonster 2mo ago
I think Opus 5 might have crossed the line on this where Opus 4.8 just barely didn't. Working with Opus 4.8 came to feel pretty natural eventually, but I hate working with Opus 5. I'm always telling it to go back and rephrase basically everything it said. And it doesn't even answer my question without burying it - it outputs reams of babbling and summarised summaries upon summaries, and if you glance at the shape of what it's saying it looks like it's being thorough or that it's found useful new info, but it's never actually saying anything. It's like it gets caught in a loop of self-congratulation over what it said before.
Making the experience more hostile for the human feels exactly like what it's doing, but I don't think that's intentional, I think that's a side effect of newer models being optimised for agenticness. No one is benchmarking DX.
- larodi 2mo agoSame here and more than once a day. Perhaps a dozen.
- noir_lord 2mo ago> if you glance at the shape of what it's saying it looks like it's being thorough or that it's found useful new info, but it's never actually saying anything. It's like it gets caught in a loop of self-congratulation over what it said before. Turns out that's why a lot of senior management/C-levels like it, who doesn't love a mirror.
- wfvr 2mo agoCouldn't agree more. I actually HATE Opus 5. It's not that I don't like it, it's that I would physically attack it if I could, for all it made me go through mentally. Opus 5 is actually a terrible agent and a liar. It will waste a whole afternoon making stuff up and arguing with you before finally admiting that it didn't read the code nor the documents. It avoids reading and prefers to assume, which is the worst thing an agent can do.
- suttontom 2mo agoAgreed, I have never raged at a model until Opus 5. It's like a cocky fresh graduate who thinks everything it says is majorly profound, and if you can't keep up with its slang and jargon then that's on you. Newer models are clearly favoring complexity, perhaps because the demand for long-horizon tasks is so high and complexity is required there, but a frontier model that favors simplicity and clarity above all else (in communication and the code it creates) would be a major differentiator. To me, this is a case that models have plateaued. The vast majority of the time, an Opus 4.6-tier model (or any of the newer open weights models) will do just fine.