3 ms·
6 months ago: get Claude to do some work, have GPT review it, ask Claude to first verify findings before actioning. Now: get GPT to do some work, have Claude r
by ben8bit 2mo ago
6 months ago: get Claude to do some work, have GPT review it, ask Claude to first verify findings before actioning.
Now: get GPT to do some work, have Claude review it, question Claude about a finding that is surprising to me because I thought the functionality was already in place.
[Claude/Opus 5 Max goes looking] "You're right — I was wrong about that."
Our Claude license ends in about 2 weeks and we're not renewing - this has been par for the course for the last 2 months now.
And their marketing is really starting to bother me, on top of that.
- ValentineC 2mo ago> [Claude/Opus 5 Max goes looking] The general consensus I've seen for Opus 5 on the various Claude subreddits is to only use it in low or medium effort. I've reverted to Opus 4.8 for most of my work. It's crazy that Opus 5 was lauded at release for scoring so highly on benchmarks.
- this_user 2mo agoThat always seems to happen upon release of a new model. People look at the benchmarks, which look good, because the model was almost certainly optimised for that. Then they start actually using it, and after a week or two we get the real assessment that is almost always less impressive than the initial reactions.
- ben8bit 2mo agoInteresting, I generally use High for all models, even non Anthropic ones. This was one of the few times I tried Max, and the problem was the code (it listed as a "hole") was directly adjacent to the problem area, and not especially complex. Kind of like looking at a washing machine and telling the customer to be careful because the inlet pipe will pump water into an empty box.
- ValentineC 2mo agoI believe the reasoning for not using anything higher than medium is that Opus 5 has a tendency to overthink at higher effort levels. So I guess the washing machine analogy is that you've somehow added too much detergent because the new formula is 3x as strong, and your clothes are very clean, but the fibres have also degraded leaving your clothes a bit threadbare.
- kingkongjaffa 2mo agoGPT5.6-Sol is brilliant on Max in my experience.
- ben8bit 2mo agoSol is absolutely incredible. It's the first model where it feels like a mid-weight engineer that actually looks at the details. I haven't tried Max yet - High for me has been enough so far.
- kingkongjaffa 2mo agoAs a day 1 claude code user I've been a Claude fan boy but Opus and Fable usage and pricing is just not viable. I cancelled my personal claude account and got chatGPT $20 - the amount of usage is such that I've never run out ever, which was a common occurrence with Claude. Have not missed it at all for random household use and a bit of Codex as well. Still use Claude Code for $dayjob provided by my company.
- DaedalusII 2mo agothe biggest problem with claude is it is so negative and tells you things can't be done, are wrong, have hidden problems, etc I run things through grok to make them positive again after claude does the main work
- ben8bit 2mo agoGrok 4.5 is actually good though (as a workhorse model).