2 ms·
All Claude models are huge suck ups. The "you're absolutely right" meme is real even if that exact phrase doesn't show up as much anymore. I don't want to star
by resonious 4mo ago
All Claude models are huge suck ups. The "you're absolutely right" meme is real even if that exact phrase doesn't show up as much anymore.
I don't want to start a fight or anything but IME Codex has a bit more of a spine. If you point out something weird, it sometimes gives a good reason for it. Whereas Claude will always say "whoopsie you're right as always sir" even when it's me who missed something.
- teaearlgraycold 4mo agoRight now the thing I get from Opus 4.8 is a ton of “That’s a good instinct”. Also >50% of its closing statements begin with “Clean.”
- herdymerzbow 4mo agoI only use free AI chats to help me with my learning, but often I direct its responses neutral and to refrain from providing any encouraging language, or value judgements. It tends to get rid of these 'you're absolutely right' comments when I point out a mistake. But your comment just made me think whether this tendency for LLMs to resort to flattery when found out is a built in strategy to distract the user from the error prone fragility of much of the output? It's perhaps a stretch to think these canned responses were put in strategically, but the result is that the user's attention may be deflected to contemplating their own superior knowledge and insight, and bask in the glory of all that, but then forgot to appreciate that 'Hey, chatLLM is just making all this stuff up/doesn't know which way is up/or down!'
- pyridines 4mo agoIME it's Claude that pushes back, and Codex that just does the thing. It's happened once or twice where I've told Claude bluntly and directly "do this" and it responded "no, here's why that's a bad idea..." Maybe it's just my CLAUDE.md. Not sure if there are sycophancy benchmarks for coding agents
- mcintyre1994 4mo agoI find the same. Someone posted this benchmark here: https://petergpt.github.io/bullshit-benchmark/viewer/index.v2.html https://petergpt.github.io/bullshit-benchmark/viewer/index.v... It measures whether models push back on bullshit prompts or just go along with it, and Claude models are all the top performers.
- QuiEgo 4mo agoI’ve had this experience as well. I love Codex for doing code reviews, it takes a way more direct, less passive tone when calling out issues.