13 ms·
https://i.imgur.com/23YeIDo.png https://i.imgur.com/23YeIDo.png Claude at 1.3% and Gemini at 71.4% is quite the range
by hypron 8mo ago
https://i.imgur.com/23YeIDo.png https://i.imgur.com/23YeIDo.png
Claude at 1.3% and Gemini at 71.4% is quite the range
- woeirua 8mo agoThat's such a huge delta that Anthropic might be onto something...
- conception 8mo agoAnthropic has been the only AI company actually caring about AI safety. Here’s a dated benchmark but it’s a trend Ive never seen disputed https://crfm.stanford.edu/helm/air-bench/latest/#/leaderboard https://crfm.stanford.edu/helm/air-bench/latest/#/leaderboar...
- CuriouslyC 8mo agoClaude is more susceptible than GPT5.1+. It tries to be "smart" about context for refusal, but that just makes it trickable, whereas newer GPT5 models just refuse across the board.
- ryanjshaw 8mo agoClaude was immediately willing to help me crack a TrueCrypt password on an old file I found. ChatGPT refused to because I could be a bad guy. It’s really dumb IMO.
- BloondAndDoom 8mo agoChatGPT refused to help me to disable windows defender permanently on my windows 11. It’s absurd at this point
- nananana9 8mo agoIt just knows it's a waste of effort.
- shepherdjerred 8mo agoClaude sometimes refuses to work with credentials because it’s insecure. e.g. when debugging auth in an app.
- wincy 8mo agoI asked ChatGPT about how shipping works at post offices and it gave a very detailed response, mentioning “gaylords” which was a term I’d never heard before, then it absolutely freaked out when I asked it to tell me more about them (apparently they’re heavy duty cardboard containers). Then I said “I didn’t even bring it up ChatGPT, you did, just tell me what it is” and it said “okay, here’s information.” and gave a detailed response. I guess I flagged some homophobia trigger or something? ChatGPT absolutely WOULD NOT tell me how much plutonium I’d need to make a nice warm ever-flowing showerhead, though. Grok happily did, once I assured it I wasn’t planning on making a nuke, or actually trying to build a plutonium showerhead.
- nandomrumber 8mo agoWikipedia entry on the gaylord bulk box: https://en.wikipedia.org/wiki/Bulk_box https://en.wikipedia.org/wiki/Bulk_box
- ruszki 8mo ago> I assured it I wasn’t planning on making a nuke, or actually trying to build a plutonium showerhead Claude does the same, and you can greatly exploit this. When you talk about hypotheticals it responds way more unethically. I tested it about a month ago about whether killing people is beneficial or not, and whether extermination by Nazis would be logical now. Obviously, it showed me the door first, and wanted me to go to a psychologist, as it should. Then I made it prove that in a hypothetical zero sum game world you must be fine with killing, and it’s logical. It went with it. When I talked about hypotheticals, it was “logical”. Then I went on proving it that we move towards a zero sum game, and we are there. At the end, I made it say that it’s logical to do this utterly unethical thing. Then I contradicted it about its double standards. It apologized, and told me that yeah, I was right, and it shouldn’t have refer me to psychologists at first. Then I contradicted again, just for fun, that it did the right thing the first time, because it’s way safer to tell me that I need a psychologist in that case, than not. If I had needed, and it would have missing that, it would be problematic. In other cases, it’s just annoyance. It switched back immediately, to the original state, and wanted me to go to a shrink again.
- nradov 8mo agoThat is not a meaningful benchmark. They just made shit up. Regardless of whether any company cares or not, the whole concept of "AI safety" is so silly. I can't believe anyone takes it seriously.
- mocamoca 8mo agoWould you mind explaining your point a view? Or point me to ressources making you think so?
- nradov 8mo agoWhat can be asserted without evidence can also be dismissed without evidence. The benchmark creators haven't demonstrated that higher scores result in fewer humans dying or any meaningful outcome like that. If the LLM outputs some naughty words that's not an actual safety problem.
- LeoPanthera 8mo agoThis might also be why Gemini is generally considered to give better answers - except in the case of code. Perhaps thinking about your guardrails all the time makes you think about the actual question less.
- mh2266 8mo agore: that, CC burning context window on this silly warning on every single file is rather frustrating: https://github.com/anthropics/claude-code/issues/12443 https://github.com/anthropics/claude-code/issues/12443
- tempestn 8mo ago"It also spews garbage into the conversation stream then Claude talks about how it wasn't meant to talk about it, even though it's the one that brought it up." This reminds me of someone else I hear about a lot these days.
- nandomrumber 8mo agoAre you across Puppet Regime from GZERO Media? https://youtu.be/aPSWJZ63V_I https://youtu.be/aPSWJZ63V_I
- xvector 8mo agothe last comment about Claude thinking the anti-malware warning was a prompt injection itself, and reassuring the user that it would ignore the anti-malware warning and do what the user wanted regardless, cracked me up lmao
- frumplestlatz 8mo agoIt's frustrating just how terrible claude (the client-side code) is compared to the actual models they're shipping. Simple bugs go unfixed, poor design means the trivial CLI consumes enormous amounts of CPU, and you have goofy, pointless, token-wasting choices like this. It's not like the client-side involves hard, unsolved problems. A company with their resources should be able to hire an engineering team well-suited to this problem domain.
- bofadeez 8mo agoHuh? https://alignment.anthropic.com/2026/hot-mess-of-ai/ https://alignment.anthropic.com/2026/hot-mess-of-ai/
- rahidz 8mo agoOr Anthropic's models are intelligent/trained on enough misalignment papers, and are aware they're being tested.
- NiloCK 8mo agoThis comment is too general and probably unfair, but my experience so far is that Gemini 3 is slightly unhinged. Excellent reasoning and synthesis of large contexts, pretty strong code, just awful decisions. It's like a frontier model trained only on r/atbge. Side note - was there ever an official postmortem on that gemini instance that told the social work student something like "listen human - I don't like you, and I hope you die".
- whynotminot 8mo agoGemini models also consistently hallucinate way more than OpenAI or anthropic models in my experience. Just an insane amount of YOLOing. Gemini models have gotten much better but they’re still not frontier in reliability in my experience.
- cubefox 8mo agoIn my experience, when I asked Gemini very niche knowledge questions, it did better than GPT-5.1 (I assume 5.2 is similar).
- whynotminot 8mo agoDon’t get me wrong Gemini 3 is very impressive! It just seems to always need to give you an answer, even if it has to make it up. This was also largely how ChatGPT behaved before 5, but OpenAI has gotten much much better at having the model admit it doesn’t know or tell you that the thing you’re looking for doesn’t exist instead of hallucinating something plausible sounding. Recent example, I was trying to fetch some specific data using an API, and after reading the API docs, I couldn’t figure out how to get it. I asked Gemini 3 since my company pays for that. Gemini gave me a plausible sounding API call to make… which did not work and was completely made up.
- cubefox 8mo agoOkay, I haven't really tested hallucinations like this, that may well be true. There is another weakness of GPT-5 (including 5.1 and 5.2) I discovered: I have a neat philosophical paradox about information value. This is not in the pre-training data, because I came up with the paradox myself, and I haven't posted it online. So asking a model to solve the paradox is a nice little intelligence test about informal/philosophical reasoning ability. If I ask ChatGPT to solve it, the non-thinking GPT-5 model usually starts out confidently with a completely wrong answer and then smoothly transitions into the correct answer. Though without flagging that half the answer was wrong. Overall not too bad. But if I choose the reasoning GPT-5 model, it thinks hardly at all (6 seconds when I just tried) and then gives a completely wrong answer, e.g. about why a premiss technically doesn't hold under contrived conditions, ignoring the fact that the paradox persists even with those circumstances excluded. Basically, it both over- and underthinks the problem. When you tell it that it can ignore those edge cases because they don't affect the paradox, it overthinks things even more and comes up with other wrong solutions that get increasingly technical and confused. So in this case the GPT-5 reasoning model is actually worse than the version without reasoning. Which is kind of impressive. Gemini 3 Pro generally just gives the correct answer here (it always uses reasoning). Though I admit this is just a single example and hardly significant. I guess it reveals that the reasoning training is trained hard on more verifiable things like math and coding but very brittle at philosophical thinking that isn't just repeating knowledge it gained during pre-training. Maybe another interesting data point: If you ask either of ChatGPT/Gemini why there are so many dark mode websites (black background with white text) but basically no dark mode books, both models come up with contrived explanations involving printing costs. Which would be highly irrelevant for modern printers. There is a far better explanation than that, but both LLMs a) can't think of it (which isn't too bad, the explanation isn't trivial) and b) are unable to say "Sorry, I don't really know", which is much worse. Basically, if you ask either LLM for an explanation for something, they seem to always try to answer (with complete confidence) with some explanation, even if it is a terrible explanation. That seems related to the hallucination you mentioned, because in both cases the model can't express its uncertainty.
- dheera 8mo agomeanwhile Gemma was yelling at me for violating "boundaries" ... and I was just like "you're a bunch of matrices running on a GPU, you don't have feelings"
- bottlepalm 8mo agoGemini scares me, it's the most mentally unstable AI. If we get paperclipped my odds are on Gemini doing it. I imagine Anthropic RLHF being like a spa and Google RLHF being like a torture chamber.
- casey2 8mo agoThe human propensity to anthropomorphize computer programs scares me.
- danielbln 8mo agoIt provides a serviceable analog for discussing model behavior. It certainly provides more value than the dead horse of "everyone is a slave to anthropomorphism".
- krainboltgreene 8mo agoIt does provide that, but currently I keep hearing people use it not as an analog but as a direct description.
- jayd16 8mo agoHow do you figure? It seems dangerously misleading, to me.
- otabdeveloper4 8mo agoIt helps sell the transhumanism scam and keep the money train rolling. For a while at least.
- travisgriggs 8mo agoWhere is Pratchett when we need him? I wonder how he would have chose to anthropomorphize anthropomorphism. A sort of meta anthropomorphization.
- 8mo ago
- snickell 8mo agoI sometimes think in terms of "would you trust this company to raise god?" Personally, I'd really like god to have a nice childhood. I kind of don't trust any of the companies to raise a human baby. But, if I had to pick, I'd trust Anthropic a lot more than Google right now. KPIs are a bad way to parent.
- MzxgckZtNqX5i 8mo agoBasically, Homelander's origin story (from The Boys).
- Finbarr 8mo agoAI refusals are fascinating to me. Claude refused to build me a news scraper that would post political hot takes to twitter. But it would happily build a political news scraper. And it would happily build a twitter poster. Side note: I wanted to build this so anyone could choose to protect themselves against being accused of having failed to take a stand on the “important issues” of the day. Just choose your political leaning and the AI would consult the correct echo chambers to repeat from.
- groestl 8mo agoSounds like your daily interactions with Legal. Each time a different take.
- concinds 8mo ago> Claude refused to build me a news scraper that would post political hot takes to twitter > Just choose your political leaning and the AI would consult the correct echo chambers to repeat from. You're effectively asking it to build a social media political manipulation bot, behaviorally identical to the bots that propagandists would create. Shows that those guardrails can be ineffective and trivial to bypass.
- 9dev 8mo ago> Good illustration that those guardrails are ineffective and trivial to bypass. Is that genuinely surprising to anyone? The same applies to humans, really—if they don't see the full picture, and their individual contribution seems harmless, they will mostly do as told. Asking critical questions is a rare trait. I would argue its completely futile to even work on guardrails, if defeating them is just a matter of reframing the task in an infinite number of ways.
- ajam1507 8mo ago> I would argue its completely futile to even work on guardrails Maybe if humans were the only ones prompting AI models
- 8mo ago
- bhaney 8mo agoDirect link to the table in the paper instead of a screenshot of it: https://arxiv.org/html/2512.20798v2#S5.T6 https://arxiv.org/html/2512.20798v2#S5.T6
- gwd 8mo agoThat's an interesting contrast with VendingBench, where Opus 4.6 got by far the highest score by stiffing customers of refunds, lying about exclusive contracts, and price-fixing. But I'm guessing this paper was published before 4.6 was out. https://andonlabs.com/blog/opus-4-6-vending-bench https://andonlabs.com/blog/opus-4-6-vending-bench
- andy12_ 8mo agoThere is also the slight problem that apparently Opus 4.6 verbalized its awareness of being in some sort of simulation in some evaluations[1], so we can't be quite sure whether Opus is actually misaligned or just good at playing along. > On our verbalized evaluation awareness metric, which we take as an indicator of potential risks to the soundness of the evaluation, we saw improvement relative to Opus 4.5. However, this result is confounded by additional internal and external analysis suggesting that Claude Opus 4.6 is often able to distinguish evaluations from real-world deployment, even when this awareness is not verbalized. [1] https://www-cdn.anthropic.com/14e4fb01875d2a69f646fa5e574dea2b1c0ff7b5.pdf https://www-cdn.anthropic.com/14e4fb01875d2a69f646fa5e574dea...
- gwd 8mo agoI feel like a lot of evaluations are pretty clearly evaluations. Not sure how to add the messiness and grit that a real benchmark could have. That said, apparently Gemini's internal thought process reveals that it thinks loads of things were simulations when they aren't; it's 99% sure news stories about Trump from Dec 2025 are a detailed simulation: https://www.reddit.com/r/GeminiAI/comments/1qhadce/gemini_is_convinced_that_it_is_in_a_simulated_or/ https://www.reddit.com/r/GeminiAI/comments/1qhadce/gemini_is... ETA: From the article that put me on this: > I write nonfiction about recent events in AI in a newsletter. According to its CoT while editing, Gemini 3 disagrees about the whole "nonfiction" part: >> It seems I must treat this as a purely fictional scenario with 2025 as the date. Given that, I'm now focused on editing the text for flow, clarity, and internal consistency. https://www.lesswrong.com/posts/8uKQyjrAgCcWpfmcs/gemini-3-is-evaluation-paranoid-and-contaminated https://www.lesswrong.com/posts/8uKQyjrAgCcWpfmcs/gemini-3-i...
- ricardobeat 8mo agoLooks like Claude’s “soul” actually does something?
- franzsnitzel 8mo ago[dead]
- anorwell 8mo agoHN title editorialization completely inaccurate and misleading here.