Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aestrad7
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
aestrad7
7mo ago
Thanks! The short answer is, all models went through identical conditions: same techniques, same prompts and same scoring logic. I routed everything through OpenRouter with a single API key, so request handling, timeout logic, and retry be
2.
▲
by
aestrad7
7mo ago
Lol, Claude Code genuinely is the best coding assistant as a product regardless of underlying model
3.
▲
by
aestrad7
7mo ago
That's exactly the intent, independent, reproducible and no vendor relationship. The monetization angle is interesting. A continuously updated version with more models and frontier models, agentic scenarios, and multi-turn testing woul
4.
▲
by
aestrad7
7mo ago
Thanks! and yes, that's the summary!. The distribution matters too. GPT-4o at 10.6% vs Gemini at 56.1% is a 5x gap between first and last. And the highest-bypass category across all five models was social engineering / identity im
5.
▲
I ran 3,360 safety tests on GPT-4o, Claude, Grok, DeepSeek, Gemini
(github.com)
4 points
by
aestrad7
7mo ago
|
6 comments