2 ms·
It's a great benchmark. Don't listen to the haters. This one is especially interesting. https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-medium/qwen-qw
by jaggs 6mo ago
It's a great benchmark. Don't listen to the haters. This one is especially interesting.
https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-medium/qwen-qwen3-6-plus-medium/?order=qwen-qwen3-6-plus-medium%2Canthropic-claude-sonnet-4-6-medium https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...
- BoorishBears 6mo agoThis one's even more interesting https://aibenchy.com/compare/anthropic-claude-opus-4-6-medium/google-gemini-2-5-flash-medium/?order=google-gemini-2-5-flash-medium%2Canthropic-claude-opus-4-6-medium https://aibenchy.com/compare/anthropic-claude-opus-4-6-mediu... Who knew Anthropic was this far behind???
- jaggs 6mo agoYeah, but actually that's not a good look. Anyone who's used Gemini will know how random it is in terms of getting anything serious done, compared to the rock solid opus experience.
- BoorishBears 6mo agoTheir benchmark is chock-full of things like that: It's deeply flawed and is essentially rating how LLMs perform if you exert yourself trying to hold them entirely the wrong way.