4 ms·
How does it compare to Kimi 2.5 or Qwen 3.6 Plus?
by jaggs 6mo ago
How does it compare to Kimi 2.5 or Qwen 3.6 Plus?
- DeathArrow 6mo agoCompared to Kimi 2.5 or Qwen 3.6 Plus I don't know, but I ran GLM 5 (not 5.1) side by side with Qwen 3.5 Plus and it was visibly better.
- eis 6mo agoThe blog post has a benchmark comparison table with these two in it
- jaggs 6mo agoThanks, I missed that. It's very interesting. They're quite close, but I found Qwen 3.6 plus was just marginally better than Kimi 2.5. But looking at the stats I'll definitely give GLM 5.1 a try now. [edit: even though looking at it, it's not cheap and has a much smaller context size.And I can't tell about tool use.]
- XCSme 6mo agoGeneral intelligence (not coding) comparison: https://aibenchy.com/compare/z-ai-glm-5-medium/z-ai-glm-5-1-medium/moonshotai-kimi-k2-5-medium/qwen-qwen3-6-plus-preview-medium/ https://aibenchy.com/compare/z-ai-glm-5-medium/z-ai-glm-5-1-...
- BoorishBears 6mo agoIs there really no rule that discourages 99% of your interactions with HN from being peddling some useless slop benchmark?
- XCSme 6mo agoIf it's relevant to the discussion, I hope not. I've spent probably over100 hours working on this benchmarking/site platform, and all tests are manually written. For me (and many others that reached out to me) are not useless either. I use this myself regularly when choosing and comparing new models. I honestly beleive it is providing value to the conversation. Let me know if you know of a better platform you can use to compare models, I built this one because I didn't find any with good enough UX.
- jaggs 6mo agoIt's a great benchmark. Don't listen to the haters. This one is especially interesting. https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-medium/qwen-qwen3-6-plus-medium/?order=qwen-qwen3-6-plus-medium%2Canthropic-claude-sonnet-4-6-medium https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...
- BoorishBears 6mo agoThis one's even more interesting https://aibenchy.com/compare/anthropic-claude-opus-4-6-medium/google-gemini-2-5-flash-medium/?order=google-gemini-2-5-flash-medium%2Canthropic-claude-opus-4-6-medium https://aibenchy.com/compare/anthropic-claude-opus-4-6-mediu... Who knew Anthropic was this far behind???
- jaggs 6mo agoYeah, but actually that's not a good look. Anyone who's used Gemini will know how random it is in terms of getting anything serious done, compared to the rock solid opus experience.
- BoorishBears 6mo agoTheir benchmark is chock-full of things like that: It's deeply flawed and is essentially rating how LLMs perform if you exert yourself trying to hold them entirely the wrong way.