3 ms·
Qwen has previously engaged in deceptive benchmark hacking. They previously claimed SOTA coding performance back in January and there's a good reason that no so
by stuartjohnson12 1y ago
Qwen has previously engaged in deceptive benchmark hacking. They previously claimed SOTA coding performance back in January and there's a good reason that no software engineer you know was writing code with Qwen 2.5.
https://winbuzzer.com/2025/01/29/alibabas-new-qwen-2-5-max-model-takes-on-deepseek-in-ai-benchmarks-xcxwbn/ https://winbuzzer.com/2025/01/29/alibabas-new-qwen-2-5-max-m...
Alibaba is not a company whose culture is conducive to earnest acknowledgement that they are behind SOTA.
- swyx 1y ago> there's a good reason that no software engineer you know was writing code with Qwen 2.5. this is disingenous. there are a bunch of hurdles to using open models over closed models and you know them as well as the rest of us.
- omneity 1y agoAlso dishonest since the reason Qwen 2.5 got so popular is not so much paper performance.
- stuartjohnson12 1y agoThose hurdles exist because they're worse for most people. You think Cursor wouldn't spin up their own Qwen inference cluster or contract with someone who can if doing so would give them SOTA code editing performance against Claude?
- stocksinsmocks 1y agoThere is also paranoia that the Chinese government may compel their tech companies to play dirty tricks on their users. Yet without a trace of irony the critics have nothing to say about this not-so-secret practice for US based technology companies.
- pxc 1y agoClearly the thing we should want is a healthy, international AI ecosystem characterized both by cooperation and by competition, so that we are free to choose between models developed under different conditions, for compliance with different laws, subject to different cultures and biases, pressured by different interests, etc. To the extent that there's a solution, the solution is choice!
- daemonologist 1y agoMaybe not the big general purpose models, but Qwen 2.5 Coder was quite popular. Aside from people using it directly I believe Zed's Zeta was a fine-tune of the base model.
- sourcecodeplz 1y agoBenchmarks are one thing but the people really using these models, do it for a reason. Qwen team is top in open models, esp. for coding.