3 ms·
> What’s improved? Language consistency: fewer CN/EN mix-ups & no more random chars. It's good that they made this improvement. But is there any advantages at
by sbinnee 1y ago
> What’s improved? Language consistency: fewer CN/EN mix-ups & no more random chars.
It's good that they made this improvement. But is there any advantages at this point using DeepSeek over Qwen?
- IgorPartola 1y agoI wish there was some easy resource to keep up with the latest models. The best I have come up with so far is asking one model to research the others. Realistically I want to know latest versions, best use case, performance (in terms of speed) relative to some baseline, and hardware requirements to run it.
- exe34 1y ago> asking one model to research the others. that's basically choosing are random with extra steps!
- throwup238 1y agoResearch not spit out the answer based on weights. Just ask Gemini/Claude to do deep research on /r/LocalLLama and HN posts.
- Jgoauh 1y agohave you tried https://artificialanalysis.ai/ https://artificialanalysis.ai/
- JimDugan 1y agoDumb collation of benchmarks that the big labs are essentially training on. Livebench.ai is the industry standard - non contaminated, new questions every few months.
- IgorPartola 1y agoThanks! Are the scores in some way linear here? As in, if model A is rated at 25 and model B at 50, does that mean I will have half the mistakes with model B? Get answers that are 2x more accurate? Or is it subjective?
- esafak 1y agoI believe the score represents the fraction of correct answers, so yes.
- alexeiz 1y agoIt says the best "coding index" is held by Grok 4 and Gemini 2.5 Pro. Give me a break. Nobody uses those models for serious coding. It's dominated by Sonnet 4/Opus 4.1 and GPT-5.
- __mharrison__ 1y agoI use Aider heavily and find their benchmark to be pretty good. It is updated relatively frequently (a month ago, which may be an eternity in AI time). https://aider.chat/docs/leaderboards/ https://aider.chat/docs/leaderboards/
- comrade1234 1y agoMIT license that lets you run it on your own hardware and make money off of it.
- coder543 1y agoQwen3 models (including their 235B and 480B models) use the Apache-2.0 license, so it’s not like that’s a big difference here.
- coder543 1y agoThey seem fairly competitive with each other. You would have to benchmark them for your specific use case.
- twotwotwo 1y agoThe fast Cerebras thing got me to try the Qwen3 models. I couldn't get them working all that well: they had trouble using the required output format and following instructions. On the other hand, benchmarks say they should be great, and it sounds like maybe some people use them OK via different tools. I'm curious if my experience was unusual (it very much could be!) and I'd be interested to hear from anyone who's used both.