4 ms·
DeepSeek models have such good benchmark performance, amazing pricing, and the team over there seems to be widely considered impressive. I just haven't found t
by habosa 12d ago
DeepSeek models have such good benchmark performance, amazing pricing, and the team over there seems to be widely considered impressive.
I just haven't found them to be very good? I've had a ton more success with the GLM models (since 5.2 anyway). Maybe I'm just holding it wrong, DS models seem to get stuck in loops or tell me nonsense. GLM feels like budget Claude.
- tired_star_nrg 12d agoIt’s funny, in oh-my-pi I’ve had much better luck with DS 4.1 Flash than GLM 5.3 flash. Glm seems to just constantly reason in circles before attempting a tool call
- tacomagick 12d agoDepends on a lot of factors, AFAIK Claude Code (TUI Harness) does not work well with Deepseek, but Opencode is surprisingly good.
- gpugreg 12d ago> DS models seem to get stuck in loops I have never had looping issues with DeepSeek models. Which provider/serving framework and harness are you using?
- pimeys 12d agoMy experience is complete opposite from yours. The V4 Flash was already quite good, but V4.1 is really very close to SOTA in my books. I've eval'd these models for weeks against Gemini 3.8, Kimi K3 and Opus 5, and DeepSeek absolutely wins these evals. It's really crazy. We price per token our customers. If Kimi was about 30% of the price of Opus 5 for the same quality, DeepSeek is 1/10th of a price of Kimi K3. We've come down in price so much that I seriously cannot recommend other models before they reduce their pricing. And what is really interesting is its programming ability. As I've said in my previous comments, I use agents a lot in my work. Since early Opus days until now I have 7-8 agents working in parallel for different tasks. Rust, design, GEPA, evals, analysis. For a long time Kimi K3 was the best model for this work, and before that GPT 5.5. But I still can't really believe how well DeepSeek works here. I really try to find faults from it, trying to see that it must be doing sloppy work and be worse than the others. But it does not. It finishes every task I give to it. And the cost per task is under a dollar, usually 15-30 cents. In comparison the same task with Kimi would be 3-15 dollars; sometimes closing to 100. And before that with GPT 5.5 a 800 dollar task was not uncommon if I spent days evaluating models. Now it's less than a dollar. For me if the other providers will not drop their prices dramatically in the coming weeks I see no reason to use them. Even with a 200 dollar subscription, paying per token for DeepSeek is better value. My harness: https://omp.sh/ https://omp.sh/
- stavros 12d agoDS 4.1 Flash did manage to hack my watch: https://github.com/skorokithakis/amazfit-neo-hacks https://github.com/skorokithakis/amazfit-neo-hacks Unfortunately, it made a mistake and bricked it. I handed over to Astra but noticed that my watch was dead, it must have died immediately before. Too bad, unlucky me.
- trollbridge 12d agoHave you tried Flash V4.1?
- rpdillon 12d agoWhat harness are you using? I've had incredible luck with OMP and Deepseek v4 Flash.
- copperx 12d agoOMP and DS are a match made in heaven.
- copperx 12d agoAre you using the API directly from Deepseek? Otherwise, performance will suffer greatly. If you're using OpenRouter make sure to pin Deepseek as the provider. If you're using OpenCode Go who knows who the provider is.
- 0xbadcafebee 12d agoHow are you holding it? You'll get wildly different real-world results from different inference providers, for one. For two there's the harness and what you do with it. In general you'll need to provide more direction to lightweight models, better guardrails, more planning, tighter goals.