6 ms·
I'm using Deepseek-v4-pro as my main model and this is sometimes pretty annoying, I have to do some easy boring task, think "I'll just leave the agent to do it
by flexagoon 4mo ago
I'm using Deepseek-v4-pro as my main model and this is sometimes pretty annoying, I have to do some easy boring task, think "I'll just leave the agent to do it and go take a nap", but it's already done writing the code before I even walk away from the computer
- RussianCow 4mo agoDo you mean Flash and not Pro? I haven't tried it personally, but according to OpenRouter, the fastest DeekSeep V4 Pro providers are only ~50tps. That's slower than Claude Opus. https://openrouter.ai/deepseek/deepseek-v4-pro?sort=throughput https://openrouter.ai/deepseek/deepseek-v4-pro?sort=throughp...
- specproc 4mo agoYeah, flash is crazy fast, but I've found performance variable.
- binary0010 4mo agoFlash is amazing if you know the domain really well. E.g. occasionally it makes the dumbest mistakes you've ever seen and can't correct them. However it's fairly rare, and if you know the domain really well, occasionally popping in the code and pushing it towards the correct solution takes like 20seconds or whatever. So the speed you can move with flash + high domain knowledge beats opus by a mile in my experience. I tried to switch back to 4.8 for a bit when it came out, feels so bad waiting 20mins for a mediocre solution when I could have had everything complete - with multiple iteration cycles - in flash in like 3-5mins.
- addozhang 4mo agoYes, you don't need much domain knowledge to use Opus, but it's just way too expensive.
- 59nadir 4mo agoFor losers who can't put together a program to save their life, have no real skills and were always not really interested in programming (hence their poor skills), renting a robot buddy to do it for them is a good deal, until the buddy cuts in materially into their salary, and until their bosses realize that they really just have robot operators on staff instead of people who can actually do things.
- Induane 4mo agoIt's nice when I want to be lazy though. Or when I'm working two contract gigs. I can spec things out for one and turn it loose and trust it. Then work more closely with deepseek on the other project.
- sarjann 4mo agoI don't think token speed matters as much when a lot of tokens are needed to achieve a task. E.g. artificial analysis benchmarks where deepseek v4 is one of the biggest token burners to go through the benchmark.
- brianwawok 4mo agoBoth matter.
- flowbarai 4mo ago[flagged]
- SwellJoe 4mo agoIn recent benchmarking I've been doing, DeepSeek V4 Pro was the fastest of 21 models, by a comfortable margin (https://swelljoe.com/html/bench-report-final.html https://swelljoe.com/html/bench-report-final.html). Faster than Claude Opus 4.8, which was the second fastest (Mistral doesn't count because it seems to have refused to participate). But, it's a limited data set, just a few benchmark runs of a limited set of tasks. It's entirely possible I happened to be calling the API at its least busy time and maybe Claude got hit during a busy time.
- flexagoon 4mo agoNo, I mean Pro. I use it through OpenCode Go so I don't know what provider it uses under the hood, but it's very fast in my experience.
- thecopy 4mo agoDS through OpenRouter is significantly slower than direct from DS platform in my experience
- tmaly 4mo agoThis reminds me of the Peter / Boris comments on writing loops to keep the agents busy.
- deleted 4mo ago[deleted]
- deleted 4mo ago[deleted]
- throwaway67678 4mo agoAgent mania setting in It's also pretty funny sometimes how it gives weird future roadmap estimates ("part 2 - 3 weeks, part 3 - 2 months", etc.) and when you tell it to actually do those changes it's pretty much done in half an hour
- smith7018 4mo agoI've long believed those numbers were faked by Anthropic/OpenAI to serve as a form of advertisement. The estimates are impossible to verify and their ability to do "2 days of work" in 10 minutes will presumably make the user go "Wow, I just saved SO much time!" Plus, the unnecessary text eats up the users' tokens so it helps the companies on the backend, as well.
- leodavi 4mo agoI agree with you that labs are benefiting from those outputs but I'm skeptical that labs are purposefully training the models to produce those outputs. Raw pre-training data includes plenty of conversations between professional builders and some of those include estimates. I believe the outputs are a training coincidence with consequences that are opportunitistic for the labs.
- AgentMasterRace 4mo agoAll the models have broken estimates. They're trained heavily on jira and GitHub tasks and issues, that's why their estimates are human.
- esperent 4mo agoEven for humans the estimates are way off, unless it's based on data that has some serious padding. That said, it'll often say "2 days of work" and then complete the coding in 30 minutes, and while that's amusing, afterwards, I'll need to manually test, or send to other people for review, or realize the agent only actually did half the work and I need to do a second pass (or a third etc.) and then often getting the feature in does genuinely take two days.
- behnamoh 4mo agoSame. How can DeepSeek serve the V4-Pro at such high speeds despite the sanction?
- rubyn00bie 4mo agoThe sanctions only “prevent” them from directly buying NVidia’s latest and greatest in the sense that NVidia can’t sell directly to them. Essentially, there are companies now who are in a country without the sanctions, they buy from NVidia (or a partner), and then ship them off to China. For the orgs in China doing this, there’s zero legal risk besides having foreign customs service intercept the shipment and losing the goods. For NVidia there is zero incentive to care, as long as they look like they do, because sales are sales. You can bet Jensen ain’t losing sleep over it. GamersNexus had a really good investigative piece (~3hrs long) on this where they went to China and met with grey market sellers. That piece absolutely pissed off NVidia and resulted in a fight with Bloomberg too. Deepseek may be also be running inference on oodles of Chinese hardware but it wouldn’t surprise me for a second if they just acquired Blackwell chips through the grey market. The original Deepseek models were all trained using NVidia chips if I remember right.
- seewhydee 4mo agoThat wouldn't explain why Deepseek is fast relative to other Chinese providers, especially considering that they're reportedly ahead of the curve among Chinese companies in moving off Nvidia. I think their quant fund background has more to do with it. Their models are clearly designed with performant inference clearly in mind.
- ljosifov 4mo agoYes, it's performant, and esp performant at non-trivial context depths. DeepSeek-V4 DS4 (and Flash - DS4F) drop tok/s speed much less than the rest. On my M2 Max it took context depths of 768K to drop tok/s to ~10 tok/s. https://x.com/ljupc0/status/2062457314414587996 https://x.com/ljupc0/status/2062457314414587996 Other local models I've checked drop to unusable speeds way sooner. Only other model with similarity favourable curve I've tried is nemotron-cascade-2-30b-a3b. But it's a small model, way dumber than DS4F. Coding agents use cases have large context depths. The rate of decline is as important as the headline number.
- binary0010 4mo agoI exclusively use deepseek v4 flash now, completely stopped using slow models like Claude. Basically I never have to wait - yes I have to tell it little corrections occasionally (but I know the domain really well so that's not an issue), but it's so much faster than anything else it's kinda crazy. I love the super fast speeds with high involvement development cycle. I actually enjoy using agentic development flows for the first time now - whereas with Claude I absolutely hated it. That 5 to 20 min wait after every prompt absolutely killed my desire to even want to work at all.
- SwellJoe 4mo agoDeepSeek is the fastest model in the benchmarks I've been doing (https://swelljoe.com/post/will-it-mythos/ https://swelljoe.com/post/will-it-mythos/). Followed not so closely by Opus 4.8 and even less closely by Gemini 3.5 Flash and GPT 5.5. I've been really impressed with it, so far. It's also among the best at doing the work, though still trailing the frontier models from Anthropic and OpenAI.
- anschl 4mo agoNice benchmark, thanks! Which quants did you choose for the self hosted models?
- SwellJoe 4mo ago8-bit on that one (unsloth 8_K_XL). But, the next post compares all common quantizations of Qwen 3.6. I have another coming in a day or so for Gemma 4 with the 4-bit QAT version, which is very surprising (in a good way, Gemma 4 is impressive for this task).
- throw-the-towel 4mo agoFWIW, for me just today it got itself into silly rabbit holes twice, and both times I had to fix things myself. Scarily, this is something I catch myself doing as well.
- andai 4mo agoWith Flash it's basically instant for smaller tasks, yeah.
- znpy 4mo ago> I have to do some easy boring task, think "I'll just leave the agent to do it and go take a nap", but it's already done writing the code before I even walk away from the computer the way software engineering works these days reminds me a lot of factory workers on production lines that just sit in front of a production line all day and take out faulty items and/or perform a single step in the production of goods.
- abustamam 4mo agoTake the nap anyway, just say it took all afternoon :)