10 ms·
DeepSeek V4 Pro 0813
https://api-docs.deepseek.com/ https://api-docs.deepseek.com/
https://artificialanalysis.ai/models/deepseek-v4-pro https://artificialanalysis.ai/models/deepseek-v4-pro
https://twitter.com/ChrisGPT/status/2087572834650407024 https://twitter.com/ChrisGPT/status/2087572834650407024, https://xcancel.com/ChrisGPT/status/2087572834650407024 https://xcancel.com/ChrisGPT/status/2087572834650407024
- linzhangrun 2mo ago[flagged]
- aabdi 2mo agohttps://api-docs.deepseek.com/quick_start/pricing/ https://api-docs.deepseek.com/quick_start/pricing/ Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.
- swiftcoder 2mo agoHow does it stack against the updated Deepseek Flash version?
- k__ 2mo agoAround 5 percentage points better. (E.g., 87% instead of 82%)
- Gecko4072 2mo agoSo not worth it over flash? Even at ~7x the size it isn't worth the price hike. Flash may be a monster of a model due to all the RL it received from free usage everywhere.
- k__ 2mo agoI tried the previous Pro model and in the end it was 50% more expensive than the previous Flash. Wasn't worth it.
- saaga 2mo agoYea that's what I was thinking. Flash is nuts. I find I have to be a more precise and specific with it but damn. It's crossed a threshold of production grade coding for sure. I was running a session over a couple days and it didnt cross a dollar lol.
- networked 2mo agoI haven't tried DeepSeek V4 Pro 0813 yet. Recent experience tells me that larger models are worth it in non-obvious ways. MiMo-V2.5-Pro solved problems that DeepSeek V4 Flash 0731 couldn't solve for me: for example, adding a live counter for elided reasoning lines to a terminal-based coding harness. You wouldn't be able to tell from the scores on their respective Artifical Analysis page (https://artificialanalysis.ai/models/mimo-v2-5-pro https://artificialanalysis.ai/models/mimo-v2-5-pro, https://artificialanalysis.ai/models/deepseek-v4-flash https://artificialanalysis.ai/models/deepseek-v4-flash). I like the DeepSeek V4 models, though. They critiqued my engineering decisions better than MiMo, and they seem to have a distinct aesthetic in the SVGs they write.
- trollbridge 2mo agoInteresting - I've been dropping into MiMo-V2.5-Pro-UltraSpeed whenever Flash seems to be "stuck" and it usually figures it out. I use UltraSpeed just because I'm so frustrated by then that I'm impatient. I still find 5.6-Sol can solve some things neither of those can, but it's so slow (and it's so hard to trace / debug the reasoning) that I just let it run overnight.
- networked 2mo agoWhat about 5.6 Terra and especially Luna? Luna scores pretty high on benchmarks and seems to have different habits (like a denser pattern of tool use) and blind spots. I'm trying out a development workflow where I generate mundane code with MiMo and Luna (and soon V4 Pro 0813?) and have Opus 5, which is running on only a Pro subscription, review and refactor it. I'm not sure it will justify the context switching, but it's an interesting exercise.
- trollbridge 2mo agoTerra and Luna are fine, but they’re quite slow (OAI seems to be really slow lately) and don’t have the reasoning traces. My workflow really depends on them or I can’t switch models effectively.
- npn 2mo agoI still believe this is not the full potential of pro models. I expect they will release another checkpoint later this year.
- eli 2mo agoOpus 5 medium to Opus 5 max is only 3 points, if that puts it in context
- sparkling 2mo agodeepseek-v4-flash feels so fast and snappy, i'm loving it. Happy to trade speed for the the 5% degraded benchmarking performance.
- k__ 2mo agoI wouldn't exactly call it snappy, but faster than Pro, yes.
- ericd 2mo agoSingle request depth on vllm with dspark, I'm getting ~200 tps, I'd say it's pretty snappy.
- JacobAsmuth 2mo agoWell sure but you're running on tens of thousands of dollars of hardware.
- ericd 2mo agoIt's much faster than other models on that same hardware in the same size class. I've tested a few, it's by far the fastest I've tested. And it wasn't tens* until recently. Didn't expect this to be one of my best performing assets this year.
- k__ 2mo agoI get like 80.
- ericd 2mo agoWhat's your setup? Happy to try to point you in the direction that worked for me.
- saaga 2mo agoI feel the same too. I like the speed. I'm also a big fan of glm 5.2 fast. I can't wait for like 2000 t/s on these haha.
- pixelesque 2mo agoI've found Pro to be a lot better per "task" than the recently released Flash for code reviews and things (via OpenRouter running in pi.dev). Flash makes a lot more initial mistakes, and then has to re-check stuff, and produces much more output compared to Pro. It often gets to the correct result eventually, but the output volume is often 5x more than for Pro, and the initial outputs are often wrong, with the first few saying something wrong (like there's a bug, or the code won't compile when it does), and then saying things like "Wait, let me re-check:", or "Actually, looking at it more carefully:" and then it thinks a bit more and eventually gets to the right answer.
- swiftcoder 2mo agoyeah, I've definitely noticed one has to be quite precise to keep Flash on the straight-and-narrow
- RALaBarge 2mo agoEvery plan and every code checkpoint finds me saying "Check with Grok and Fable latest to critique our strategy/code review" with pretty much every model. I havent ran into any deal breakers with the new Flash version yet (like it not running a tool properly or coming back with something completely daft)
- surgical_fire 2mo agoI use a plan -> implement wotkflow for this reason. pro plans, flash implements. I am super happy with how flash behaves like that.
- JacobAsmuth 2mo agoPer token. You need to look at pricing per task.
- trollbridge 2mo ago... which still comes out cheaper, since DeepSeek caches so much more. I keep track of my token consumption even on subscription plans and my equiv. cost for my 5.6-Sol usage is around $4000-$8000 a month.
- wmedrano 2mo agoSome benchmarks are out. Seems a bit weaker than Opus 4.8 but at least 10x cheaper once you account for verbosity. https://artificialanalysis.ai/?models=claude-opus-5%2Cclaude-opus-4-8%2Cdeepseek-v4-pro%2Cdeepseek-v4-flash%2Cclaude-opus-5-medium%2Cclaude-opus-5-high https://artificialanalysis.ai/?models=claude-opus-5%2Cclaude... Wonder how much more they'll squeeze out.
- xynelius 2mo agoIf that wasn't impressive enough, it's actually ~60x cheaper if you take into account the typical cache-read/input/output split in agentic coding, and the deep discount for cache reads offered by DeepSeek. Opencode has some public data on the typical split [1]: For DeepSeek V4 Pro the typical split is 750 in, 290 out, 82k cached. Cost per request for V4 Pro: $0.000875 per request. Equivalent Opus cost (w/o taking into account cache write costs): $0.052 per request. [1] https://opencode.ai/docs/go/#usage-limits https://opencode.ai/docs/go/#usage-limits
- taosx 2mo agoI created a simulation for coding harnesses based on my own pi sessions. When taking into account all factors, DS-v4-Pro is cheaper than gpt-5.6-luna due to caching. Look at the bill segments difference for cache read cost and uncached cost between deepseek and the other models. At this point is cheaper to use ds-v4-pro than the luna models from openai. ignore the numbers except the classic and keep in mind that classic is based on pi with the only change limiting tool output to 10kb https://harness.eveid.com/lazy-harness-cost-simulation https://harness.eveid.com/lazy-harness-cost-simulation * I built this for getting an initial estimate between different checkpoint/ compaction methods for the harness.
- RALaBarge 2mo agoHey this looks good! Maybe consider adding a hover-over popup for the rectangles explaining what each thing means to a lay person. I see it at the bottom, but that is below the fold.
- taosx 2mo agoDone, I'll take any other suggestions and apply them later, I will also split it a bit for different usecases as this was initially a throwaway prototype but found it useful. Basically it needs a bit more human touch.
- HDBaseT 2mo agoCan we have a conversation about subscription plans for a minute? I don't mean to hype up the US AI firms, but if a ChatGPT $200/m subscription can get you $16,000 in effective API costs, doesn't effectively every model get destroyed by the subsidized Claude/ChatGPT models? Both in price and intelligence.
- segmondy 2mo ago... and mere mortals can run this at home or rent a GPU, you can't do so with Sol or Fable.
- deleted 2mo ago[deleted]
- indigodaddy 2mo ago@dang - Pls merge this with https://news.ycombinator.com/item?id=49274018 https://news.ycombinator.com/item?id=49274018
- CharlesW 2mo agoEmail hn@ycombinator.com with anything you want HN mods to see. They're incredibly responsive.
- LeonKnst 2mo agoI find it interesting how much adoption seems to be influenced by momentum. Some of these Chinese models are surprisingly capable, but developers often default to the models that are already established as the “industry standard
- HawtAds 2mo agoHacker News is very Bay Area/US tech centric where spending a few hundred a month on AI is just pocket change. The weaker AI models with more questionable data retention policies are popular in developing countries. I think the new Facebook muse model will be similarly popular.
- BlackRabbit1 2mo agoA lot of it/infrastructure departments aren't aware that you can use Asian models hosted within the US or even EU.
- spacebanana7 2mo agoIn an enterprise setting Chinese models are often discouraged due to political risk. They don't want to need to remove a model that's deeply embedded in their stack. And it's entirely feasible that the US gov bans federal contractors from using them in the next 6 months for example, or that EU AI safety rules effectively ban them too.
- BlackRabbit1 2mo agoThere are EU/US providers offering Deepseek/Qwen/Kimi/etc.-as-a-Service. With zero ties of their infrastructure to China. Fully compatible with the well known Antrophic API. You only have to replace the URL and your key.
- odo1242 2mo agoBased on what the political climate looks like nowadays it's entirely possible the US bans federal contractors from associating with any company that uses the models themselves, regardless of data provenance or where they are hosted. Or they create AI safety rules that make it impossible to release open source models (for example, making it so that closed-source models can be evaluated with a harness but open-source models need to pass the benchmark with the weights alone, which isn't really possible). Or they just declare Chinese models a security risk like TikTok (claiming that the model would be trained to respect Chinese interests). It may not be likely but it's definitely possible enough to be something people worry about.
- scrlk 2mo agoBenchmarks: | Benchmark | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2 | Kimi-K3 | Opus-4.8 | Fable 5 | | | 0813 | 0731 | Preview | Preview | | | | (w/ fallback) | |--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------| | HLE (wo/w tools) | 42.7/60.0 | 37.8/51.5 | 37.7/48.2 | 34.8/45.1 | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.3/63.0 | | Terminal Bench 2.1 | 87.9 | 82.7 | 72.1 | 61.8 | 81.0 | 88.3 | 85.0 | 88.0 | | NL2Repo | 61.5 | 54.2 | 38.5 | 39.4 | 48.9 | - | 69.7 | - | | Cybergym | 83.3 | 76.7 | 52.7 | 38.7 | - | 80.0 | 78.3 | 83.1 | | DeepSWE | 62.7 | 54.4 | 12.8 | 7.3 | 46.2 | 67.5 | 58.0 | 70.0 | | Toolathlon-Verified | 74.1 | 70.3 | 55.9 | 49.7 | 59.9 | 76.5 | 76.2 | 77.9 | | Agents' Last Exam | 25.7 | 25.2 | 16.5 | 15.8 | 23.8 | 27.6 | 25.7 | - | | AutomationBench (Public) | 31.8 | 25.1 | 12.8 | 10.8 | 12.9 | 30.8 | 27.2 | 29.1 | | DSBench-FullStack | 71.1 | 68.7 | 41.8 | 37.0 | 61.8 | 73.7 | 71.6 | 77.2 | | DSBench-Hard | 67.2 | 59.6 | 31.1 | 25.8 | 54.5 | 63.0 | 71.7 | 68.3 | Source: https://reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepseek_v4pro0813_benchmarks/ https://reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepseek_v4...
- parsimo2010 2mo agoThe timing looks like they are trying to take the wind out of Qwen's sails by releasing this on the same day that Qwen released the weights of Qwen3.8-max. Or maybe it's coincidence... For comparison I looked at Qwen's claimed benchmarks for Qwen3.8-max (https://qwen.ai/blog?id=qwen3.8 https://qwen.ai/blog?id=qwen3.8). Assuming each published set of benchmarks is believable, it looks like v4 Pro 0813 is better on average but overall performance is comparable. Pro 0813 is much cheaper. If you don't need vision capabilities then you don't have much reason to use Qwen3.8-max. - 43.6 on HLE (Presumably without tools). Pro 0813 is a little worse. - 86.6 on Terminal Bench 2.1. Pro 0813 is better. - 55.9 on NL2Repo. Pro 0813 is better. - 27 on Agent's Last Exam. Pro 0813 is a little worse. - 72.5 on Toolathon-Verified. Pro 0813 is better. - 56.6 on DeepSWE 1.1. If the DeepSWE listed for Pro 0813 is the same version, then Pro is better. - 27.3 on AutomationBench. If the AutomationBench (Public) listed for Pro 0813 is the same, then Pro is better. I guess we do need to wait to see if the upcoming DS pricing increase is enough to change the value proposition. As it is now, they could double or triple prices and it still would be a better value to use DS. I bet they know that.
- Gecko4072 2mo agoCurrently burning money quickly on official deepseek api. They are also increasing pricing starting today. V4 Flash 0731 still feels like the most outstanding model of the past few months and probably to come.
- Jsttan 2mo agoWhat is the new price through?
- Gecko4072 2mo agohttps://api-docs.deepseek.com/quick_start/pricing/ https://api-docs.deepseek.com/quick_start/pricing/ edit: there are banner announcements saying v4 flash pricing will increase first then overall by an undetermined amount
- book_mike 2mo agoWhat I care about is whether the model is capable of the tasks I give it at the lowest cost. Right now I'm using Kimi-K3/GLM-5.2/Minimax. Sonnet is great but I burn through the tokens too fast. Opus 5 set to max is amazing and more intelligent than all of us. .998 of the time I don't need that kind of intelligence. I just need the job done.
- okamiueru 2mo agoHow do you define intelligence? I encounter that kind of sentiment all too often, and I have to assume we go by wildly different understanding of what that might entail.
- bikemike026 2mo agoIf you read Opus 5's output, it is beyond the comprehension of virtually all engineers and developers. That is what I mean by intelligence. Math, science, and engineering are all contained in one model. We may be experts in one field. The model is an expert in everything that humans know.
- logicchains 2mo agoYou mean Fable 5 right? Opus 5 makes lots of stupid mistakes about anything that requires any domain knowledge.
- hgoel 2mo agoI don't think that's because of its "intelligence". It speaks obtuse techbro-ese: stringing together words that sound smart to obscure the simplicity of the thing it's describing. In many ways it's the opposite of intelligence. Opus 5 and Fable 5 in particular suffer from this issue at worse level than most models in the same class.
- greenchair 2mo agoyep, it is so bad i had to create rules to cut down on the techbro language and domain slang.
- Readerium 2mo agoV4 Pro has vision correct?
- trollbridge 2mo agoNo.
- coredog64 2mo agoSaw somewhere that they don't believe vision advances the AGI work they're doing, so it's not on the roadmap.
- alecsm 2mo agoI've been using the last Deepseek Flash update for a week and I'm amazed. It was a capable model for easy tasks but now it looks like it can do some heavy development for peanuts. I can't wait to try this new one.
- coredog64 2mo agoIME I can't trust it to write it's own plans from a spec, but if I give it a detailed execution plan written by Opus, it's fast and cheap (if chatty) in executing it.
- stavros 2mo agoThis is what I do, and it works fantastically well. Just make sure you have Opus/GPT review after.
- polski-g 2mo agoInteresting. I use Flash for making the plans and GPT for execution.
- Kadin 2mo agoDepending on the language you're writing in and the problem domain, the smaller models can do dramatically better or worse. I suspect in the future we'll see language-specific small models. "Coding" is still pretty broad as an activity. It'd be nice to be able to load up a model specific to, say, class-based Python and run it on-device.
- eru 2mo agoHarmonic's Aristotle is a sort-of language specific model for Lean, if you want to see the future you described today.
- jatora 2mo agoflash for plans?! i don't understand why you wouldnt use something far stronger for the most load bearing point of the project
- yipinwong 2mo agoWorse than Luna but more expensive than Luna. Sticking with Luna without sending my data to Deepseek (China)
- comandillos 2mo agoAt least you can run it for relatively cheap hardware. I guess OpenAI doesn't let you do that.
- Eueudhsbsj32 2mo agoUnless you're Chinese, why would you care if they see your data? As an American, I'd much rather have my data kept outside the country than here where companies and the government have a lot more leverage over me.
- deleted 2mo ago[deleted]
- segmondy 2mo agoThis! It's always amusing when folks say "But China", my data in the hands of my government and their billionaire friends is more than dangerous than in China. I mean, if it's an IP sort of thing then go local.
- akman 2mo agoI do think this question comes up a lot-- I can understand why. For some well-explained reasons, check out https://darioamodei.com/essay/the-adolescence-of-technology https://darioamodei.com/essay/the-adolescence-of-technology and search for "CCP".
- Eueudhsbsj32 2mo agoSo is your concern more about reducing the risk of an authoritarian China "winning" the AI race? And less about reducing the risk of your data being used against you personally? To me, the risks of an individual helping China to continue to develop their AI by being a customer is pretty marginal compared with the personal risks of my data being used against me.
- nthypes 2mo agoStill behind Kimi-K3 in almost half of the benchmarks
- segmondy 2mo agoMuch easier & cheaper to run than Kimi
- jklmnopqrstuvw 2mo agoTested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project. Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug. Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.
- hugmynutus 2mo agoNullius in verba
- computerex 2mo agoRepeat the test like 5 times for each model and see the results.
- epolanski 2mo ago+1, a single test means little.
- jklmnopqrstuvw 2mo agoI don't think so. I specifically kept this PR to test model capabilities, and I've already tested a bunch of models. Current test results show that the more advanced the model is, the easier it passes. For example, GPT-5.5 Medium fails the test(has bug), but High passed.
- seunosewa 2mo agoDo it a second time at least.
- computerex 2mo agoThey are causal autoregressive models, the output is sensitive even to the implementation nuances in inference. Even 1 token that's badly selected could throw off the entire answer.
- 2mo ago
- eshack94 2mo agoIt appears that the only available endpoint (as of this writing) requires enabling "Allow paid endpoints that train on request data" in the OpenRouter privacy settings. I hope additional paid providers will become available that don't require training on data.
- cdolan 2mo agoThat is likely because Deepseek themselves is the only host. In 24-48 hours there will be other options I presume
- jubilanti 2mo agoTheir privacy policy doesn't forbid them from just straight up publishing your raw prompts as training data. My threat model is that anything I POST to DeepSeek I treat as public to the web, as much as a public GitHub repo is.
- nullbyte 2mo agoEven though cost-per-token is low, Deepseek v4 tends to burn an immense number of tokens to accomplish tasks.
- freakynit 2mo agoJust tested through openrouter.. gave exactly same task.. the task was to scan existing repo, and generate a single docker-compose file to deploy behind a caddy server, where certain port ranges are already used, the service demands widlcard certificates to be provisioned from outside, and postgre needs to be built-in one... Tested this model, and gpt-5.6-terra-high. Results: this one had few issues. terra: none. These results are consistent with my past observations with the latest flash version as well. What benchmarks say, vs what I've been observing are different. They are good till the project is simple... not anymore.
- shimman 2mo agoI've always wondered if I was using containers wrong because none of them I've ever had to create were complicated. Maybe it's because I choose tools that make local development easy (Go + sqlite + various CLTs) or maybe it's because I never hard to interact with this on the professional side outside of making images for our projects (which still weren't complicated for the reasons above). LLMs make containers in a pretty workable format for me (still hand tweak the env variables for a sanity check). How exactly does it struggle here and why does postgres need to be built? Were the needs beyond what you get in a base image?
- freakynit 2mo agoThis was the repo: https://github.com/amalshaji/portr https://github.com/amalshaji/portr And this was my gh issue: https://github.com/amalshaji/portr/issues/308 https://github.com/amalshaji/portr/issues/308 And below was my prompt: """ give me single docker-compose file that i can run on my server to run current project... you can read README.md , and then, this relevant page: https://docs-custom-reverse-proxy.portr-docs.pages.dev/docs/server/custom-reverse-proxy https://docs-custom-reverse-proxy.portr-docs.pages.dev/docs/... ... this was the result of me raising github issue: https://github.com/amalshaji/portr/issues/308 https://github.com/amalshaji/portr/issues/308 ... you can use gh cli to fetch the details and comments... i already have a caddy server running on my vps... and i will create wildcard certificates myself using certbot.. the domain name will be helloportr.xyz ... also, ports up to 9019 are already taken... ask me if anymore info is needed... """ You can try yourself and let me know of what you got.
- cjg007 2mo agoBefore DeepSeek-V4-Pro-0813's price goes up, I expect a surge of frantic traffic — hope the servers can hold up.
- cjg007 2mo agoDeepSeek raised their prices. Oh my god.
- simonw 2mo agoNice bicycle chain, the little basket with a fish didn't show up in the right place: https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F5d8474b8a9c1316c21ab5162b564f8ff https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
- wolttam 2mo agoI think I saw a better overall composition out of Flash 0731 Effort on this one?
- simonw 2mo agoDefault effort for OpenRouter. I'll try a grid of efforts... Wow, the low, medium, and high pelicans came out in surprisingly different styles: https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fc1108a380593547c2def5863bca63160 https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
- segmondy 2mo agoexciting. it's almost like 3 models in one. that variety would matter when trying to solve a creative problem.
- throwaway9af11m 2mo ago[flagged]
- simonw 2mo agoIf I live my life on the basis that some people don't share my sense of humor, and hence I should avoid doing anything funny that might be misunderstood, my life will be a lot less fun.
- throw10920 2mo agoYou're doing great. Don't let insanely low-effort (negative-effort, as in making others dumber rather than having no effect?) comments like from the above throwaway affect your actions.
- Myzura 2mo agoThis model is not very good at coding, but it is quite good at research, evaluation and action, I don't write code, but it really goes head-to-head with the most expensive models in searches such as stock market and forex
- ai_fry_ur_brain 2mo ago[dead]
- nimsarajay 2mo agoI'm Satisfied with this model (in opencode)
- Palmik 2mo agoWhy does this link to OpenRouter, which has no useful information on its own? Linking to the official API or the benchmarks would make more sense: - https://api-docs.deepseek.com/ https://api-docs.deepseek.com/ - https://x.com/ChrisGPT/status/2087572834650407024/photo/1 https://x.com/ChrisGPT/status/2087572834650407024/photo/1 (officially posted on WeChat, this is just one of many reposts)
- alexwwang 2mo agoDon’t you find the official document website was out of service for a long time since the new model was published soon?
- andreagosto 2mo ago[dead]
- echelon 2mo agoMoreover, OpenRouter is NOT Open Source, fair source, source available, etc. It's a proprietary cloud service that got first place in the API aggregation distribution game. Link to DeepSeek!
- ljlolel 2mo agoTrustedRouter is hosted and full opensource!
- zamadatix 2mo agoOpen doesn't always refer to the code. Just like their previous project, it refers to an open marketplace where anybody can sign up to sell access to models. But it'd still be nice to post to wait an extra minute to find some other page/new url from deepseek for it instead of posting that it exists somewhere.
- XCSme 2mo agoAgain, I will wait until there's a provider that doesn't train on prompts before I will benchmark.
- dakolli 2mo agopsst.. they all do. Also, what kind of IP are you protecting, are you protecting some crazy discovery, nothing you're throwing at them is special, they aren't going to steal your CRUD pomodora app. If anything Deepseek is the only company I'd want to consent to training on my data, they're by far the most altruistic. Atleast they give back all their IP in the form of research and open source weights. It's not like they're hoarding your data for them to make money, they're basically giving everything out for free. The only reason you even have the option of waiting for another provider is because they release weights. They're releasing all their IP, which is a trillion times more valuable than anything you're providing, you people are just greedy and oddly self centered.
- diydsp 2mo ago>are you protecting some crazy discovery Yes. If someone figured out my current project they would have a huge scoop.
- ai_fry_ur_brain 2mo ago[dead]
- XCSme 2mo agoI "trust" what they say on OpenRouter for the provider, for some it says they retain prompts, for other that they retain but can also use them for training. It's not any crazy IP, just my own benchmarks/tests, once they are in the training set it defeats the purpose of the tests, and I have to make new ones.
- LeBit 2mo agoThe good thing is that there seems to be quite a lot. Let’s just wait a bit for this one.
- moritzwarhier 2mo agoIs having padded version numbers with a leading zero a common thing? Wondering, sorry if it's a dumb triviality to ask. Is this even a (sub-)version number? I mean the major version is clearly 4.
- almyk 2mo agoIt is the date of the release
- moritzwarhier 2mo agothanks... XD
- gorfxx 2mo agoits for the month and day the model released in 2026. 0813 -> Aug 13th. I assume that if they use it internally the padded 0 makes finding the newest model easier cause all the numbers for the date line up instead of the zig zag you get without it once you get to 2 digit months.
- ernsheong 2mo agoThese people can't version control properly, V4.1 or V5 would be more appropriate.
- gigatexal 2mo agoWelp gonna give Deepseek more money. This is very cheap indeed. And I’ve been using them and kimi for a bit now not via open router but on my own and have found them on part with sonnet 5 though sonnet 5 these days I think has gotten worse. At work I had to move to Fable to get decent work results.
- monster_truck 2mo agoHave been letting it spin pretty hard (~$12.50 for 2B, 50% cache hits) on my traffic simulator/distributed physics engine all day, it's found some pretty significant gains without introducing any new problems. I'm happy
- blahyawnblah 2mo agoCan you tell me about your engine?
- monster_truck 2mo agoIt's ~a traffic simulator with the fidelity of a rally sim like DiRT Rally or Asseto. Full engine/drivetrain, grip modelling with tire deformation, suspension, aero, collisions, etc. I just really like Rally Racing and Car Accidents The absurd goal was to be able to simulate all of the active traffic in NYC, so 3-500,000 cars, without using any of the macro flow corner cutting that you see commercially or academically. Initially thought that one machine was not going to be enough to do this in realtime so I spun off a very big fork and built out a webtransport stack to split the effort over a local network. It was promising until I also wanted to cover the highway in thousands of giant beach balls [1] [2]. In the process of leveraging codegen to shrink the data that needed to be relayed by >10000x (packing and dynamically updating sparse continous arrays of floats), threw in a fuckton of LOD work (both spatial & temporal), statistical aggregation, and a lot of differential equation bullshit to derive LUTs. It's at the point where a base model M2 mini can handle much more than that by itself. All of the networking effort paid large dividends in cross thread coordination and lock-free data passing. Went from struggling to fit each of the sim kernels for just 32 cars @ 60hz (~11ms) to 0.03, 0.0003ms p99s with the corresponding jumps in car count (north of 32k/thread). Multiplayer works well enough, it falls apart where it should (~128 people or LLMs driving around in the same square quarter mile, and a great deal more NPCs, latency permitting) The remaining work is making it look and sound cool as fuck. Building out a physically-based audio synthesis engine that simulates the pulses of exhaust gas starting from the cylinder count/size/firing order, intake and exhaust count and size, header and exhaust configuration, and another one for the tire sounds, and another one for the collisions, and then rendering those out to wavetables so it scales and frees up the cycles for occlusion. And then getting it to visually render and control performantly in a browser tab, lots of instancing and shader work. There's also some bullshit cooking that uses the motion estimation built into GPU video compression engines for ... other purposes at stupid low latency. I figured that's what Waymo had to be doing so I let it rip It's fucking nuts, I'm having so much fun :) It will be done when it's finished [1]: https://i.imgur.com/BDQSuLv.png https://i.imgur.com/BDQSuLv.png [2]: https://i.imgur.com/CpreOWT.png https://i.imgur.com/CpreOWT.png
- Perenti 2mo agoGraphs without labels and/or scales on the axes are useless. I know less after viewing that page than before, but I got to see some pretty lines that I guess must mean something.
- CGamesPlay 2mo agoThe only graphs that don't have axes require a mouse to use. They show the values on hover—or if you tap the fullscreen button, that version also has axes. One graph is a 3-month time series showing a single day, so it looks like it doesn't have an X axis, but it does.
- anthropic-dario 2mo ago[dead]
- roysting 2mo ago[dead]
- halyconWays 2mo agoYou mean to tell me it's been 10 hours and there's no unsloth quant? I've been running 0731 and love it.
- jocelyner 2mo ago[dead]
- jeffmcjunkin 2mo agoArtificial Analysis page is up: https://artificialanalysis.ai/models/deepseek-v4-pro https://artificialanalysis.ai/models/deepseek-v4-pro
- LLLDP 2mo ago[dead]
- xbmcuser 2mo agoDid you guys read the fine print they plan to increase prices significantly in the future
- mrbnprck 2mo agoTake a look at openrouter there are a few alternative deepseek model providers that follow an even cheaper pricing strategy. Funnily even if deepseek themselves increase price 2-3x they are still more affordable.
- madhu_ghalame 2mo ago[dead]
- schmorptron 2mo agoThe past DeepSeek models and now these new checkpoints score very badly on the ArtificialAnalysis AA-Omniscience and hallucination rate benchmarks. I wonder where that's from? Maybe they're overindexing on coding even more than others? I can't say I've noticed it in my (coding) usage so far, has anyone seen it make up potential root causes or other speculative stuff more than other models?
- gunalx 2mo agoYeah. I stopped ising deepseek v4 flashbecause it is awful (even worse than my local qwen3.6 35B model) at multilingual prose.
- isqueiros 2mo agoWorking on a language related app makes me realize that all these supposed language models don't have many good language benchmarks
- mordae 2mo ago0731 is definitely tuned for coding. I mean: https://gertlabs.com/rankings?mode=agentic_coding https://gertlabs.com/rankings?mode=agentic_coding But it is also a decent translator from English to Czech in my experience.
- big-chungus4 2mo agoSo flash is 52 points on artificial analysis, and pro is 53
- WASDx 2mo agoThis was a disappointed to me. Why would I use pro over flash now? Is there some area where the difference is significant?
- lionkor 2mo agoThe idea is, I believe, that the Pro model is a larger model (more parameters, or less quantization) in general. What implication that has, I couldn't tell you. For tasks like pondering on something, reviewing code, etc. I use Pro, just because it feels like the right model for that.
- irthomasthomas 2mo agoThat is more an indictment of AA than DS
- WASDx 2mo agoOn DeepSWE it's now 53% vs 63% which is one of the coding benchmarks I trust the most. DS own measurements also show a more significant increase so I suspect AA might update when they release an article. Surprisingly DeepSWE currently shows a lower total cost for pro so that might also update I guess. As usual, don't trust the benchmarks and try for yourself.
- magekinnarus 2mo agoDS Pro is what I hoped it would be. I have used it as an auditor for a couple of implementations, and the work is solid. This will allow me to split the work between 5.6 Sol and DS Pro.
- tomveber 2mo ago[dead]
- nicman23 2mo agomannnn i just downloaded 0731 ffs
- minraws 2mo agoDeepseek V4 Pro 0813 is the most unreliable model I have tried, it works on pass@3 shockingly well you can get it to match Sol or Fable perhaps in task done, but it's horrendous at pass@1 very prone to going wrong and doing horribly at most benches. I am not sure what it is buy I suspect it might be GRPO.
- seunosewa 2mo agoCould you try setting the temperature very low e.g. 0.0?
- Hardd 2mo agoBased on my experience so far, compared to previous models, DeepSeek V4 Pro achieves results equal to or even better than before, but at a lower cost.
- thunfischtoast 2mo agoSounds like something a DeepSeek V4 Pro bot would say
- twelvechairs 2mo agoYou are absolutely right! Would you like me to respond in a more naturalistic way for hackernews denizens?
- KoolKat23 2mo agoPassed your personal turing test.
- Hardd 2mo agolol,That is indeed the case; I used it for the translation.
- floki165 2mo ago[flagged]
- swingboy 2mo agoI’ve found Flash 0731 to be pretty great recently. I do feel like a lot of the models are pretty close in terms of ability. I often run code through multiple different models _and_ harnesses for code reviews and they all typically find the same things as each other.
- arj 2mo agoBeen testing this on hobby project https://github.com/arj03/seedkernel/ https://github.com/arj03/seedkernel/. Latest flash was a big step up. Pro feels really slow compared. Claude opus is still better day this level.
- arj 2mo agoThat said. It is really good at security review at max settings. Just ran one, it came to 0.15$
- hnmuuatsno 2mo ago[dead]
- arggjarvs 2mo ago[flagged]
- hemkeshr 2mo agoI had high expectations for V4 Pro, especially since DeepSeek V4 Flash 0731 performed so well compared with other Flash models. What a letdown.
- erichocean 2mo agoBizarre, it has a nearly identical improvement as the flash model.
- aqme28 2mo agoWait, what are you disappointed by? Seems like a significant jump in performance, and it competes handily with other models.
- deleted 2mo ago[deleted]
- simjnd 2mo agoDeepseek V4 Flash 0731 was such a massive jump in capability for such a small model (and price), that I'm a bit disappointed by this release. I keep my agents on tight leashes, using them very interactively for bouncing off ideas, architecture, and then writing code (especially prototyping) and Flash has been crushing everything I ever needed it to do. Maybe my ambitions are too tame compared to people needing Fable / Sol grade models, but I'm probably staying on Flash and not moving on to Pro for the foreseeable future.
- srigi 2mo agoPeople when OW LLM looks good in benchmarks: benchmaxxxed People when OW LLM looks mediocre in benchmarks: disappointed
- ctx_wrangler 2mo ago[flagged]
- code51 2mo agoOpenRouter is a place with zero support. Suddenly get a big debt on your account with nobody to respond. As an early adopter of OpenRouter, I'm afraid they are in shambles.
- htrp 2mo agoIsn't that the model for all of the labs though?
- rochansinha 2mo agoOfficial deepseek announcement- https://x.com/deepseek_ai/status/2087864585504305397?s=20 https://x.com/deepseek_ai/status/2087864585504305397?s=20
- Palmik 2mo agoThe link should be changed to this.
- jannishan 2mo agoWhy do I feel that the Pro version's effects are inferior to Flash's? Is it just my imagination?
- m00dy 2mo agoDeepSeek is moving to peak/off-peak API pricing. Off-peak rates are 50% of peak rates. Peak: 01:00–04:00 UTC and 06:00–10:00 UTC Off-peak: all other hours New pricing takes effect August 16, 2026 at 16:00 UTC. Model Period Cache hit Cache miss Output (input / 1M) (input / 1M) (/ 1M) deepseek-v4-flash Off-peak $0.007 $0.22 $0.66 deepseek-v4-flash Peak $0.014 $0.44 $1.32 deepseek-v4-pro Off-peak $0.022 $0.66 $1.98 deepseek-v4-pro Peak $0.044 $1.32 $3.96 For batchable workloads, scheduling outside those two UTC windows cuts token costs in half.
- Ruca_AI 2mo ago[flagged]