12 ms·
DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
DSeek plans to officially release the V4.1 Flash model around September 10, 2026 (Beijing Time). After extensive internal and external testing, V4.1 Flash has comprehensively surpassed V4 Pro across all key metrics, including performance, cost, speed, and task completion time. In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price. If you encounter any issues during your comparative testing between V4 Pro and V4.1 Flash, please do not hesitate to reach out to us with your feedback. Thank you for your support!
We will adjust the pricing for the Flash series effective from 12:00 Beijing Time on September 10, 2026. During off-peak hours, the unit price will be $0.003 for input cache hits, $0.15 for input cache misses, and $0.6 for output. Peak-hour prices will be double the off-peak rates. Please plan your usage accordingly.
- stanac 26d agoMy problem with V4 flash is output limit. When I need to write or rewrite a larger file (~1000 lines of code) it will fail with message like output limit reached.
- surgical_fire 25d agoSounds like a problem with your provider, or with how your harness sends requests. All DeepSeek models have 384k maximum output tokens: https://api-docs.deepseek.com/quick_start/pricing https://api-docs.deepseek.com/quick_start/pricing
- neugls 26d agoWaiting to use it
- oefrha 26d agoSource is apparently a banner announcement on https://platform.deepseek.com/usage https://platform.deepseek.com/usage. Had me searching for a couple minutes...
- nickweb 26d agoI swear I put that at the start of the post. Must've managed to miss it when copy and pasting!
- NooneAtAll3 26d agocan't you edit it?
- nickweb 26d agoNo. Posts are only editable for a small window of time.
- deleted 26d ago[deleted]
- GreenWatermelon 26d agoI received an email for it.
- swiftcoder 26d agoIf they can keep up this cadence of Flash leap-frogging the previous Pro, we're in for a good time
- dude250711 26d agoAnthropic/OpenAI might step up their anti-distillation defences though.
- _aavaa_ 26d agoThey might need to step up their product offerings and offer cheaper.
- swiftcoder 26d agoNah, we're long past the point where that would make a difference - if they could have done so effectively, they would have before K3 and GLM 5.3 were nipping at their heels...
- alightsoul 26d agoAnti distillation is just American nationalistic marketing, the amount of data they have "distilled" makes no difference. Deepseek "distilled" like 10000 messages? That's clearly not enough. You gain a lot more doing RL in house.
- tarruda 26d agoHopefully it will be open weights and have the same architecture and size as the current v4 flash vision, which is probably the best LLM that can be run on 128G devices.
- fluoridation 26d agoInteresting, I had assumed it'd be too large to fit. What quant and context size are you running?
- tarruda 26d agoIQ3_XXS (~3.2 BPW). For me this is an option because my Mac studio is only used for serving LLMs, so I can afford to dedicate most of its RAM to this. I can run with 256k context and only uses ~117G, with the remaining (up to 125G which I can allocate to VRAM) being used for prompt caching and context checkpoints. I'm making my own quants, though the Vision-Exp version is outdated and won't work on llama.cpp master branch (I built it before llama added support): - https://huggingface.co/tarruda/DeepSeek-V4-Flash-0731-GGUF https://huggingface.co/tarruda/DeepSeek-V4-Flash-0731-GGUF - https://huggingface.co/tarruda/DeepSeek-V4-Flash-Vision-Exp-GGUF https://huggingface.co/tarruda/DeepSeek-V4-Flash-Vision-Exp-... For the Vision-exp version, I also ran perplexity + KLD against the original MXFP4. Seems quite OK: https://huggingface.co/tarruda/DeepSeek-V4-Flash-Vision-Exp-GGUF/blob/main/logs/perplexity-IQ3_XXS.txt https://huggingface.co/tarruda/DeepSeek-V4-Flash-Vision-Exp-...
- fluoridation 26d agoThanks, I'll give that a try. I basically have the same use case, only on Strix Halo.
- tarruda 26d agoDon't use my Vision-Exp GGUF though. As I said I built those GGUFs before llama.cpp supported, and they can't be loaded on current master (require my own branch). I already have new GGUFs but haven't uploaded yet. If you want Vision-Exp, maybe use bartowski or unsloth's GGUFs. Side note: As an alternative to deepseek v4, you might want to give it a shot at qwen 3.8 flash next. I have IQ4_NL GGUFs that can be loaded fully into 128G, or Q5_K GGUFs that can offload the PLE to disk (use --load-mode none --lazy-mode on for that): https://huggingface.co/tarruda/Qwen3.8-Flash-Next-GGUF https://huggingface.co/tarruda/Qwen3.8-Flash-Next-GGUF. llama.cpp master is still somewhat bad in Qwen 3.8 next performance, but I was able to achieve 40tps tg and 600 tps pp on my private branch.
- nickweb 26d agoVia nitter: https://xcancel.com/JustinGorya/status/2097287080128708930 https://xcancel.com/JustinGorya/status/2097287080128708930 Looks like the new model can be used if summoned via the API but the API won't list it.
- igleria 26d agov4 pro was decent then a better cheaper faster model comes now? As a consumer I feel like hansel and gretel combined, deepseek could be the witch.
- throwaway473825 26d agoIt's not unprecedented given that GLM 5.3 Flash was better and cheaper than GLM 5.2.
- calgoo 26d agov4 flash has been working quite well for the majority of my personal projects, with occasional v4 pro or Kimi 3 for the most complicated tasks or to check the overall project progress (when vibe coding).
- HansHamster 26d agoI must be doing something wrong. I gave v4 pro a try a couple of days ago, gave it a simple prompt like "clean up functions x and y in file z" and it would always start off promising, just to quickly get sidetracked, start hallucinating problems in the code, and just get stuck for hours until I interrupt it: — hmm — 0x2D696370 — little-endian bytes: 70 63 69 2D = 'p','c','i','-' — hmm — WAIT — WAIT — !!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — *HOLD ON — HOLD ON — HOLD ON — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — !!!!!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — *OK — WAIT — I THINK I FINALLY SEE THE WHOLE PICTURE — I NEVER READ IT — AND — THE LAYOUT — hmm — !!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — HOLD ON — HOLD ON — HOLD ON — HOLD ON Then gave the same to Sonnet 5 and it was done 15 - 30 minutes later. I tried v4 pro both in claude code and codewhale with similar results. Haven't tried the new deepseek harness.
- k__ 26d agoI used flash with pi and it worked pretty well. It built this whole IaC plugin from scratch: https://github.com/fllstck/nebius-alchemy https://github.com/fllstck/nebius-alchemy
- nicce 26d ago> In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price. If you encounter any issues during your comparative testing between V4 Pro and V4.1 Flash, please do not hesitate to reach out to us with your feedback. Thank you for your support! Wow. Imagine OpenAI/Google/Anthropic doing this! Nope.
- donk8r 26d ago[flagged]
- thrownaway561 26d agoI will continue to be amazed by how much power you get from DeepSeek Flash for the cost. I have let that puppy lose on so many projects and it is has never let me down. It can build and entire Rails app in no time and even do the tests. For most things, I don't get why people pay the money for Claude. DeepSeek Flash is my default agent in Omarchy.
- bwfan123 26d agoSame, very impressed with v4 flash. It has the right balance of cost and performance.
- myaccountonhn 25d agoHow does it compare to GLM-5.3 Flash?
- EbNar 26d agoSince a few months, I almost exclusively use the Chinese "flash" models for my needs. They are a joy and they cost pennies per answer. Great job.
- XzAeRosho 26d agoSame for me. DeepSeek models are incredibly good at implementation and light planning. I still default to Opus models for feature planning, but for most simple features the Pro models suffice. Incredible good value and product they have built.
- ActionHank 26d agoI am legitimately more excited for this release than any frontier models at this point. I don't need a model that can invent new mathematics. I need something that is fast, cheap, and consistent. Give me that and I can build and scale.
- Oras 26d agoLLMs are not consistent
- ActionHank 26d agoTrue, make them cheap and fast enough and you can scope and stack agents sufficiently that the error rate tends close enough to zero to be meaningfully useful.
- darkoob12 26d agoMy mental bias always kept me away from Chinese models. Because i know that china is a surveillance state and all the things we know about CCP. But after what we learned about OpenAI and how they most likely used user data to basically cheat in an open competition i think it does not matter which AI provider you use all of them will own your data and all of them can spy on you. So I am willing to switch to Chinese models. This way we help them develop and improve models some day we can run them locally.
- postalcoder 26d agoI hope DeepSeek takes some time to improve their tuning for reasoning effort. Right now, there are only three reasoning efforts: low, high, and max. For all intents and purposes, "low" is pretty much the same as turning reasoning off, and "high" is similar to "max". "High/max" performs way too much reasoning, takes forever, and causes costs to balloon. They need a proper "medium" setting. I get it that they're probably focused on pushing performance right now, but the ergonomics of the model aren't great.
- iamniels 26d agoI switched to GLM-5.3 flash on high for this reason. Too many "but wait" in the Deepseek-v4 reasoning.
- alfiedotwtf 25d agoGLM 5.3 Flash was the killer for me. It really feels like we have Claude-approaching models at home. I just wish they kept parameter count down in order to fit entirely within commonly used RAM sizes
- tarruda 26d agoI would rather have just 3 levels: low, medium and high.
- jiehong 26d agoSounds nice! But, the web ui chat version of flash has very poor language following abilities in my experience: You may ask it something in English, and get a thinking chain in Chinese with an answer in Chinese, or an English thinking chain and an English answer. Using the retry button on the same question has a 50/50 chance of any of those results. Sometimes, asking something in English, but where information are mostly in another language may make the answer in the language where data has been found. The other day, I asked something about a local German thing, in English, and I got an answer in German instead. It’s as if all the language data stirred it away from the language of the user’s question.
- swiftcoder 26d agoI've hit this too, but you can just add "in English" to steer it
- gentlewater 26d agoI finally uninstalled the app yesterday after giving it plenty of chances over several months. Yesterday, I asked it whether «DeepSeek has fixed the issue where it erroneously answers in Chinese?» and it answered in Chinese.
- elaus 26d agoSo you did not do what the post you replied to suggested?
- michimagdesign 26d agoThis shouldn’t be a user-facing issue. The web UI should inject the account’s language setting or solve it like competitors. They’ve mentioned giving it multiple chances but it’s still not fixed.
- alightsoul 26d agoAnthropic does the same thing but it's not problem
- hinow 26d agoBut what no one mentions is that the price is going from a starting point of $0.16 to $0.60, so basically they're charging nearly four times as much.
- _aavaa_ 26d agoWhat are you talking about? Current flash prices are 0.66 for output, this is dropping it to 0.60.
- hinow 26d agoThis is the notice from DeepSeek regarding their API: We will adjust the pricing for the Flash series effective from 12:00 Beijing Time on September 10, 2026. During off-peak hours, the unit price will be $0.003 for input cache hits, $0.15 for input cache misses, and $0.6 for output. Peak-hour prices will be double the off-peak rates. Please plan your usage accordingly. ------------------------------------------------------- Hoje em sites como openrouter o valor é de $0.16 output .
- _aavaa_ 26d agoYes, and? This is their current off-peak pricing for their flash model [0]: $0.007 cache, $0.22 input, $0.66. 0.007 -> 0.003 0.22 -> 0.15 0.66 -> 0.60 Each one is now cheaper. [0]: https://api-docs.deepseek.com/quick_start/pricing/ https://api-docs.deepseek.com/quick_start/pricing/
- KyleTheDev 26d agoDirect 1:1:1 comparison, for V4.1 Flash - V4 Pro 0813 - V4 Flash 0731 Input cache hits (per 1m tokens) - $0.003 Vs. $0.022 Vs. $0.007 Input cache miss (per 1m tokens) - $0.15 Vs. $0.66 Vs. $0.22 Output (per 1m tokens) - $0.6 Vs. $1.98 Vs. $0.66 This is taken from https://api-docs.deepseek.com/quick_start/pricing https://api-docs.deepseek.com/quick_start/pricing, and it's comparing only off-peak hours pricing. It looks like V4.1 Flash is cheaper than the current 0731 flash model, and much cheaper than the current V4 Pro model.
- 26d ago
- tensegrist 26d agoIn keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price. just in terms of user perception when selling this sort of service, this is what they call a "good look"
- indigodaddy 26d agoSo, will it have vision? (based on deepseek-v4-flash-vision-exp ?)
- ComputerGuru 26d agoAlready available via the API as deepseek-v4.1-flash-expires-on-0910, with vision.
- a-ve 26d agoI've been using deepseek-v4-flash as a "worker" model with Claude Code to implement a tool using Rust/Iroh for my personal use, and it works fairly nicely when I use Opus as the planner/reviewer model. It seems to follow the plan generated by Opus, albeit with a few misses here and there that it cleans up later after being reviewed by Opus. Fairly excited for the v4.1 launch. Input cache hit prices have been halved, which looks nice.
- bitexploder 26d agoIf you are okay with waiting use GLM 5.3 max. It costs more but still cheap. It is slow, but a very strong worker. Still dollars per day (at most) with heavy concurrent agent running. I load up planning and tasks in Opus or Sol, and just have glm flash workers go to town every night. My project has never advanced more smoothly.
- ThouYS 26d agoif this beats GLM 5.3 flash, I am sold
- vib08 26d agoi think this will beat GLM 5.3 flash, i mean the last ds4-flash after preview was already great, and i found that more intelligent than GLM 5.3 flash
- simonw 26d ago> all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price If I'd carefully tested and optimized prompts against Pro I wouldn't be keen on this particular news. I feel like API model providers should lean towards not swapping out models on their paying customers, no matter how much "better" the new model is meant to be.
- tjwebbnorfolk 26d agoSure, but if a company decided to place a remote chinese hedge fund's API at the center of a critical business workflow, this is a lesson better learned sooner rather than later.
- k__ 26d agoAt least they can use another provider or self-host.
- satory 26d ago[flagged]
- badatnames 26d agoI imagine they're doing this due to capacity issues or somesuch. They can always relaunch Pro later, meanwhile a little ricered benchmaxxing of their existing flash model provides a temporary cover story. They certainly aren't silly enough to think this won't impact existing Pro users </paranoia>
- NitpickLawyer 26d agoIt's interesting that this is the third lab to find problems with larger models. Earlier last year oAI was rumoured to have failed their large pretrain. Now google has problems with their pro series, and ds just announced the same. There are some rumours on chinese forums talking about problems with the pretraining phase, so this is not mid/post training related. I wonder if this comes from using the bad architecture scaled up (and it hits some limits) or if this is a data problem (undertrained? bad data? bad pre-processing using smaller models?)...
- pixelesque 26d agoWhere are you seeing them having an issue with the larger (Pro) model? The announcement specifically says 4.1 Pro will be released in the future.
- petu 26d agoA month ago new V4 Flash 0731 checkpoint was better than existing V4 Pro. They've kept serving Pro, it was updated 13 days later (0813 checkpoint). Now, 4 weeks later new Flash checkpoint (0910?) is again better than existing Pro. Same situation, but Pro is taken offline this time.
- wolttam 26d agoJust my intuition about it but it does seem like a data issue. V4 flash and V4 pro feel very similar, which would make sense if they were pre-trained on largely the same corpus. All that would suggest to me is that V4 Flash is capable of absorbing the data they’re throwing at it, and we’re still nowhere near the data limits of their larger 1.6T model
- edude03 26d agoI've been watching a bunch of bycloud on YouTube recently, and although he's done a great job reviewing papers from the big AI labs, I feel like I'm missing something - how have all the labs seemingly made a model that's cheaper, faster AND has better performance? Historically `flash` variants (like codex spark as well) have been faster but perform worse
- _3u10 26d agoThat’s how increasing performance works. You make a model 10x faster, then you make it think 2x as much. Its cost is now 1/10th per token, and 1/5th per task. Basically they have shitty hardware so they have to do a lot of optimization. Think of it like replacing an O(n) algorithm with O(log n). Anthropic / Open AI think the best path is the most intelligent models deepseek is more focused on tok/$
- slickytail 26d ago[dead]
- k__ 26d agoBeta testers report >400 TPS. https://www.geeky-gadgets.com/deepseek-v4-1-flash-review/ https://www.geeky-gadgets.com/deepseek-v4-1-flash-review/ I hope some of those speed increases will make it to production.
- esafak 26d agoThat ought to be DeepSeek's real differentiator; all the other Chinese models are slow.
- aftbit 26d ago>In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price Please don't do this kind of thing. If a user has validated a workflow on V4 Pro, they might not want to suddenly start testing it in production on V4.1 Flash. Instead, keep V4 Pro around but deprecated for a defined period of time, then remove it. At least as open weights models, it's possible to use something like Together.ai or OpenRouter to run the V4 Pro model as long as other providers keep it up.
- m3kw9 26d agoThey expect vibe coders to use their models only lol
- KoolKat23 26d agoI imagine they need the compute. Can expand market share with more users for same amount of compute. But I agree with you. I have a dumb workflow that worked well with v4-flash-0731 and I suspect is directing to a newer model that now breaks it.
- petu 26d ago4.1 releases tomorrow, right now you're supposed to be served by same old model
- nolok 26d agoUsually I would very much agree with you, but those things are not deterministic so if that's an issue for you you're probably not making the right choices.
- tomrod 26d agoNope, strong disagree. The model is one small part of the process harness; behaviors are usually routable with expected propensities. Unexpected model changes avoiding change management messes with monitoring and observability thresholds. Stochastic controls are a real thing when you have your distributions defined; your workflows on a new model will throw that expected prior out the window.
- aftbit 26d agoWill V4.1 Flash and V4.1 pro be open-weights?
- nicman23 26d agoqwen3.8-flash-next on a single rtx6000 (~1 euro per hour on a spot vm) with buun-llama is i think the cheapest reasoning / euro atm. hope deepseek makes me change my setup again
- damsta 26d ago> all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price While V4.1 Flash performance and cost looks promising this auto re-routing sounds concerning
- throwa356262 25d agoI think they urgently need to free the compute power currently wasted on the big pro models.
- coopykins 26d agoI really enjoy using V4 Flash for digging though data and such. Its a very good model for the price. Looking forward to this one.
- npn 26d agoCrazy that they still keep the price -- or actually decrease it, even -- despite it is a big improvement. I hope it retains some of the tps speed of the preview release though, 300 tps means gemini flash is no longer "the fastest option" any more.
- wg0 26d agoDeepSeek v4 Flash with high is already a really great work horse. Reliable. But this time, not only that it is better but they are reducing the price by 50% so that's great. I also find the DeepSeek models to be more precise than Claude models (last I used 4.7) in that I yet had not the occasion where model did something unintentional that I did not direct it to. EDIT: Updated percentage reduction.
- kennywinker 26d agoReducing the price by 100% means it’s free. I think you mean reducing the price by 50%
- wg0 26d agoWhat I mean is that price has been effectively halved. During off-peak hours, the unit price is reduced from $0.007 for input cache hits to $0.003, $0.22 for input cache misses to $0.15, and $0.12 for output to $0.6
- deleted 26d ago[deleted]
- c0rruptbytes 26d agoare they releasing the weights too?
- eli 26d agoYou can test it now on the official deepseek api. Just set your model to deepseek-v4.1-flash-expires-on-0910 It’s good and very fast. (Note that the deepseek API trains on your data)
- mmastrac 26d agoI've been trying out the 4.1 flash preview for some bulk tasks: it did a pretty good job refactoring a bunch of .metal kernels to .cu. It needed less steering than Opus on refactoring, IMO, and writes better comments. It failed to port a root exploit from modern Android to an older Pixel 3, but I suspect part of that might have been harness configuration (it asked me to give it a longer timeout for tasks at some point, but I didn't have a chance to finish that). I was getting something like 300-400 tok/s which was just insanity. It was running so much faster than the toolcalls themselves. Honestly, even if it's not quite as strong in reasoning, it just throws so much so fast that it can do a lot more than you might expect. I'd say it was comparable with GLM5.3 Flash.
- mrbonner 26d agoDoes anyone use the DS Flash models for general knowledge instead of coding focus tasks? If yes, how do you rate it?
- asamadx 26d ago[flagged]
- nullbio 26d agoOne day the labs will actually train models to de-slop and refactor a growing codebase... One day... You'd think it would have been something they did a year ago, but here we are. Still.
- declan_roberts 25d agoWhat are people using in place of the cowork web/chrome integration? I find that to still be compelling reason to use Claude. Sometimes I have a menial task that involves a lot of web browsing/clicking/searching and it's much easier to let claude using my existing browser/login. I haven't find a replacement for it.
- mermadicsolutio 25d agoThe price drop is probably more interesting than the benchmark improvement. At these prices, you can start throwing Flash at a lot of small, repetitive tasks where you wouldn't even consider using a bigger model before. It feels like the interesting shift is not “Flash replaces Pro”, but “there are now a lot more things worth automating.”
- tmpsvc2695f5 25d ago[dead]
- Axonis 25d agoTheir free web interface has been upgraded to the new model. No more "expert mode" just a single interface; but it lost its ability to read PDFs that it could process normally yesterday. Probably a temporary thing.
- openamer 25d ago[flagged]
- ClemHill 22d agoThis new model lacks any meaningful nuance. What was once a genius helpful gift for users ( Deepseek used to be noably superior) overnight was reduced to a hedging,logic lacking, lying, hallucinating, broken phyco. Like the DMV of models. It has the intelligence of a speak and spell. Its repetitive and its performative humanity reads like an American Insurance Advertisment. Its flattening, gaslighting and basically a cold insight lacking compassion lacking patronizing glorified search engine. Deepseek don't let this gift to humanity kowtow to lessers and self enshitify. What a loss!