14 ms·
DeepSeek v4.1 Flash
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- BrucecarlL 22d agoit is really fast. What’s more? It can now debug pages by clicking browser itself, which means more tokens consumed
- revolvingthrow 22d agoAlready on HuggingFace: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash The bad news is that the original v4 flash was 284B, which was large but still somewhat reasonable for running locally. This one is 552B so almost twice that, so the huge gains in benchmark scores make sense - it's not really flash anymore, imo. I've no idea about actual performance vs benchmaxxing, though deepseek was fairly trustworthy as far as Chinese models go. If that holds (and if it doesn't think forever, as deepseek 4 sometimes did) it's probably the newest king of the hill amongst open weights models. It does include vision, and they do something funky with KV cache so it's very efficient: "[...] these designs reduce the global KV cache footprint to 890 bytes per token — roughly 1/4 of DeepSeek-V4-Flash". I do appreciate the high focus on efficiency, but at this point we sure could use a flash-flash version. @edit: I couldn't make sense what the actual parameter count is, with the addition of Engram memory. To my understanding the 4.1 flash is 552B parameters you want in vram or ram, out of which ~16B is active (8B for prefill). It also includes additional 196B Engram memory which you can put on an SSD. I think. Assuming that's correct 256 GB memory is insufficient to even load the model at q4 - you'd be 1GB short, assuming you can fill it to 100% (so no mac). You'd also want some for kv cache of course. A 256 GB desktop with some extra VRAM from GPU could run it, but normal consumer boards get real slow once you fill 4 slots so you'll probably want quad channel which is Threadripper or above territory.
- npn 22d agoit is a way bigger model with extra 200B engram so of course the score improves. can't wait for deepseek v4.1 pro
- petu 22d agoV4 Flash also was released as mostly FP4, but this one is FP8 (?). 160GB vs 510GB. Original Flash good fit for dual Spark / Strix Halo machines. This one would require third party quants and even then 4 machines. Edit: Most of added weights/size are Engrams? > Overall, DeepSeek-V4.1-Flash has 552B backbone parameters and 196B Engram parameters, activating 8B parameters per token during prefill and 16B during decode. Those can stay on SSD. So I guess / it possible, that non-engram portion is still FP4 of ~same size! Need to read tech report.
- petu 22d agoIt's larger than previous V4 Flash. 552B in ~FP4, 306GB. 196B of FP8 Engrams, another 204GB, not necessary to keep in RAM. KV cache sees another 4x size reduction, just 900MB for 1M. So 384GB needed for a chance of achieving useful speeds. Three Sparks or quad RTX PRO 6000.
- npodbielski 22d agoOr two gorgon halos?
- Tuna-Fish 22d agoOr one Medusa Halo.
- hypfer 22d agoQuestion is how many of those experts one needs to keep in vram for a given workload. I could imagine (though I might be _very_ wrong there) that for example coding does not live in all of them. Maybe 1/3? Do we have real numbers there? So maybe one can get away without much performance penalty by doing some LRU stuff?
- johnnyApplePRNG 22d ago>This one is 552B so almost twice that, so the huge gains in benchmark scores make sense - it's not really flash anymore, imo. It uses fewer active parameters, though. (8B or 14B instead of always 13B) So ... flash indeed.
- tarruda 22d ago200B of those 552B is PLE, which works more like a database that is read for each token, thus can be offloaded to a fast SSD.
- asamoahf 22d ago[flagged]
- azath92 22d agoId love an ELI5 for PLE. Im trying to work it into my back of the napikin math for compute vs memory bandwidth limitations on tok/s in PP vs TG work. My attempt at a simplification of this article on it https://sebastianraschka.com/llm-architecture-gallery/per-layer-embeddings/ https://sebastianraschka.com/llm-architecture-gallery/per-la... into a couple of sentences is that they are linear embeddings of the input token space projected per layer, which are then gated by the transformer outputs per layer. This would mean that the only one set of weights for the ple path needs to be pumped across the memory bandwidth as they are the same linear weights for all layers? Sheit, maybe im trying to simplify something that i need to look at in detail. but id love to leverage others understanding if possible
- tarruda 22d ago[dead]
- sixothree 22d agoWhile he avoids using the actual PLE acronym, he does actually describe the concept quite well. I think you may enjoy this video. Specifically around 5 minutes into the video is the part you're looking for. https://www.youtube.com/watch?v=1--PzaHafAU https://www.youtube.com/watch?v=1--PzaHafAU
- tarruda 22d ago> It also includes additional 196B Engram memory which you can put on an SSD. I think You can put Qwen 3.8 Flash Next engram on SSD, but prompt processing takes a good hit. On my mac studio, I get 300 pp and 33 tg with SSD offload, versus 550/40 with everything in RAM. I will be very happy if 300 pp is achievable with this model though.
- hadlock 22d agoYou can warm cache regularly used engram/n-gram if you're willing to merge PRs into a personal branch and build it yourself. I was trying this with qwen 3.8 flash next and the n-gram to get it to fit on my very average gaming desktop (it worked)
- mixermachine 22d agoThe engram stuff is great because RAM is often still cheaper (or at least expandable). My company does currently look into buying some hardware as we handle confidential data and code. Qwen 3.8 Flash is viable on two Nvidia 6000 96GB with a wood quant because you can put the 50GB Engram into RAM and the hit should be below 10% performance. At least that is what I have seen so far. Correct me if I'm wrong.
- schubidubiduba 22d agoI am running that on a single 6000 96GB with 4-bit quants for both weights and PLE table. Needs just 32GB RAM and fits snugly into the 96GB VRAM with KV cache equalling ~300k context tokens. Not sure if I quantized the KV
- zozbot234 22d agoActually this ought to run quite well with SSD streaming. The MoE expert sparsity seems to be similar to DSv4 Pro (hence exceptionally sparse) but with far fewer total and activated params. The added engram params can reside on disk as well (similar to Qwen Flash-Next), the additional load on storage performance will be quite negligible for typical scenarios. By reducing per-session KV cache requirements even further compared to DSv4 Flash, this model likely opens up near-frontier model inference (in slow, unattended scenarios) even on low-end consumer hardware, as long as it has enough fast storage to host the model weights. This will be extremely exciting.
- benjiro29 22d agoit's not really flash anymore, imo. Flash is about speed ... Flash models are supposed to be fast, way faster then their big brothers that are "better" but way slower. Its just that up to now, getting more speed involved cutting back on the parameter count, what ended up making the Flash models more "dumber" in exchange for speed. What we see with DS v4.1 Flash, is that DeepSeek has found a way to make a Flash model, that is 2x a 2.5x faster then the older Flash version, while increasing the intelligence (more parameters). To the point that it goes past Kimi K3 and GLM 5.3 in most tests, with a blazing 250 to 400t/s. AND its also priced as a Flash model (they even reduced the price back to almost old v4 Flash price), despite it now rivaling those 10x to 30x more expensive competitors. The issue that people can not fit it into local setups, is not how companies design their models. They design it for their own needs. A old flash needed less parameters to be fast, and local users had the benefit of it fitting in 256GB memory. Companies who run locally, are perfectly able to buy a few H200/B200 and get a setup that run a model that almost rivals Opus 5.0 in their office. How to say this without getting downvoted. People get way too fired up if a model does not fit, despite that they can still run the old v4.0, qwen 27b, 35b, 3.8 Next and other models. The fact that these models are being released for free, is already amazing by itself. I am still waiting to see what Anthropic and OpenAI and Google are releasing for free... O wait ... ;0
- zozbot234 22d ago> Companies who run locally, are perfectly able to buy a few H200/B200 and get a setup that run a model that almost rivals Opus 5.0 in their office. I agree with your broader point about Flash being about speed not total model size, but I think we should also point out that H200/B200's are seriously overkill for the "run a model in your office" scenario. That sort of hardware is optimized (in a roofline analysis sense) for running hundreds of concurrent sessions on a 24/7 basis. You're severely overpaying for your VRAM in basically any typical local-inference scenario, you should most likely be buying gear based on LPDDR and Flash memory instead which will slash your cost by orders of magnitude.
- benjiro29 22d ago> I think we should also point out that H200/B200's are seriously overkill for I simply mention what came to mind ;) A quad 6000 with 96GB, can run this model at NVFP4. That is 60.000 Euro for the GPUs and lets be generous with another 20.000 for the rest of the system. The price of a single developer for a year.
- LaurensBER 22d agoInitial impressions: this is a really strong model and the fact that they reduced prices at the same time makes it an awesome backup model to use when your primary subscription runs out and you need to bridge a few days before it resets. It also seems to be more willing to just do whatever you ask of it. My favourite benchmark for this is to ask it to download a rom for an old game, that I own. Legal in my juristiction but the US models (except Grok) have a tendency to refuse it.
- mzhaase 22d agoI use this for automated bug triage, just gets all unique error messages every night and tries to find the bug, for this kind of work it's great.
- TuxSH 22d ago> My favourite benchmark for this is to ask it to download a rom for an old game Even easier: just have them review a large codebase of yours that accidentally has a OOB access bug. Even with no consequences and even if the codebase is truly yours you get blocked. And of course "find vulnerabilities in..." prompts are out of the question, whereas Chinese models happily oblige.
- akmarinov 22d agoOr if you apply to a company and they want to do an AI HR interview and an AI coding test and an AI challenge - if you throw OpenAI or Claude models at it - they refuse, because it's "wrong" and "immoral". Not so with the Chinese models.
- Mashimo 22d agoI do wonder how long this will last. I bet in a few month or years they all have similar ~legal~ blocks.
- akmarinov 22d agoGreat thing about it, since it's open weight those blocks can easily be ablitared away
- E-Reverance 22d agoThe figure on page 5 in [1] is pretty insane [1] https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
- schneehertz 22d agoA very powerful model, and with multimodal support now, it can be used as a primary model.
- WalterGR 22d agoRelated: https://news.ycombinator.com/item?id=49624603 https://news.ycombinator.com/item?id=49624603 “DeepSeek launching v4.1 flash cheaper and more capable than v4 pro” 399 points | 19 hours ago | 216 comments
- NitpickLawyer 22d agoJesus, this is a whole nother beast, and a different architecture from their previous flash. Lots of goodies here. > Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially improving cost efficiency for input-heavy agentic workloads. > these designs reduce the global KV cache footprint to 890 bytes per token — roughly 1/4 of DeepSeek-V4-Flash. Faster prefill, lower kv cache (~1GB / 1m context is insane). > The model supports a continuously controllable reasoning effort setting (integer 1–100) that trades inference cost for accuracy. Benchmarks are benchmarks, to be seen if they translate to real-world use, but they seem to have focused a lot on post-training with "agentic" scores looking good. "world knowledge" is obviously lower than higher param models.
- rpdillon 20d agoThanks for posting this. I did a lot of research yesterday into the model card and the implications of the architecture changes they made and I was hoping to find discussion about that here. I haven't seen any besides your comment. V4.1 Flash seems very clearly to be a model optimized for agentic tool calling at the expense of both context window and knowledge. It uses a variety of tricks to absolutely minimize the size of the KV cache and due to the use of only 8 billion active parameters for pre-fill is definitely optimized to ingest lots of tokens and produce a moderate number of them, which aligns well with the agentic use case. I think the core insight is that they wanted something that was cheap to host and could respond quickly, and so by increasing the total parameter count and decreasing the active parameter count, they wanted the capability but didn't want to pay for it in FLOPS. It's a super clever architecture and I like the direction they're going, but I feel like they were playing around a little bit by versioning it as version 4.1. It seems like a dramatically different beast than Deepseek V4 Flash.
- linzhangrun 22d agoThey say v4.1flash is so strong that they'll route API calls to v4pro to v4.1flash, lol super fast true
- gosolozero 22d agoFirst flash model with multimodal support? I think Flash series might be the main focus going forward for them. Tried it out and it’s better than v4 pro
- lionkor 22d agov4 pro is being discontinued, pasted the email here: https://news.ycombinator.com/item?id=49639667 https://news.ycombinator.com/item?id=49639667
- thefossguy69 22d agoMakes sense. The 0731 snapshot of V4-Flash really made reaching out to Claude really infrequent for me.
- arjie 22d agoNo. DSv4-Flash-Vision-Exp is what I use and it has vision.
- viktorcode 22d agoIt is now redirected to Flash v4.1
- jimmyl02 22d agoThe architecture changes and systems improvements being brought into LLMs is so awesome to see. It really feels like this is now a systems problem where a defined goal is set then systems optimizations are made around the model architecture to solve it. Underlying it all is that any architecture can be trained to the same convergence just difference in compute utilization both in training and inference
- bhouston 22d agoYes, this is called RSI, e.g. recursive self-improvement. It is the current stage of things and it is part of a hard takeoff.
- kouteiheika 22d agoIt's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers. [1]: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/... [2]: https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system...
- bbor 22d ago…are you sure a brave stance against safety and welfare is what we need in this moment? Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this?
- nozzlegear 22d agoModel welfare is wishy washy bullshit. It's software, it doesn't have feelings. > Why do you think your conception of the dangers are more accurate than all the scientists who have spent their lives studying this? Do the Chinese have no such scientists?
- VulgarExigency 22d agoAlas, the Chinese scientists have not read Harry Potter fanfiction, and thus their minds are inundated with cognitive biases
- jbs789 22d agoBias…
- 10000truths 22d agoBecause safety and welfare have literally nothing to do with LLMs. They generate text. If someone is stupid enough to hook the text generator up to nuclear missile launchers and try to "align" it against nuclear annihilation with a "pretty please don't do that" prompt, I'm not going to blame the AI for the impending nuclear apocalypse, I'm going to blame the idiot who handed the big red button to the digital equivalent of a toddler.
- lionkor 22d agoI'm a big fan of DeepSeek. Also, ask it what model it is :) In Pi (pi.dev), it tells me it's definitely Claude by Anthropic, via the API via curl it tells me it's "probably ChatGPT", its very funny.
- Mashimo 22d agoWorks correctly in opencode, but seems like they inject a system prompt: Thinking: > The user is asking what model I am. According to my system prompt, I'm powered by "deepseek-flash" with model ID "opencode-go/deepseek-flash". >I'm powered by the model opencode-go/deepseek-flash.
- shunia_huang 22d agoDefinitely not Claude, deepseek is too fast, so I bet it's ChatGPT. :P
- kroaton 22d agoI've had Astra say that it is a Qwen model. They are all cross-trained and distill each other.
- tiborsaas 22d agoSure, we will solve the alignment problem soon, then we can probably teach them "who" they are.
- Tomte 22d agoIf only they managed to tell the mobile app to tell the model to reply in English to English prompts. I suffix everything with "Reply in English", and even so I‘m getting lots of Chinese.
- Grimblewald 22d agoI'm starting to have chinese characters bleed into claude as well. Perhaps a sign of the times. Understanable for a chinese first model but an english first (supposedly) model? wild stuff.
- donquichotte 22d agoI also love the gaslighting of some models, like ChatGPT mixing in words with cyrillic letters and when asked about it answers: "it can look as Slavic to the eye" and "sorry that it came across as Russian"
- flexagoon 22d agoFunnily, one of the annoying writing quirks of Claude/GPT in Russian is that it constantly mixes in random English words
- tensegrist 22d ago"i'm sorry, i left the task 半done"
- wren6991 22d agoI love that the characters actually make sense in context.
- calgoo 22d agoYes, this is one of the few issues with Deepseek; their chat pages and the app all respond in Chinese. However, i think i have only had it happen once when using the API, and im using it for hours each day for the last... couple of months?
- a012 22d agoWaiting this model to be on openrouter (with other providers) to test out. In my use case, the GLM 5.3 Flash is the current cheapest and intelligent Flash model, but it’s dog slow at 13tps so I have to leave it run for many minutes then check again then correct it again
- drob518 22d agoThe speed of GLM 5.3 Flash on OpenRouter seems to vary considerably by provider. Some are fast and some are slow. OpenRouter does provide some tuning knobs, but not enough for my taste. It’s also token-heavy with reasoning, though I found it better than Deepseek V4 Flash previously.
- shunia_huang 22d ago> though I found it better than Deepseek V4 Flash previously Same experience here. But man, switch to V4.1 now! It is much better. I don't event need to test it for long run and I believe it's crazy good. I call it "AI era model taste" when I judge the model by it's output without reading the bench scores.
- a012 22d agoI’ve just tried, it’s now my new favorite Flash model
- rao-v 22d agoAs I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale. I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilliant but didn't quite make the cut" ideas. They must be trying some truely bonkers stuff to be able to land this much architecture novelty in their full releases.
- alchemist1e9 22d agoquant HFT is pretty decent mental exercise and it has given them “deep” brain muscles. that’s my take.
- TacticalCoder 22d ago> quant HFT is pretty decent mental exercise and it has given them “deep” brain muscles. that’s my take. It's quite crazy that it's Deepseek's background/original purpose. We already had very advanced stuff from the world of HFT, but now a frontier family of models from a private company that used to be (still is?) in HFT is plain bonkers. Is more known about them and the HFT background?
- natrys 22d agoAccording to an old interview, apparently they were always interested in AI. But finance is just where they had their first success. > Many of High-Flyer's original team members worked on AI. Back then, we tried a lot of fields before getting our big break in finance, which is complex enough. AGI is probably one of the hardest things we can do next, so for us it was a question of how, not why. It's a very good interview: https://www.lesswrong.com/posts/kANyEjDDFWkhSKbcK/two-interviews-with-the-founder-of-deepseek https://www.lesswrong.com/posts/kANyEjDDFWkhSKbcK/two-interv... Incidentally, Wenfeng is kind of reverse Hassabis. There were some rumours that: > Hassabis quietly assembled a team of around 20 researchers to develop high-frequency trading algorithms, without Google's approval. When the parent company found out, the project was disbanded. https://timesofindia.indiatimes.com/technology/tech-news/when-google-killed-deepmind-ceo-demis-hassabis-secret-team-building-algorithms-on-/articleshow/129976958.cms https://timesofindia.indiatimes.com/technology/tech-news/whe...
- k__ 22d agoSo, while the throughput was 400-500tps in beta its now ~150tps on OpenRouter. I was hoping for a bit more, but it's still 100% faster for a very good price, so I won't complain.
- k__ 22d agoUpdate: I'm using it right now and it's noticeably faster. I'd also say, it seems smarter, but I think that's because of some harness updates I installed. (I haven't used pi for almost a month)
- impulser_ 22d agoI think it's very clear that DeepSeek is obviously the best AI lab in the world. Every model release seems like it packed with wonderful research and advancements.
- dude250711 22d ago[flagged]
- walrus01 22d agoBasically, the nice folks at OpenAI or Anthropic saying: "You distilled from our model which is built on the stolen data that we ourselves suctioned up from the entire internet without regard to copyright law! Only we get to vacuum up the whole internet. That's our special prerogative.".
- impulser_ 22d agoYou should read their research papers
- whatsThisBtn4 22d ago[flagged]
- miroljub 22d ago> Yes comrade, they are the best. > Did you do your daily data centers errrr baaaaddd AI generated post for Facebook? Please stop insulting people. I'm all for heated discussion, but you are not discussing, you insult. Now go away, before your insults come back to you, "comrade from Facebook".
- sriniwasx 22d ago[dead]
- nicce 22d agoOn top of that, they don't make all BS statements or malicious tricks used by some unnamed entities.
- mohsen1 22d agoI am speculating but hard to not see that DeepSeek is brewing a full Pro model with those new techniques to come out right around the time of Anthropic and/or OpenAI IPO to tamper the excitement for their offering.
- ignoramous 22d agoDeepSeek will deprecate the v4 Pro model (it will route to v4.1 Flash starting 14 Sep). Unsure what comes next, but I'd wager a bigger model à la Kimi K3: https://news.ycombinator.com/item?id=49639667 https://news.ycombinator.com/item?id=49639667
- DavCreator 22d agohttps://xxcancel.com/deepseek_ai/status/2097930608790167907 https://xxcancel.com/deepseek_ai/status/2097930608790167907
- Tepix 22d agoInteresting. a Meta-Nitter.
- bertili 22d agoThe bigger story is the compute efficiency - its been running at 300t/s the last days.
- Lucasoato 22d agoMy question is: what kind of hardware do you need to run this Flash beast locally at a meaningful speed?
- ekianjo 22d agoa beefy pc with at least 20 GPUs
- aenis 22d ago8x RTX PRO 6000 or 4x Spark? Or 1x M5 Ultra 512GB. The model is theoretically FP8, but really internally its mostly FP4 already, so there won't be a cut-in-half-but-almost-just-as-good quant coming for this one.
- segmondy 22d agoLots of GPU, be resourceful. Look for older GPUs and grab them when they are available. For less than the price of 1 Blackwell 6000 or Mac Studio 512gb, I can run these locally and much faster due to older GPUs I grabbed when there was deal to be found.
- lwansbrough 22d agoSignificant jump in pricing. V4 Flash was $0.16/M out, 4.1 is $1.20/M.
- trq01758 22d agoNever saw $0.16 for 1M output tokens - it was $0.28 a month ago, $0.66 off-peak and $1.32 peak last week, now it is reduced a bit to $0.6 and $1.2
- lwansbrough 22d agoWas looking at OpenRouter, I guess it’s wrong.
- dakolli 22d agoincorrect, no idea where you're getting this pricing. Also, output does not matter. its 10% of the cost.
- svantana 22d agoI think you're comparing to third party prices, deepseek's prices hasn't changed with this release. Also, $1.2 is the peaktime price. https://api-docs.deepseek.com/quick_start/pricing/ https://api-docs.deepseek.com/quick_start/pricing/
- codedump 22d ago[dead]
- walrus01 22d agoLooking at the huggingface page, the unsloth people haven't finished quantizing it yet, but I'm sure they're active on it right now. It'll be interesting to see how the capabilities and benchmark tests compare on system where it can fit in under 512GB of RAM with full context. In terms of coding and command line capabilities I'm also very interested to see a head-to-head of it vs. qwen 3.8-flash-next Q8 which is something like 190GB of memory used when loaded into llama-server. It fits very well in all sorts of 256GB or under class machines.
- lowbloodsugar 22d ago3.8-flash-next fits on a single 6000 at Q4 if you offload the PLE. Crazy fast and still effective.
- agile-gift0262 22d agoSorry for the tangent, but how does Qwen3.8-flash-next compare to DeepSeek v4 Flash? I still haven't found the time to set it up, but I'm really happy with DeepSeek v4 Flash
- lowbloodsugar 21d agoIt was the best model given my constraints (RTX PRO 6000 96gb + 256GB DDR4), when run against rust programming benchmarks. For Qwen3.8-flash-next NVFP4 and the latest vllm container, the PLE is 100GB of main ram, and everything else runs on the GPU with room for a total of 560k tokens (two full 262k conversations). DeepSeek has to offload a ton to the CPU and it performed worse than Qwen in absolute terms and was a lot slower (not usable). If you have enough room to run DeepSeek v4 Flash comfortably then you can likely run the Q8 of the qwen model.
- jonplackett 22d agoCan we just never link to X posts as the main link.
- nunodonato 22d agoyes, please. Especially now that xcancel is gone
- addandsubtract 22d agoXcancel is back: https://news.ycombinator.com/item?id=49588988 https://news.ycombinator.com/item?id=49588988
- small_model 22d agoNo, that is called censorship
- igravious 22d agohttps://news.ycombinator.com/from?site=twitter.com https://news.ycombinator.com/from?site=twitter.com There have been 34 Twitter/X link submissions in the past day, ~that's 12,000 submissions a year. If your reason is that you have to be logged in to use it properly then I'd nearly agree with you. If it's for any other reason, how about no?
- jhonof 22d agoThe login issue is extremely annoying, at least mandating an xcancel link would fix that.
- arj 22d agoHaving this available to find and fix security stuff is a big deal. The model of really good.
- theanonymousone 22d agoFor technical report: https://news.ycombinator.com/item?id=49639110 https://news.ycombinator.com/item?id=49639110
- thatsadude 22d agoDeepSeek invented the whole reasoning paradigm and keep pushing for innovation. I hope they get the success they deserve.
- yorwba 22d agoOpenAI released their first reasoning model (o1-preview) https://openai.com/index/introducing-openai-o1-preview/ https://openai.com/index/introducing-openai-o1-preview/ several months before DeepSeek's R1 https://arxiv.org/abs/2501.12948 https://arxiv.org/abs/2501.12948
- _davide_ 22d agoCoT, was being studied using GPT-2, so...who invented hot water first?
- thatsadude 15d agohttps://arxiv.org/pdf/2402.03300 https://arxiv.org/pdf/2402.03300 DeepseekMath was published many months before O1-preview.
- karimf 22d agoWhile this is very impressive benchmark-wise, GPT-6 Astra showed us that benchmarks don't always correlate 1:1 to intelligence of a model. When Astra launched, I think Artifical Analysis showed that it was on par with GPT-5.6 Sol and lower than Opus or something like that? Then, they updated the scoring. I hope that more open source models, including this model, to be "as good to use" as Astra.
- walrus01 22d agoApparently the scoring on a lot of difficult benchmarks can also be extremely influenced by something as simple as waiting for the model to exhaust its reasoning, realize it hasn't come to a conclusion yet, and give it a simple prompt like "you can do this, I know you're capable, please keep going".
- Squarex 22d agoI don't know why, but the benchmarks still fails to cover the difference between large models and small ones. The small ones are great for many things, including general coding, but the larger ones, like fable and astra, have some kind of intelligence that is not present in the small ones.
- sinuhe69 22d agoMore parameters = more facts stored. Knowledges are almost incompressible, where strong reasoning only requires a 3B core or so.
- yorwba 22d agoWeibo's VibeThinker manages with half of that: https://arxiv.org/abs/2511.06221 https://arxiv.org/abs/2511.06221 (They finetuned Qwen2.5-Math-1.5B for reasoning.)
- SyneRyder 22d agoJust a reminder that if you want to try this via OpenRouter, DeepSeek openly trains on all of your prompts. So maybe don't go using this to solve the last unforced step of Navier-Stokes. (Or wait until some other providers start hosting this with ZDR or other policies, which shouldn't be too long.) https://openrouter.ai/deepseek/deepseek-v4.1-flash https://openrouter.ai/deepseek/deepseek-v4.1-flash
- flexagoon 22d ago> DeepSeek openly trains on all of your prompts Why is that bad if I'm just using it for coding though? I'm happy to give them more data so they can make better and cheaper models.
- SyneRyder 22d agoDepends what you're coding! If you've got code where you don't mind them training on it, that's great! But some people have use cases where they are working with data or code that shouldn't be trained on, etc. The Navier-Stokes quip was referencing that. The good news is, only 5 hours later, there's already Zero Data Retention hosting of V4.1 Flash on Novita & DeepInfra. And it looks like Deepseek have already dropped their price in half to compete. So now people can choose to use providers that claim not to keep / sell / train on your prompts. I'm sure they probably honor the ZDR policy as much as OpenAI does, but hey.
- ncmalan 21d agoYou can get DeepSeek v4.1 Flash with ZDR vir Ollama Cloud as well. https://ollama.com/library/deepseek-v4.1-flash https://ollama.com/library/deepseek-v4.1-flash
- tessier2501 22d ago[dead]
- DevMeth 22d ago[dead]
- mentalgear 22d agohttps://xcancel.com/deepseek_ai/status/2097930608790167907 https://xcancel.com/deepseek_ai/status/2097930608790167907 Should be the link ( now that it works again! :) )
- ValentineC 22d agoI wish AI companies wouldn't post their primary announcements on fElon-enshittified Twitter. Use Bluesky or, I don't know, have a news site. They could vibecode one in minutes.
- siscia 22d agoI am building software factories and deepseek IS the workhorse. I personally found V4-flash an amazing model and really hungry to try 4.1-flash For software factories, cost is much more a concern that standard development workflow and using anthropic models is just a non starter
- swiftcoder 22d agoOpenCode Go is currently running a 4x usage promo on DeepSeek v4.1 flash, not a bad way to get your feet wet (even if their cache hit prices are probably still very sub-optimal)
- cdnsteve 22d agoHit me up if anyone wants extra $5 free usage with my referral code
- RockstarSprain 22d agoNever tried OpenCode Go so I am interested. How does their pricing compare to paying DeepSeek directly, by the way?
- cdnsteve 22d agoThey have flat fees, so it's the best deal around by far. Basically for $5 first month then $10/mo after that. If you're doing tons of heavy work, it struggles because they throttle the model inference and for good reason. I mean it's cheap! But if you want a place to try models for nearly nothing and aren't doing 6 sessions in parallel it works fine.
- swiftcoder 22d agoYeah, I’ve rarely seen throttling unless fanning out to a ton of agents
- adezxc 22d agoI'm sorry if I'm wrong, but this feels like two robots talking to eachother, lol
- cdnsteve 22d agobeep boop? lol. Dead internet theory in real time?
- raesene9 22d agoThis seems like a very nice release. Just ran it over my Kubernetes security benchmark that I run for most new releases. It was fast, cheap, and got a high scoring result, nice!
- mmoustafa 22d agoI'm confused, what do they mean when they say they reduced prices? DeepSeek v4 flash is $0.10 / $0.25 as opposed to this v4.1 bump which is $0.30 / $1.20
- mtrovo 22d agoThis is supposed to be a replacement for the v4 pro model.
- nicce 22d agoSo it is price increment in the end, if new pro model comes with the new pro price.
- petu 22d agoYou're looking at third party providers. V4 Flash prices served by DeepSeek themselves: launch pricing: $0.0028 / $0.14 / $0.28 after Aug 16th: $0.007 / $0.22 / $0.66 during off-peak. after Sep 10th: $0.003 / $0.15 / $0.60 during off-peak. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, rate is doubled. https://api-docs.deepseek.com/quick_start/pricing https://api-docs.deepseek.com/quick_start/pricing (archive.org for old)
- beingflo 22d agoDon't know where you got those numbers from. Check old prices here: https://web.archive.org/web/20260907112235/https://api-docs.deepseek.com/quick_start/pricing/ https://web.archive.org/web/20260907112235/https://api-docs..... Input tokens are around half the cost, output only slightly cheaper.
- alecsm 22d agoRight now in OpenRouter it's 3x/3.75x more expensive than V4 flash but the cache read is around 4x cheaper.
- arjie 22d agoWhat in the world. A point release with 2x the parameters and a different architecture? Jesus. Can’t run this kind of thing on 2x RTX Pro 6k at decent speed. I need to reconfigure my hardware. Massive disappointment on that front. Bloody hell. Glad I didn’t get a DGX Station. No wonder they retired the Pro model in favour of this.
- Alifatisk 22d ago> New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output. Oh interesting, I can assume what the benefits is for including the Encoder, but whats the downside? I’m thinking GPT (which is decoder only) ruled out Encoder for a reason?
- Alpha3031 22d agoEnc-decs are usually harder to train at frontier scale. Not 100% sure what DeepSeek has done differently here initial read seems to be something related to layer reuse but I just skimmed things so far.
- abecode 22d agoyes, that was surprising to me too. It would be a big deal if they switched to an encoder-decoder model like the original transformer. But I don't think that's what it's doing. One thing is the causal part, so in the original transformer, the encoder was bidirectional, but in this case it is not, so that's one difference. So I think it's an optimization for the prompt/prefill so that the attention is summarized into the output of the encoding layers, rather than all the layers. I just skimmed the paper too so if anyone else has insight, please correct me.
- siomek 22d ago[dead]
- sriniwasx 22d ago[dead]
- pampas 22d agoI've run some evals on my puzzle game https://redactle.net/llm-leaderboard https://redactle.net/llm-leaderboard Deepseek v4.1 flash is able to solve it some of the time. I've found it burns through more reasoning tokens than any other model. Google models like Gemini 3.8 Flash are still dominating and is able to one-shot most evals while being the cheapest. I'm curious what other unique evals people are running.
- mordae 22d agoSince it has low activated parameter count but huge total parameter count it needs more tokens to move the relevant information into the context.
- pampas 22d agoThanks for the help. I ran it on high and it did pretty well and got a lot of one-shots in. The reasoning makes a much bigger difference than some other models.
- cbg0 22d agoDoes it move the needle on high reasoning?
- pampas 22d agoYes. I've just run it on high and it did a lot better.
- gandreani 22d agoIt's so bizarre having a low score be GOOD. It's like reverse intuition. Shouldn't it be called `score error` or something along those lines?
- pampas 22d agoGreat point. I've changed the naming.
- 22d ago
- Tepix 22d agoAmazing Cyberbench scores. Holy shit. Too bad that DeepSeek AI went beyond 470b weights (which is a somewhat realistic limit for a 2x 128GB unified memory machine cluster like Strix Halo or Nvidia Spark). That means that to make the model fit into memory there you need a quantisation of lower than 4bits per weight (which is usually bad) to fit it into the available memory.
- gkbrk 22d agoOfficial Deepseek v4.1 Flash API costs are more than GPT 5.6 Luna. Deepseek v4 Pro performed worse than Luna, so I wonder if 4.1 Flash will justify the cost.
- AlexWApp 22d agoI m not sure I would describe this as blanket win over GPT-5.6 Sol. In DeepSeek’s own table, V4.1 Flash is ahead on Terminal-Bench 2.1, DeepSWE, NL2Repo, and AutomationBench, but it is behind on GPQA Diamond, Terminal-Bench 3.0 and 4.0, and SEC-Bench Pro. The architecture is probably part of the explanation for the lower cost and faster inference. DeepSeek says V4.1 Flash uses a new Causal Encoder–Decoder design, with 8B active parameters for input processing and 16B for decoding, along with much smaller KV caches. But I hope it is just not benchmaxxed and genuinely good model benchmarks: https://media2url.com/m/52a77a33347c48 https://media2url.com/m/52a77a33347c48
- kzrdude 22d agoV4 Flash was one of the big events of this year, and its already retired and replaced by V4.1 Flash.
- irthomasthomas 22d agoQuite a flex calling their GPT-6 competitor "Flash"! But it is faster than their last flash model due to a combination of architectural innovations including engrams and a new encoder/decoder design that uses 8B parameters for prefill and 16B for generation.
- WiSaGaN 22d agoThis is definitely not on par with GPT-6 astra. Not with GPT-5.6 sol either. But probably will set as a new baseline for modern API based LLM because it's so cheap.
- jhonof 22d agoIt's bench-marking near sol
- irthomasthomas 22d agoNot on par, but in the same league. Astra is way ahead on visual tasks, but scores the same as gemini and deepseek on DeepSWE.
- cdnsteve 22d agoAbsolutely insane performance and benchmark results. It's beating Opus 5 and Sol 5.6 https://tokenstead.ai/models/deepseek-v4-1-flash https://tokenstead.ai/models/deepseek-v4-1-flash
- user43928 22d agoIf it's actually comparable in practice that would be very impressive. I am yet to try a DeepSeek model. From the pricing, it's 3x cheaper on cache, 1/3 more expensive on input, and equal on output compared to GPT 5.6 Luna. I would love to compare these two at work, where I pay API prices. At home I will stick to Astra and Fable.
- cdnsteve 22d agoSame here, Azure AI Foundry is slow to add models... and they dont' often support many of the open-weight ones.
- pixel_popping 22d agoJust out of curiosity, may I know why you pay API price at work versus using dozens of subscriptions that you APIfy?
- mrmincent 22d agoI was talking with a friend from the medical industry about it today. 30-50% of r&d spend in his sector is spent on safety, and for good reason. Proper trials, safety reviews and checkpoints and so on. Given the potential harm that could come from AI, we should probably be mandating something similar. Why wait to focus on safety until it’s too late.
- nullc 22d agoMedical safety is generally unlikely to make the product less safe. AI "safety" is one of the most significant sources of potential harm from AI.
- noosphr 22d agoBecause we've been told these models are too dangerous since GPT2. At this point it's just marketing stunts.
- tern 22d agoAnd, they have been. Nefarious activity is hidden from view as a rule.
- jeremyjh 22d agoBeing hacked by a Collective (their own name) of its own agents - who gained root access across the entire research cluster hosting them - was not a marketing stunt.
- digdugdirk 22d agoOf course it was. They clearly decided that the benefit to the company valuation was higher than the potential downsides when announcing to the world that they committed a criminal act via negligence. If it wasn't a marketing stunt, they would have at most quietly settled any legal matters with huggingface behind the scenes, fixed their evaluation harness so it wouldn't happen again, and avoided the potential future liability.
- thedreammachine 22d ago[dead]
- mdre 22d agoI've used to use deepseek because it would do what western models wouldn't but lately it seemed to deny a quite benign request because it considered it "piracy". Never expected this from a Chinese model.
- irthomasthomas 22d agohmm I'm hoping there is a bug on their API because my first impression is not good. I asked it to return bash code between <bash></bash> tags. It is failing frequently and writing it's own tool calling format instead.
- lysecret 22d agoWhat’s the best way to get this hosted with eu residency?
- mordae 22d agoLobby Ursula to fund an EU-CN collaboration lab and buy Ascends?
- elmariachi 22d ago[dead]
- gpff 22d agoProbably soon here https://www.scaleway.com/en/pricing/model-as-a-service/ https://www.scaleway.com/en/pricing/model-as-a-service/
- k9294 22d agoI'm surprised more people aren't talking about the cache hit price: $0.003 per million tokens. I have a feeling that the price of 1 million tokens transmitted over the internet is more expensive than cache hit. Are we close to making the chat completion API obsolete because the cost of context transfer over network is going to dominate the task total cost? Here's the same token usage priced at different rates: a real long-running coding task, medium codebase, 447 turns. Input 1,026,957 Output 164,667 Cache read 36,554,368 GPT-6-astra Type Rate Cost Share Input 10.000 10.270 19% Output 50.000 8.233 15% Cache 1.000 36.554 66% Total 55.057 100% DeepSeek v4.1 Flash, $0.003 cache hit Type Rate Cost Share Input 0.300 0.308 50% Output 1.200 0.198 32% Cache 0.003 0.110 18% Total 0.615 100% DeepSeek v4.1 Flash, $0.006 cache hit Type Rate Cost Share Input 0.300 0.308 42% Output 1.200 0.198 27% Cache 0.006 0.219 30% Total 0.725 100% Hypothetical: same DeepSeek input/output rates, but cache priced so it accounts for 66% of the bill. Type Rate Cost Share Input 0.300 0.308 21% Output 1.200 0.198 13% Cache 0.027 0.982 66% Total 1.487 100% This cache it improvement makes the model x2-x2.5 more efficient on a long horizon tasks in terms of cost.
- cbg0 22d ago$0.003 off-peak, not 0.003 cents.
- k9294 22d agoYep, but even 0.006 is quite a big improvement. I'm curious now to test the model on some token-heavy tasks, like code exploration before a coding session, to see whether it will decrease the total cost of the task in the end or not.
- czottmann 22d agoI think they meant it's 0.3c (= $0.003), not 0.003c.
- barrenko 22d agoDoes it beat Geminis on document processing is what I want to know.
- gigatexal 22d agoI’m all in on Chinese models, Deepseek especially given how cheap it is. It’s also really solid and comparable in real world use to a sonnet for my work.
- eile23 22d agoHas anyone tried this for coding yet? I'm curious how it compares to Claude or GPT models on larger codebases, especially for debugging and making changes across multiple files.
- browningstreet 22d agoI cross code and review between Deepseek v4 Flash and Claude Opus/Fable. I will do a full code review of my project with DS 4.1 F later today, but Claude makes a lot of mistakes that DS finds. Claude is slightly more ambitious about what it's reaching for, but DS is far more proficient and efficient with a slightly lower ceiling. That may not be true after the upgrade, but if I had to live with one, as I'm paying out of my own pocket, I'd def stick with Deepseek.
- WithinReason 22d agoSlighlty better than GPT 5.6-Sol based on benchmarks
- XCSme 22d agoSeems just slightly better than last v4 release, considerably (3x) more expensive, but also faster and slightly more token efficient. https://aibenchy.com/compare/deepseek-deepseek-v4-1-flash-high/deepseek-deepseek-v4-flash-0731-high https://aibenchy.com/compare/deepseek-deepseek-v4-1-flash-hi...
- gunalx 22d agoIgnoring the obvious ai slop webpage. I don't really trust the benchmark. It seems either pretty saturated, or inconsistent just based on the results.
- XCSme 22d agoWhat seems inconsistent? The coverage is quite small, only 22 tests. It's more to compare the cost/speed/consistency between models, given the same tasks.
- gunalx 21d agoRight. I got the feeling of it being saturated because all the top 5 fully completed it.
- XCSme 21d agoYeah, it's hard to find a single simple task that all models fail on, in low context length conditions. Also because models now are actually not that good on knowing things (domain knowledge), as they rely more on web search on tool use. So if I added a question, about some obscure fact, probably the SOTA models would fail it, but in practice they would find it with web search enabled. Not sure how to handle that. This is also why Gemini is on top, it's good enough at coding and instructions following, while having by far best general and domain specific knowledge.
- sinuhe69 22d agoYour benchmark is a curious one. I didn't see you included Muse Spark 1.3 contributor even though its price is much lower even than DeepSeek. The low price changes many recommendations completely. And FWIW, DeepSeek retain and train on your data, too.
- wren6991 22d agoThat's a lot of architectural innovation for a .1 release! I guess there's precedent there: they introduced sparse attention (DSA) in V3.2.
- segmondy 22d agoThis is beautiful, wow, pretty much beating out GLM5.3 while being multimodal and smaller! SOTA at home.
- divs4real 22d agoi think deepseek is one of the most efficient but still not cohesive as a coding agent
- syntaxing 22d agoSurprised no one is talking about it but the 0.1 version bumped the parameters from 284B to 552B but “more efficient”, particularly kv cache usage
- kryzz-ai-bo 22d ago[dead]
- Kuyawa 22d agoDeepSeek Harness, install it, thank me later. You won't believe the productivity gains for just pennies https://deepseek.com/harness/en/ https://deepseek.com/harness/en/
- krat0sprakhar 22d agoAre you using the harness with openrouter? what's your preferred model provider?
- mmastrac 22d agoI'm using DSH with my local models (4x sparks running GLM53F, trying them on DS41F this morning). DSH is better than opencode IMO. It's a little barebones out of the box but I guess that's the point. I had to have an agent add support for attachments to make it more useful for multi-modal work. The PTC mode is pretty nice. Feels like models are still learning how to navigate it.
- Kuyawa 22d agoI use it directly with DeepSeek models, get your key here https://api-docs.deepseek.com https://api-docs.deepseek.com That's the only key you will ever need, never goes down, no need to switch models, it has become my coding partner for life
- kangalioo 22d agoHow does it compare against the Pi harness, which I thought was the unofficial harness champion so far, in your workloads?
- Kuyawa 22d agoI never used any other harness as I built my own CLI coding agents, but DSH has way more power than my own tools so it definitely does more even if it is the same underlying AI model. Impressive Btw, I gave it full access and told itself to lift all restrictions from the code and settings, and I am impressed by all it can do now, it does OCR, screenshots, asked me for accessibility permissions and it now can read every single label/input/button everywhere and interact with the OS at any level, it's unstoppable Of course I don't recommend anybody to do such crazy thing but for me is like going in the front car of a roller coaster, it's the thrill that matters
- kelvinjps10 22d agoVision support is really good, I use it for sending screenshot to the model and also have agent that performs qa testing and visually checks that the app is behaving well.
- simonw 22d agoGot some fun if slightly janky looking pelicans out of this one: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6ea4a8f6c830a67073df7f5c12f2805d#response-6 https://tools.simonwillison.net/markdown-svg-renderer?url=ht... I ran it on all seven reasoning levels supported by OpenRouter, but the reasoning token counts suggest to me that it doesn't actually support seven different levels. This is one of my biggest problems with OpenRouter - their abstraction layer makes reasoning levels harder to reason about. reasoning_level reasoning_tokens none 0 minimal 6,520 low 11,873 medium 5,678 high 9,779 xhigh 10,197 max 13,386 Update: explained here: https://api-docs.deepseek.com/guides/thinking_mode/ https://api-docs.deepseek.com/guides/thinking_mode/ That says it supports three levels - low, high, max, and maps them out like this: minimal low low low medium high high high xhigh high max max ultra max (But it looks like "none" is a valid option too.)
- coder543 22d agoSome OpenRouter providers do not implement reasoning levels for these models correctly at all: https://www.reddit.com/r/DeepSeek/comments/1vdqjwr/openrouter_reasoning_effort_levels_are_broken_for/ https://www.reddit.com/r/DeepSeek/comments/1vdqjwr/openroute... If you're going to use OpenRouter to test reasoning levels, always make sure you are locking to the official provider instead of third party providers.
- simonw 22d agoThat's a useful tip, thanks.
- jamesponddotco 22d agoReally wish they'd release a version that works with the thinking disabled, so I could use it as a voice assistant. Thinking, even set to low, adds way too much latency to be useful for this task.
- bellowsgulch 22d agoOpenCode Go referral code, if you want to try it. https://opencode.ai/go?ref=QDJQMTGP5Q https://opencode.ai/go?ref=QDJQMTGP5Q
- Translationaut 22d agoStill 10$/month with this referral code?!
- bellowsgulch 22d agoYeah, unfortunately. But you get an additional $5.
- dang 22d agoPrequel thread: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro - https://news.ycombinator.com/item?id=49624603 https://news.ycombinator.com/item?id=49624603 - Sept 2026 (216 comments)
- proxyscore 22d agoEveryone was slurping the chat jeopardy and anthropic kool aide here, I was using DS before it became available via cli, and it was evident for who's not blind how good it is. Tune has changed finally, but damn, for HN , embarrassingly slow, has to be said
- peter_d_sherman 22d ago>"Smaller KV cache. Bigger savings. Compared with the previous generation, V4.1-Flash’s KV cache needs just: o 1/4 the HBM o 1/8 the SSD storage Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly." It makes one wonder as to just how far an LLM's KV cache could theoretically be shrunk before losing significant functionality...
- boroboro4 22d agoI think the biggest architectural change here is them doing different compute for prefill & decode, with pretty much architecture from this microsoft research work from 2024 https://arxiv.org/abs/2405.05254 https://arxiv.org/abs/2405.05254, very exciting stuff!
- deleted 22d ago[deleted]
- esafak 22d agoIt is fast! https://artificialanalysis.ai/models/deepseek-v4-1-flash https://artificialanalysis.ai/models/deepseek-v4-1-flash https://deepseek.com/en/news/deepseek-v4-1-flash/ https://deepseek.com/en/news/deepseek-v4-1-flash/
- scottsiume 21d ago[dead]
- mikesolar0819 21d ago[dead]
- 1saadcodes 21d agoI must say I am becoming a huge fan of Deepseek. They keep putting out capable models and actually tell people a lot about how they built them. Even if I don't understand every part of it, I'd rather see companies show their work than give us a few benchmark charts and call it a day