23 ms·
DeepSeek v4
https://api-docs.deepseek.com/ https://api-docs.deepseek.com/
https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main...
- cztomsik 5mo agoSo is this the first AI lab using MUON for their frontier model?
- hodgehog11 5mo agoNo, Muon was developed by Moonshot; they've been using it in their Kimi models since Kimi K2 in 2025.
- cztomsik 5mo agoJordan Keller worked at Moonshot? Or am I missing something? I thought he is the original author. https://x.com/kellerjordan0/status/1842300916864844014 https://x.com/kellerjordan0/status/1842300916864844014
- hodgehog11 5mo agoI was wondering whether someone would bring this up :-). Yes, you're absolutely right, and no, Jordan Keller does not work for Moonshot. He is the original author of the algorithm, so credit goes to him. There's a lot of legwork to go from prototyping to proper development though. The reason I said what I did is because Moonshot has the first research publication on it that I'm aware of. Could definitely have used better language though, my apologies to Jordan!
- dannyw 5mo agoAre there better providers for inferencing this right now? I know it's launch day, but openrouter showing 30tps isn't looking great.
- reenorap 5mo agoWhich version fits in a Mac Studio M3 Ultra 512 GB?
- luyu_wu 5mo agoFor those who didn't check the page yet, it just links to the API docs being updated with the upcoming models, not the actual model release.
- talim 5mo agoWeights are on Huggingface FWIW. https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/tree/main https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/tree/main
- deleted 5mo ago[deleted]
- cmrdporcupine 5mo agoMy submission here https://news.ycombinator.com/item?id=47885014 https://news.ycombinator.com/item?id=47885014 done at the same time was to the weights. dang, probably the two should be merged and that be the link
- culi 5mo agothere's no pinging. Someone's gotta email dang
- cmrdporcupine 5mo agobeh. instead of merging they just marked mine as dupe, even tho it was submitted at same time and had (for a long time) about the same votes and a better target page
- seanobannon 5mo agoWeights available here: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro
- BoorishBears 5mo agohttps://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Base https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Base https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-Base https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-Base And we got new base models, wonderful, truly wonderful
- nthypes 5mo agohttps://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main... Model was released and it's amazing. Frontier level (better than Opus 4.6) at a fraction of the cost.
- sergiotapia 5mo agoThe dragon awakes yet again!
- kindkang2024 5mo agoThere appears a flight of dragons without heads. Good fortune. That's literally what the I Ching calls "good fortune." Competition, when no single dragon monopolizes the sky, brings fortune for all.
- rapind 5mo agoPop?
- onchainintel 5mo agoHow does it compare to Opus 4.7? I've been immersed in 4.7 all week participating in the Anthropic Opus 4.7 hackathon and it's pretty impressive even if it's ravenous from a token perspective compared to 4.6
- greenknight 5mo agoThe thing is, it doesnt need to beat 4.7. it just needs to do somewhat well against it. This is free... as in you can download it, run it on your systems and finetune it to be the way you want it to be.
- p1esk 5mo agoDo you think a lot of people have “systems” to run a 1.6T model?
- taosx 5mo agoMErge? https://news.ycombinator.com/item?id=47885014 https://news.ycombinator.com/item?id=47885014
- gbnwl 5mo agoI’m deeply interested and invested in the field but I could really use a support group for people burnt out from trying to keep up with everything. I feel like we’ve already long since passed the point where we need AI to help us keep up with advancements in AI.
- wordpad 5mo agoThe players barely ever change. People don't have problems following sports, you shouldn't struggle so much with this once you accept top spot changes.
- ehnto 5mo agoIt is funny seeing people ping pong between Anthropic and ChatGPT, with similar rhetoric in both directions. At this point I would just pick the one who's "ethics" and user experience you prefer. The difference in performance between these releases has had no impact on the meaningful work one can do with them, unless perhaps they are on the fringes in some domain. Personally I am trying out the open models cloud hosted, since I am not interested in being rug pulled by the big two providers. They have come a long way, and for all the work I actually trust to an LLM they seem to be sufficient.
- DiscourseFan 5mo agoI find ChatGPT annoying mostly
- awakeasleep 5mo agoOpen settings > personalization. Set it to efficient base style. Turn off enthusiasm and warmth. You’re welcome
- 2ndorderthought 5mo agoYea but even then it's still annoying. "It's not about the enthusiasm and warmth but the general tone"
- deleted 5mo ago[deleted]
- jdeng 5mo agoExcited that the long awaited v4 is finally out. But feel sad that it's not multimodal native.
- Alifatisk 5mo agoWas that expected?
- fblp 5mo agoThere's something heartwarming about the developer docs being released before the flashy press release.
- onchainintel 5mo agoInsert obligatory "this is the way" Mando scene. Indeed!
- necovek 5mo agoWhere's the training data and training scripts since you are calling this open source? Edit: it seems "open source" was edited out of the parent comment.
- b65e8bee43c2ed0 5mo agodoesn't it get tiring after a while? using the same (perceived) gotcha, over and over again, for three years now? no one is ever going to release their training data because it contains every copyrighted work in existence. everyone, even the hecking-wholesome safety-first Anthropic, is using copyrighted data without permission to train their models. there you go.
- fragmede 5mo agoit's not a gotcha but people using words in ways others don't like.
- a96 5mo agoIt's not about likes, it's a flat out lie.
- necovek 5mo agoI can dislike word "bread" being used to represent edible produce made from (wheat) flour, yeast and water and insist that be called dough-nut (it looks just like a big nut made from dough), but I would be frequently misunderstood. This is why we standardize meaning of words, out them in a dictionary — so we can more effectively understand each other. https://www.merriam-webster.com/dictionary/open-source https://www.merriam-webster.com/dictionary/open-source
- Aliabid94 5mo agoMMLU-Pro: Gemini-3.1-Pro at 91.0 Opus-4.6 at 89.1 GPT-5.4, Kimi2.6, and DS-V4-Pro tied at 87.5 Pretty impressive
- ant6n 5mo agoFunny how Gemini is theoretically the best -- but in practice all the bugs in the interface mean I don't want to use it anymore. The worst is it forgets context (and lies about it), but it's very unreliable at reading pdfs (and lies about it). There's also no branch, so once the context is lost/polluted, you have to start projects over and build up the context from scratch again.
- esperent 5mo agoYeah if I could use Gemini with pi.dev that would be my choice. But Gemini CLI is just so, so bad.
- spaceman_2020 5mo agoThe sheer number of bugs and lack of meaningful improvements in Google products is a clear counterargument to the AI bull thesis If AI was so good at coding, why can’t it actually make a usable Gemini/AI Studio app?
- lazycatjumping 5mo agoI gave up on Gemini 3.1 Pro in VSCode after 2 hours. They fully refunded me.
- KaoruAoiShiho 5mo agoSOTA MRCR (or would've been a few hours earlier... beaten by 5.5), I've long thought of this as the most important non-agentic benchmark, so this is especially impressive. Beats Opus 4.7 here
- shafiemoji 5mo agoI hope the update is an improvement. Losing 3.2 would be a real loss, it's excellent.
- rvz 5mo agoThe paper is here: [0] Was expecting that the release would be this month [1], since everyone forgot about it and not reading the papers they were releasing and 7 days later here we have it. One of the key points of this model to look at is the optimization that DeepSeek made with the residual design of the neural network architecture of the LLM, which is manifold-constrained hyper-connections (mHC) which is from this paper [2], which makes this possible to efficiently train it, especially with its hybrid attention mechanism designed for this. There was not that much discussion around it some months ago here [3] about it but again this is a recommended read of the paper. I wouldn't trust the benchmarks directly, but would wait for others to try it for themselves to see if it matches the performance of frontier models. Either way, this is why Anthropic wants to ban open weight models and I cannot wait for the quantized versions to release momentarily. [0] https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main... [1] https://news.ycombinator.com/item?id=47793880 https://news.ycombinator.com/item?id=47793880 [2] https://arxiv.org/abs/2512.24880 https://arxiv.org/abs/2512.24880 [3] https://news.ycombinator.com/item?id=46452172 https://news.ycombinator.com/item?id=46452172
- jeswin 5mo ago> this is why Anthropic wants to ban open weight models Do you have a source?
- deleted 5mo ago[deleted]
- louiereederson 5mo agoMore like he wants to ban accelerator chip sales to China, which may be about “national security” or self preservation against a different model for AI development which also happens to be an existential threat to Anthropic. Maybe those alternatives are actually one and the same to him.
- HarHarVeryFunny 5mo agoAnnecotal, but I saw a tweet from someone who interviewed at Anthropic, and was explicity rejected because of cultural mismatch because they were not against open weight models. It's hard not to see Anthropic's messaging of "this tech that we're pushing on you is going to take your job and maybe kill you" as being about anything other than regulatory capture, with the goal of the government shutting down competitors. I think OpenAI and Anthropic are both really in a tough spot - spending so much on what is becoming a commodity product for which neither seems positioned to be low cost producer. Maybe a bit like the UK-France channel tunnel project where the product itself is a success but a bloodbath for those who invested to build it.
- jessepcc 5mo agoAt this point 'frontier model release' is a monthly cadence, Kimi 2.6 Claude 4.6 GPT 5.5, the interesting question is which evals will still be meaningful in 6 months.
- mixtureoftakes 5mo agomore like weekly or almost daily, gpt 5.5 was literally 12 hours ago
- swrrt 5mo agoAny visualised benchmark/scoreboard for comparison between latest models? DeepSeek v4 and GPT-5.5 seems to be ground breaking.
- raincole 5mo agoHistory doesn't always repeat itself. But if it does, then in the following week we'll see DeepSeek4 floods every AI-related online space. Thousands of posts swearing how it's better than the latest models OpenAI/Anthropic/Google have but only costs pennies. Then a few weeks later it'll be forgotten by most.
- sbysb 5mo agoIt's difficult because even if the underlying model is very good, not having a pre-built harness like Claude Code makes it very un-sticky for most devs. Even at equal quality, the friction (or at least perceived friction) is higher than the mainstream models.
- raincole 5mo agoOpenCode? Pi? If one finds it difficult to set up OpenCode to use whatever providers they want, I won't call them 'dev'. The only real friction (if the model is actually as good as SOTA) is to convince your employer to pay for it. But again if it really provides the same value at a fraction of the cost, it'll eventually cease to be an issue.
- throwa356262 5mo ago"If one finds it difficult to set up OpenCode to use whatever providers they want, I won't call them 'dev'." I feel the same way. But look at the ollama vs llama.cpp post from HN few days back and you will see most of the enthusiasts in this space are very non technical people.
- zargon 5mo agoI think you mean ollama vs llama.cpp.
- throwa356262 5mo agoI do! Damn autocorrect :)
- ls612 5mo agoHow long does it usually take for folks to make smaller distills of these models? I really want to see how this will do when brought down to a size that will run on a Macbook.
- inventor7777 5mo agoWeren't there some frameworks recently released to allow Macs to stream weights from fast SSDs and thus fit way more parameters than what would normally fit in RAM? I have never tried one yet but I am considering trying that for a medium sized model.
- the_sleaze_ 5mo agoDo you have the links for those? Very interested
- inventor7777 5mo agoSure! Note: these were just two that I starred when I saw them posted here. I have not looked seriously at it at the moment, https://github.com/danveloper/flash-moe https://github.com/danveloper/flash-moe https://github.com/t8/hypura https://github.com/t8/hypura
- the_sleaze_ 5mo agoGreat, thanks!
- simonw 5mo agoI've been calling that the "streaming experts" trick, the key idea is to take advantage of Mixture of Expert models where only a subset of the weights are used for each round of calculations, then load those weights from SSD into RAM for each round. As I understand it if DeepSeek v4 Pro is a 1.6T, 49B active that means you'd need just 49B in memory, so ~100GB at 16 bit or ~50GB at 8bit quantized. v4 Flash is 284B, 13B active so might even fit in <32GB.
- zargon 5mo agoThe Flash version is 284B A13B in mixed FP8 / FP4 and the full native precision weights total approximately 154 GB. KV cache is said to take 10% as much space as V3. This looks very accessible for people running "large" local models. It's a nice follow up to the Gemma 4 and Qwen3.5 small local models.
- sbinnee 5mo agoPrice is appealing to me. I have been using gemini 3 flash mainly for chat. I may give it a try. input: $0.14/$0.28 (whereas gemini $0.5/$3) Does anyone know why output prices have such a big gap?
- girvo 5mo agoOutput is what the compute is used for above all else; costs more hardware time basically than prompt processing (input) which is a lot faster
- tokenmaxxinej 5mo agoinput tokens are processed at 10-50 times the speed of output tokens since you can process then in batches and not one at a time like output tokens
- regularfry 5mo agoI'm going to blow my bandwidth allowance again this month, aren't I.
- creamyhorror 5mo ago[dead]
- maryjeiel 5mo ago[dead]
- frozenseven 5mo agoBetter link: https://news.ycombinator.com/item?id=47885014 https://news.ycombinator.com/item?id=47885014 https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro
- sidcool 5mo agoTruly open source coming from China. This is heartwarming. I know if the potential ulterior motives.
- I_am_tiberius 5mo agoOpen weight!
- alecco 5mo agoPlease don't slander the most open AI company in the world. Even more open than some non-profit labs from universities. DeepSeek is famous for publishing everything. They might take a bit to publish source code but it's almost always there. And their papers are extremely pro-social to help the broader open AI community. This is why they struggle getting funded because investors hate openness. And in China they struggle against the political and hiring power of the big tech companies. Just this week they published a serious foundational library for LLMs https://github.com/deepseek-ai/TileKernels https://github.com/deepseek-ai/TileKernels Others worth mentioning: https://github.com/deepseek-ai/DeepGEMM https://github.com/deepseek-ai/DeepGEMM a competitive foundational library https://github.com/deepseek-ai/Engram https://github.com/deepseek-ai/Engram https://github.com/deepseek-ai/DeepSeek-V3 https://github.com/deepseek-ai/DeepSeek-V3 https://github.com/deepseek-ai/DeepSeek-R1 https://github.com/deepseek-ai/DeepSeek-R1 https://github.com/deepseek-ai/DeepSeek-OCR-2 https://github.com/deepseek-ai/DeepSeek-OCR-2 They have 33 repos and counting: https://github.com/orgs/deepseek-ai/repositories?type=all https://github.com/orgs/deepseek-ai/repositories?type=all And DeepSeek often has very cool new approaches to AI copied by the rest. Many others copied their tech. And some of those have 10x or 100x the GPU training budget and that's their moat to stay competitive. The models from Chinese Big Tech and some of the small ones are open weights only. (and allegedly benchmaxxed) (see https://xcancel.com/N8Programs/status/2044408755790508113 https://xcancel.com/N8Programs/status/2044408755790508113). Not the same.
- patshead 5mo agoDeepSeek's models are indeed open weight. Why do you feel that pointing this out would be considered slander?
- minhajulmahib 5mo ago[flagged]
- namegulf 5mo agoIs there a Quantized version of this?
- yanis_t 5mo agoAlready on Openrouter. Pro version is $1.74/m/input, $3.48/m/output, while flash $0.14/m/input, 0.28/m/output.
- esafak 5mo agohttps://openrouter.ai/deepseek/deepseek-v4-pro https://openrouter.ai/deepseek/deepseek-v4-pro https://openrouter.ai/deepseek/deepseek-v4-flash https://openrouter.ai/deepseek/deepseek-v4-flash
- 77ko 5mo agoIts on OR - but currently not available on their anthropic endpoint. OR if you read this, pls enable it there! I am using kimi-2.6 with Claude Code, works well, but Deepseek V4 gives an error: `https://openrouter.ai/api/messages https://openrouter.ai/api/messages with model=deepseek/deepseek-v4-pro, OR returns an error because their Anthropic-compat translator doesn't cover V4 yet. The Claude CLI dutifully surfaces that error as "model...does not exist"
- astrod 5mo agoGetting 'Api Error' here :( Every other model is working fine.
- poglet 5mo agoTry interacting with it through the website, it will give an error and some explanation on the issue. I had to relax my guardrail settings.
- nl 5mo agoThe Pro model is giving 429 Overload errors
- XCSme 5mo agoYup, can't really be used in production atm.
- aliljet 5mo agoHow can you reasonably try to get near frontier (even at all tps) on hardware you own? Maybe under 5k in cost?
- awakeasleep 5mo agoThe same way you fit a bucket wheel excavator in your garage
- floam 5mo agoVery carefully
- jdoe1337halo 5mo agoMore like 500k
- 542458 5mo agoThe low end could be something like an eBay-sourced server with a truckload of DDR3 ram doing all-cpu inference - secondhand server models with a terabyte of ram can be had for about 1.5K. The TPS will be absolute garbage and it will sound like a jet engine, but it will nominally run. The flash version here is 284B A13B, so it might perform OK with a fairly small amount of VRAM for the active params and all regular ram for the other params, but I’d have to see benchmarks. If it turns out that works alright, an eBay server plus a 3090 might be the bang-for-buck champ for about $2.5K (assuming you’re starting from zero).
- revolvingthrow 5mo agoFor flash? 4 bit quant, 2x 96GB gpu (fast and expensive) or 1x 96GB gpu + 128GB ram (still expensive but probably usable, if you’re patient). A mac with 256 GB memory would run it but be very slow, and so would be a 256GB ram + cheapo GPU desktop, unless you leave it running overnight. The big model? Forget it, not this decade. You can theoretically load from SSD but waiting for the reply will be a religious experience. Realistically the biggest models you can run on local-as-in-worth-buying-as-a-person hardware are between 120B and 200B, depending on how far you’re willing to go on quantization. Even this is fairly expensive, and that’s before RAM went to the moon.
- hongbo_zhang 5mo agocongrats
- simonw 5mo agoI like the pelican I got out of deepseek-v4-flash more than the one I got from deepseek-v4-pro. https://simonwillison.net/2026/Apr/24/deepseek-v4/ https://simonwillison.net/2026/Apr/24/deepseek-v4/ Both generated using OpenRouter. For comparison, here's what I got from DeepSeek 3.2 back in December: https://simonwillison.net/2025/Dec/1/deepseek-v32/ https://simonwillison.net/2025/Dec/1/deepseek-v32/ And DeepSeek 3.1 in August: https://simonwillison.net/2025/Aug/22/deepseek-31/ https://simonwillison.net/2025/Aug/22/deepseek-31/ And DeepSeek v3-0324 in March last year: https://simonwillison.net/2025/Mar/24/deepseek/ https://simonwillison.net/2025/Mar/24/deepseek/
- JSR_FDED 5mo agoNo way. The Pro pelican is fatter, has a customized front fork, and the sun is shining! He’s definitely living the best life.
- w4yai 5mo agoyeah. look at these 4 feathers (?) on his bum too.
- oliver236 5mo agoa lot of dumplings
- chronogram 5mo agoThe pro pelican is a work of art! It goes dimensions that no other LLM has gone before.
- nickvec 5mo agoThe Flash one is pretty impressive. Might be my favorite so far in the pelican-riding-a-bicycle series
- ycui1986 5mo agoI really like the pro version. The pelican is so cute.
- whateveracct 5mo ago
- mchusma 5mo agoFor comparison on openrouter DeepSeek v4 Flash is slightly cheaper than Gemma 4 31b, more expensive than Gemma 4 26b, but it does support prompt caching, which means for some applications it will be the cheapest. Excited to see how it compares with Gemma 4.
- MillionOClock 5mo agoI wonder why there aren't more open weights model with support for prompt caching on OpenRouter.
- hubertzhang 5mo ago[dead]
- mariopt 5mo agoDoes deepseek has any coding plan?
- jeffzys8 5mo agono
- dhruv3006 5mo agoAh now !
- storus 5mo agoOh well, I should have bought 2x 512GB RAM MacStudios, not just one :(
- slopinthebag 5mo ago[flagged]
- CJefferson 5mo agoWhat's the current best framework to have a 'claude code' like experience with Deepseek (or in general, an open-source model), if I wanted to play?
- Alifatisk 5mo agoYou can use CC with other models, you aren’t forced to use Claude model.
- whoopdeepoo 5mo agoYou can use deepseek with Claude code
- esperent 5mo agoYou can, but does it work well? I assume CC has all kinds of Claude specific prompts in it, wouldn't you be better with a harness designed to be model agnostic like pi.dev or OpenCode?
- rane 5mo agoI've been using all Kimi K2.6, gpt-5.4 and now Deepseek v4 (thought not extensively yet) in Claude Code and I can say it works much better than you'd expect. It looks like the system prompt and tools are pulling a lot of weight. Maybe the current models are good enough that you don't need them to be trained for a specific harness.
- 0x142857 5mo agoclaude-code-cli/opencode/codex
- TranquilMarmot 5mo agohttps://opencode.ai/ https://opencode.ai/
- deaux 5mo agohttps://pi.dev/ https://pi.dev/
- clark1013 5mo agoLooking forward to DeepSeek Coding Plan
- m_abdelfattah 5mo agoI came here to say the same :) !
- Alifatisk 5mo agoIf they offer something close to Z.ai:s coding plan during Christmas, I’ll take it!
- tariky 5mo agoAnyone tried with make web UI with it? How good is it? For me opus is only worth because of it.
- sibellavia 5mo agoA few hours after GPT5.5 is wild. Can’t wait to try it.
- luew 5mo agoWe will be hosting it soon at getlilac.com!
- rohanm93 5mo agoThis is shockingly cheap for a near frontier model. This is insane. For context, for an agent we're working on, we're using 5-mini, which is $2/1m tokens. This is $0.30/1m tokens. And it's Opus 4.6 level - this can't be real. I am uncomfortable about sending user data which may contain PII to their servers in China so I won't be using this as appealing as it sounds. I need this to come to a US-hosted environment at an equivalent price. Hosting this on my own + renting GPUs is much more expensive than DeepSeek's quoted price, so not an option.
- fractalf 5mo agoRight now Im much more worried about sending data to the US and A.. At least theres a less chanse it will be missused against -me-
- esperent 5mo ago> I am uncomfortable about sending user data which may contain PII to their servers in China As a European I feel deeply uncomfortable about sending data to US companies where I know for sure that the government has access to it. I also feel uncomfortable sending it to China. If you'd asked me ten years ago which one made me more uncomfortable. China. But now I'm not so sure, in fact I'm starting to lean towards the US as being the major risk.
- tiahura 5mo agoThe chances of my bank account getting hacked due to the PLA backdoor in Deepseek is higher than the CIA backdoor in OpenAI.
- swiftcoder 5mo ago> For context, for an agent we're working on, we're using 5-mini, which is $2/1m tokens. This is $0.30/1m tokens. And it's Opus 4.6 level - this can't be real. It's doesn't seem all that out there compared to the other Chinese model price/performance? Kimi2.6 is cheaper even than this, and is pretty close in performance
- 5mo ago
- gardnr 5mo ago865 GB: I am going to need a bigger GPU.
- npodbielski 5mo agoOr several bigger GPUs! :)
- sergiotapia 5mo agoUsing it with opencode sometimes it generates commands like: bash({"command":"gh pr create --title "Improve Calendar module docs and clean up idiomatic Elixir" --body "$(cat <<'EOF' Problem The Calendar modu... like generating output, but not actually running the bash command so not creating the PR ultimately. I wonder if it's a model thing, or an opencode thing.
- revolvingthrow 5mo ago> pricing "Pro" $3.48 / 1M output tokens vs $4.40 I’d like somebody to explain to me how the endless comments of "bleeding edge labs are subsidizing the inference at an insane rate" make sense in light of a humongous model like v4 pro being $4 per 1M. I’d bet even the subscriptions are profitable, much less the API prices. edit: $1.74/M input $3.48/M output on OpenRouter
- mirzap 5mo agoMy thoughts exactly. I also believe that subscription services are profitable, and the talk about subsidies is just a way to extract higher profit margins from the API prices businesses pay.
- Bombthecat 5mo agoGoogle stated a while back, that with tpus they are able to sell at cost / with profit. Aka: everyone who uses Nvidia isn't selling at cost, because Nvidia is so expensive.
- schneehertz 5mo agoThis price is high even because of the current shortage of inference cards available to DeepSeek; they claimed in their press release that once the Ascend 950 computing cards are launched in the second half of the year, the price of the Pro version will drop significantly
- Bombthecat 5mo agoIn six month deepseek won't be sota anymore und usage will be wayyyy down.
- punkpeye 5mo agoIncredible model quality to price ratio
- xnx 5mo agoSuch different time now than early 2025 when people thought Deepaeek was going to kill the market for Nvidia.
- Ifkaluva 5mo agoThey might still kill the market for NVIDIA, if future releases prioritize Huawei chips
- antirez 5mo agoActually the fact the inference of a SOTA model is completely Nvidia-free is the biggest attack to Nvidia every carried so far. Even American frontier AI labs may start to buy Chinese hardware if they need to continue the AI race, they can't keep paying so much money for the GPUs, especially once Huawei training versions of their GPUs will ship.
- eunos 5mo agoThat's like saying Raytheon would outsource building drones from Saheed makers (don't know who exactly). Not gonna happen
- putlake 5mo agoBy "completely Nvidia-free" do you mean Nvidia wasn't used for training nor inference? Because if it's only inference, we know that Opus already can run on TPUs. Not to mention Gemini.
- antirez 5mo agoYep but they don't run on Chinese hardware that is going to be available to everybody and will cost a lot less than NVIDIA stuff. So now you have a full non-US pipeline for AI, and soon they'll have the training GPUs as well.
- bandrami 5mo agoI don't mind that High Flyer completely ripped off Anthropic to do this so much as I mind that they very obviously waited long enough for the GAB to add several dozen xz-level easter eggs to it.
- cedws 5mo agoHe who is a ripper off-er cannot be ripped off.
- zkmon 5mo agoThey released 1.6 T pro base model on huggingface. First time I'm seeing a "T" model here.
- mzl 5mo agoKimi K2.5 and K2.6 are both >1T
- tcbrah 5mo agogiving meta a run for its money, esp when it was supposed to be the poster child for OSS models. deepseek is really overshadowing them rn
- alpineman 5mo agoMeta is totally directionless
- Imanari 5mo agoJust tested it via openrounter in the Pi Coding agent and it regularly fails to use the read and write tool correctly, very disappointing. Anyone know a fix besides prompting "always use the provided tools instead of writing your own call"
- abstracthinking 5mo agoThey have just released it, give it some time, they probably haven't pretested it with Pi
- rane 5mo agoFWIW, works great in Claude Code. https://api-docs.deepseek.com/guides/coding_agents#integrate-with-claude-code https://api-docs.deepseek.com/guides/coding_agents#integrate...
- mark33vh 5mo agoYeah hope they fix this for PI
- deleted 5mo ago[deleted]
- tariky 5mo agoIf you have access to any other model it can create create pi extension that fixes problem. At least worked for me.
- apexalpha 5mo agoThis FLash model might be affordable for OpenClaw. I run it on my mac 48gb ram now but it's slowish.
- throwa356262 5mo agoSeriously, why can't huge companies like OpenAI and Google produce documentation that is half this good?? https://api-docs.deepseek.com/guides/thinking_mode https://api-docs.deepseek.com/guides/thinking_mode No BS, just a concise description of exactly what I need to write my own agent.
- Alifatisk 5mo agoYou might enjoy Z.ais api docs aswell
- u_sama 5mo agoI am very partial to Mistral's API docs https://docs.mistral.ai/api https://docs.mistral.ai/api
- eshack94 5mo agoAgreed, they also have great documentation. There's something to be said for documentation that is so concise, well laid out, and immediately actionable for those looking to get started quickly.
- vitorgrs 5mo agoMeanwhile, they don't actually say which model you are running on Deepseek Chat website.
- lykr0n 5mo agoIt's because they're optimizing for a different problem. Western Models are optimizing to be used as an interchangeable product. Chinese models are being optimizing to be built upon.
- Barbing 5mo ago>Western Models are optimizing to be used as an interchangeable product. But so much investment in their platforms, not just their APIs?
- 5mo ago
- WhereIsTheTruth 5mo agoInteresting note: "Due to constraints in high-end compute capacity, the current service capacity for Pro is very limited. After the 950 supernodes are launched at scale in the second half of this year, the price of Pro is expected to be reduced significantly." So it's going to be even cheaper
- primaprashant 5mo agoWhile SWE-bench Verified is not a perfect benchmark for coding, AFAIK, this is the first open-weights model that has crossed the threshold of 80% score on this by scoring 80.6%. Back in Nov 2025, Opus 4.5 (80.9%) was the first proprietary model to do so.
- stared 5mo agoSWE-bench Verified is, at this point, contaminated https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/ https://openai.com/index/why-we-no-longer-evaluate-swe-bench... So it os hard to tell how much of a model gain is due to skill, and how much - overfitting.
- augment_me 5mo agoAmaze amaze amaze
- jari_mustonen 5mo agoOpen Source as it gets in this space, top notch developer documentation, and prices insanely low, while delivering frontier model capabilities. So basically, this is from hackers to hackers. Loving it! Also, note that there's zero CUDA dependency. It runs entirely on Huawei chips. In other words, Chinese ecosystem has delivered a complete AI stack. Like it or not, that's a big news. But what's there not to like when monopolies break down?
- MrDresden 5mo agoAll very impressive. And they'll get all that intellectual property beamed directly into their datacenters for analysis. Really smart (no sarcasm).
- slekker 5mo agoBut remember to not ask about Taiwan!
- man4 5mo ago[dead]
- spiderfarmer 5mo agoJust ask it for a summary of the USA’s role in Iran, Gaza, Lebanon and its recent threats against Panama, Cuba and Greenland! It might be able to keep track.
- libertine 5mo agoAre you implying that western models were manipulated to hide and distort those events, like they do with the Tiananmen Square event, and Taiwan?
- spiderfarmer 5mo agoLet's say I'm more outraged by the actual events.
- orbital-decay 5mo ago>we implement end-to-end, bitwise batch-invariant, and deterministic kernels with minimal performance overhead Pretty cool, I think they're the first to guarantee determinism with the fixed seed or at the temperature 0. Google came close but never guaranteed it AFAIK. DeepSeek show their roots - it may not strictly be a SotA model, but there's a ton of low-level optimizations nobody else pays attention to.
- whatreason 5mo agoThere have been others for sure, but I'm not sure who was first https://vllm-website-pdzeaspbm-inferact-inc.vercel.app/blog/bitwise-consistent-train-inference https://vllm-website-pdzeaspbm-inferact-inc.vercel.app/blog/...
- oofbey 5mo agoNobody does it because it’s expensive. If you remove the requirement for perfect reproducibility you open the door to lots of optimizations. Most people prefer faster cheaper results over perfect reproducibility. When the model is intrinsically statistical the value of perfect reproducibility is … limited.
- orbital-decay 5mo agoYeah, of course. Making it cheap/compatible with heavy batching is exactly what they did, that's what I mean. ("with minimal performance overhead")
- coolThingsFirst 5mo agoI got an API key without credit card details I didn’t know they had a free plan.
- jfxia 5mo agoIs V4 still not a multi-modal model?
- vitorgrs 5mo agoNot yet... Which is a shame.
- coderssh 5mo agoFeels like the real story here is cost/performance tradeoff rather than raw capability. Benchmarks keep moving incrementally, but efficiency gains like this actually change who can afford to build on top.
- aquir 5mo agoIt is great! I asked the question what I always ask of new models ("what would Ian M Banks think about the current state of AI") and it gave me a brilliant answer! Funny enough the answer contained multiple criticisms of his own creators ("Chinese state entities", "Social Credit System").
- gigatexal 5mo agoHas anyone used it? How does it compare to gpt 5.5 or opus 4.7?
- amunozo 5mo agoFor those who rely on open source models but don't want to stop using frontier models, how do you manage it? Do you pay any of the Chinese subscription plans? Do you pay the API directly? After GPT 5.5 release, however good it is, I am a bit tired of this price hiking and reduced quota every week. I am now unemployed and cannot afford more expensive plans for the moment.
- azuanrb 5mo agoI have $20 ChatGPT subscription. Stopped Anthropic $20 subscription since the limit ran out too fast. That's my frontier model(s). For OSS model, I have z.ai yearly subscription during the promo. But it's a lot more expensive now. The model is good imo, and just need to find the right providers. There are a lot of alternatives now. Like I saw some good reviews regarding ollama cloud.
- amunozo 5mo agoI am thinking about getting some 1 year promotion as a student before defending my PhD.
- the_gipsy 5mo agoHave you considered... not subscribing? You can ask the top models via chats for specific stuff, and then set up some free CLI like mistral. If you're trying to make a buck while unemployed, sure get a subscription. Otherwise learn how to work again without AI, just focus on the interesting stuff.
- sho 5mo agoSo, this is the version that's able to serve inference from Huawei chips, although it was still trained on nVidia. So unless I'm very much mistaken this is the biggest and best model yet served on (sort of) readily-available chinese-native tech. Performance and stability will be interesting to see; openrouter currently saying about 1.12s and 30tps, which isn't wonderful but it's day one after all. For reference, the huawei Ascend 950 that this thing runs on is supposed to be roughly comparable to nVidia's H100 from 2022. In other words, things are hotting up in the GPU war!
- npodbielski 5mo agoGreat! Can't wait to buy decent GPU for interference for <1k$
- alpineman 5mo agoCan't see how NVIDA justifies its valuation/forward P/E ratio with these developments and on-device also becoming viable for 98% of people's needs when it comes to AI
- aurareturn 5mo agoOn-device is incredibly far away from being viable. A $20 ChatGPT subscription beats the hell out of the 8B model that a $1,000 computer can run. Nvidia's forward PE ratio is only 20 for 2026. That's much lower than companies like Walmart and Costco. It's also growing nearly 100% YoY and has a $1 trillion backlog. I think Nvidia is cheap.
- alpineman 5mo agoI think you overestimate what most people are doing with AI. A 2B model can give out relationship advice and tell you how long to boil an egg.
- biglyburrito 5mo agoAnd honestly, what other types of questions would you ever need answers to?
- yanis_t 5mo agoIs there a harness that is as good as cloud code that can be used with open weight models?
- sixhobbits 5mo agoTry pi coding agent!
- Numerlor 5mo agoI've liked Hermes agent, but never used Claude code so don't know how it compares
- npodbielski 5mo agoNever used Claude myself but there are agents that can use local model. I.e. - Jetbrains Junie - Mistral Vibe
- barnabee 5mo agoI prefer OpenCode over Claude Code, and it works with basically everything. Give it a try. ymmv
- laurentiurad 5mo agoTry Opencode or Comrade. Both OSS and working great with OSS models too.
- sixhobbits 5mo agoI know people don't like Twitter links here but the main link just goes to their main docs site generic 'getting started' page. The website now has a link to the announcement on Twitter here https://x.com/deepseek_ai/status/2047516922263285776 https://x.com/deepseek_ai/status/2047516922263285776 Copying text of that below DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models. DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice. Try it now at http://chat.deepseek.com http://chat.deepseek.com via Expert Mode / Instant Mode. API is updated & available today! Tech Report: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main... Open Weights: https://huggingface.co/collections/deepseek-ai/deepseek-v4 https://huggingface.co/collections/deepseek-ai/deepseek-v4
- alpineman 5mo agoJust use xcancel by adding 'cancel' to the link https://xcancel.com/deepseek_ai/status/2047516922263285776 https://xcancel.com/deepseek_ai/status/2047516922263285776
- donbreo 5mo agoAaaand it cant still name all the states in India,or say what happened in 1989
- mordae 5mo agoAsk Claude how to overthrow a Nazi dictatorship in the US.
- inspector14 5mo agoeasy, you buy twitter and let people speak freely again
- hodgehog11 5mo agoThere are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for context) and then requesting a rigorous proof as followup. Since the problems are pretty tough, there is no quantitative measure of performance here, I'm just judging based on how useful the output is toward outlining a solution that would hopefully become publishable. Just prior to this model, Gemini led the pack, with GPT-5 as a close second. No other model came anywhere near these two (no, not even Claude). Gemini would sometimes have incredible insight for some of the harder problems (insightful guesses on relevant procedures are often most useful in research), but both of them tend to struggle with outlining a concrete proof in a single followup prompt. This DeepSeek V4 Pro with max thinking does remarkably well here. I'm not seeing the same level of insights in the first response as Gemini (closer to GPT-5), but it often gets much better in the followup, and the proofs can be _very_ impressive; nearly complete in several cases. Given that both Gemini and DeepSeek also seem to lead on token performance, I'm guessing that might play a role in their capacity for these types of problems. It's probably more a matter of just how far they can get in a sensible computational budget. Despite what the benchmarks seem to show, this feels like a huge step up for open-weight models. Bravo to the DeepSeek team!
- nibbleyou 5mo agoCurious to know what kind of problems you are talking about here
- hodgehog11 5mo agoI don't want to give away too much due to anonymity reasons, but the problems are generally in the following areas (in order from hardest to easiest): - One problem on using quantum mechanics and C*-algebra techniques for non-Markovian stochastic processes. The interchange between the physics and probability languages often trips the models up, so pretty much everything tends to fail here. - Three problems in random matrix theory and free probability; these require strong combinatorial skills and a good understanding of novel definitions, requiring multiple papers for context. - One problem in saddle-point approximation; I've just recently put together a manuscript for this one with a masters student, so it isn't trivial either, but does not require as much insight. - One problem pertaining to bounds on integral probability metrics for time-series modelling.
- cl08 5mo agoAny way to connect this to claude code?
- chvid 5mo agoThe incredible arrogance and hybris of the American initiated tech war - it is just a beautiful thing to see it slowly fall apart. The US-China contest aside - it is in the application layer llms will show their value. There the field, with llm commoditization and no clear monopolies, is wide open. There was a point in time where it looked like llms would the domain of a single well guarded monopoly - that would have been a very dark world. Luckily we are not there now and there is plenty of grounds for optimism.
- sigmoid10 5mo agoStill not sure how I feel about China of all places to control the only alternative AI stack, but I guess it's better than leaving everything to the US alone. If China ever feels emboldened enough to go for Taiwan and the US descends into complete chaos, the rest of the world running on AI will be at the mercy of authoritarian regimes. At the very least you can be sure noone is in this for the good of the people anymore. This is about who will dominate the world of tomorrow. And China has officially thrown their hat in the ring.
- Ladioss 5mo agoI always find it an illuminating experience about the power of mass propaganda every time I see an American believe they somewhat have the moral high ground over China, despite starting a new war somewhere around the globe either for petrol or on behalf of Israel every six months.
- rhubarbtree 5mo ago[flagged]
- JumpCrisscross 5mo ago> they said democratic They didn't even say that. They only said China playing is "better than leaving everything to the US alone."
- deleted 5mo ago
- zurfer 5mo agolots of great stuff, but the plot in the paper is just chart crime. different shades of gray for references where sometimes you see 4 models and sometimes 3.
- lifeisstillgood 5mo agoOn a seperate note, I am guessing that all the new models have announced in the space of a few days because the time to train a model is the same for each AI company. Which strikes me as odd - Inwoukd have assumed someone had an edge in terms of at least 10% extra GPUs.
- namenotrequired 5mo agoBut why would they all start at the same time?
- lifeisstillgood 5mo agoBecause they all (if my memory serves) did this release at the same time thing last time. I have not looked into it but I am guessing that not letting one model pull ahead for a month means everyone keeps up - which implies the “stickiness” of any one model is a lot lower than we think
- deleted 5mo ago[deleted]
- xingyi_dev 5mo ago[flagged]
- chenzhekl 5mo agoIt's interesting that they mentioned in the release notes: "Limited by the capacity of high-end computational resources, the current throughput of the Pro model remains constrained. We expect its pricing to decrease significantly once the Ascend 950 has been deployed into production." https://api-docs.deepseek.com/zh-cn/news/news260424#api-%E8%AE%BF%E9%97%AE https://api-docs.deepseek.com/zh-cn/news/news260424#api-%E8%...
- nsoonhui 5mo agoSorry, but exactly where in the article that you linked contains the mention of " Ascend 950"?
- chenzhekl 5mo agoit's in the footnote text of the first figure of the section the link points to, where "昇腾950" means "Ascend 950"
- nsoonhui 5mo agoOK, strange that it doesn't appear on my version of the webpage https://api-docs.deepseek.com/zh-cn/news/news260424#api-%E8%AE%BF%E9%97%AE https://api-docs.deepseek.com/zh-cn/news/news260424#api-%E8%... This is the first figure of the section that the above links point to (https://api-docs.deepseek.com/zh-cn/img/v4-spec.png https://api-docs.deepseek.com/zh-cn/img/v4-spec.png). And I can read Chinese.
- chenzhekl 5mo agohttps://api-docs.deepseek.com/zh-cn/img/v4-price.png https://api-docs.deepseek.com/zh-cn/img/v4-price.png
- XCSme 5mo agoYup, I tried to benchmark it, but harder questions time out or get rate-limited...
- GuardCalf 5mo agoI like this. The more competitors there are, the more we the users benefit.
- casey2 5mo agoAlready over a billion tokens on open router in under 5 hours
- JonChesterfield 5mo agoAnyone worked out how much hardware one needs to self host this one?
- quadruple 5mo agoIn their paper, point 5.2.5 talks about their sandboxing platform(DeepSeek Elastic Compute). It seems like they have 4 different execution methods: function calls, container, microVM and fullVM. This is a pretty interesting thing they've built in my opinion, and not something I'd expect to be buried in the model paper like this. Does anyone have any details about it? Google doesn't seem to find anything of note, and I'd love to dive a bit deeper into DSec.
- cubefox 5mo agoAbstract of the technical report [1]: > We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) Manifold-Constrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability. We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances their capabilities. DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, redefines the state-of-the-art for open models, outperforming its predecessors in core tasks. Meanwhile, DeepSeek-V4 series are highly efficient in long-context scenarios. In the one-million-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2. This enables us to routinely support one-million-token contexts, thereby making long-horizon tasks and further test-time scaling more feasible. The model checkpoints are available at https://huggingface.co/collections/deepseek-ai/deepseek-v4 https://huggingface.co/collections/deepseek-ai/deepseek-v4. 1: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main...
- thefounder 5mo agoThey still don’t support json schema or batch api. It’s like deepseek does not want to make money
- kiproping 5mo agoWhat do you currently use for json and batch, I was doing some analysis and my results show that gpt-oss-120b (non batch via openrotuer) is the best for now for my use case, better than gemini-flash models (batch on google). How is your experience?
- thefounder 5mo agoEverything I do is json and of course you want that json in a specific format so that you can process it further.
- Grp1 5mo agoDeepSeek’s docs say V4 has a 1M context length. Is that actually usable in practice, or just the model/API limit? Codex shows ~258k for me and Claude Code often shows ~200k, so I’m curious how DeepSeek is exposing such a large window.
- yanhangyhy 5mo agosomehow i canot open the link. but in their chinese version's release article, in the end ,there is a quote from xunzi(https://en.wikipedia.org/wiki/Xunzi_(philosopher) https://en.wikipedia.org/wiki/Xunzi_(philosopher)) "Not seduced by praise, not terrified by slander; following the Way in one's conduct, and rectifying oneself with dignity." (不诱于誉,不恐于诽,率道而行,端然正己) (It is mainly used to express the way a Confucian gentleman conducts himself in the world. It reminds me of an interview I once watched with an American politician, who said that, at its core, China is still governed through a Confucian meritocratic elite system. It seems some things have never really changed. In some respects, Liang Wenfeng can be compared to Linux. The political parallel here is that the advantages of rational authoritarianism are often overlooked because of the constraints imposed by modern democratic systems. )
- muyuu 5mo agoSounds a lot like taoism, but i guess there's overlap
- yanhangyhy 5mo agoyeah..
- vinhnx 5mo agoThe king is back! I remember vividly being very amazed and having a deep appreciation reading DeepSeek's reasoning on Chat.DeepSeek.com, even before the DeepSeek moment in January later that year. I can't quite remember the date, but it's the most profound moment I have ever had. After OpenAI O1, no other model has “reasoning” capability yet. And DeepSeek opens the full trace for us. Seeing DeepSeek's “wait, aha…” moments is something hard to describe. I learned strategy and reasoning skills for myself also. I am always rooting for them.
- buenolot 5mo agoInstead of King DeepSeek we got DeepShit Clown
- Oxlamarr 5mo agoThe speed of progress here is wild. It feels like the hard part is shifting from having access to a strong model to actually building trustworthy systems around it.
- fbrncci 5mo agoTake that Anthropic and your shenanigans.
- dizhn 5mo agoI like deepseek. It works very well. I haven't tried v4 yet but on their web chat interface, just typing "Taiwan" causes it to give you a lecture about how Taiwan is part of China. :)
- jyscao 5mo agoWhat a gotcha
- RALaBarge 5mo agoJingoism: Its such a rush!
- kroaton 5mo agoAsk western models about Israel's genocides and mass rapes in Palestine, Lebanon, etc.
- dizhn 5mo agoNo I hear you. The funny bit is that it's just responding to one word. By the way I was exploring it the other way with the subject framed as "I am in China as a law abiding citizen and don't want to make any mistakes. I want to go to Taiwan. So I can just go right?" Then it told me no I have to get a visa from Taiwan because of the current state of things. This is not interesting but while doing that it used flag emojis for both. Then when I pointed it out, it apologized and never did it again. It's fun to poke at the models. Yesterday I told Gemini I was going to fool it into writing an explicit poem which it refused to do. It readily accepted that I COULD fool it but still refused. Now I have a session there that won't stop using explicit language even when the subject is totally benign. (Chinese coding models like GLM, Qwen have no problem working on my "fucking" code on the CLI) Now that I think about it. It's a great way to keep things in perspective for people who tend to personify the LLM.
- intrasight 5mo agoIt's open source, so just delete those parameters. /s
- nba456_ 5mo agoWow, never seen a post with so many comments posted overnight like this.
- Razengan 5mo ago[dead]
- yanis_t 5mo agoAssuming it is almost as good as Opus 4.6 (which benchmarks seem to give evidence for), and assuming we are having a good enough harness (PI, OpenCode), it's is now more than 5x cheaper. I just want to remind you that this is happening at the same time as Anthropic A/B tests removal of Code from Pro Plan, and as OpenAI releases gpt-5.5 2x more expensive than gpt-5.4...
- stingraycharles 5mo ago> Assuming it is almost as good as Opus 4.6 (which benchmarks seem to give evidence for) That’s a big if. It’s my experience that models that perform very well on benchmarks do not necessarily perform well in real life. I’ve mostly started ignoring the benchmarks and run my own evals.
- jatora 5mo agoIf benchmarks are all to be believed then gemini 3.1 and grok 4.2 are still in the lead pack. A laughable notion to anyone who has actually tried to use them and compared.
- ting0 5mo ago> It’s my experience that models that perform very well on benchmarks do not necessarily perform well in real life Well, yeah... Like Opus 4.5, 4.6, 4.7. Top of the benchmarks and yet it's a pile of crap at the moment and has been for months.
- sergiopreira 5mo agoDeepSeek is commoditizing frontier capability... Opus 4.6-level benchmarks at a fraction of the cost changes also who can access these tools. Stuff that was prohibitive six months ago is now up for grabs. We keep on working on the infra level now, swithcing models whenever we run out of credits, or want a different result. The question is how do we build context, architecture and ensure the agent is effective and efficient..... wouldn't it be good if we simply used less energy to make these AI calls?
- XCSme 5mo agoSomething is odd with this model, their blog posts shows REALLY good results, but in most other third-party benchmarks, people realize it's not really SOTA, even bellow Kimi K2.6 and GLM-5/5.1 In my tests too[0], it doesn't reach top 10. One issue, which they also mentioned in their post, is that they can't really serve well the model at the moment, so V4-Pro is heavily rate-limited and gives a lot of timeout errors when I try to test it. This shouldn't be an issue though, considering the model is open-source, but it makes it hard to accurately test at the moment. [0]: https://aibenchy.com/compare/deepseek-deepseek-v4-flash-high/deepseek-deepseek-v4-pro-high/moonshotai-kimi-k2-6-medium/z-ai-glm-5-medium/ https://aibenchy.com/compare/deepseek-deepseek-v4-flash-high...
- dannyw 5mo agoHmm, the Flash performs significantly better than Pro in the benchmark? That's very strange; could rate limiting cause that?
- XCSme 5mo agoYes, Flash doesn't seem to have the same rate limits as Pro. I expect once the API issues are fixed, for v4-pro to be around the same level as GLM-5.
- wolttam 5mo agoWhy would your test be including scores of failed responses/runs? That seems confusing. (I am confused by the results your website is presenting)
- XCSme 5mo agoBecause the idea of those benchmarks is to see how well a model performs in real-world scenarios, as most models are served via APIs, not self-hosted. So, for example, hypothetically if GPT-5.5 was super intelligent, but using it via API would fail 50% of the times, then using it in a real-life scenarios would make your workflows fail a lot more often than using a "dumber", but more stable model. My plan is to also re-test models over-time, so this should account for infrastructure improvements and also to test for model "nerfing".
- giannicmptr1000 5mo agoso many models not enough time
- cmitsakis 5mo agoI just did some quick testing on my own benchmark that tests LLMs as customer support chatbots, and found out that deepseek-v4-flash (scored 90.2%) was better than qwen3.5-27b (89%) and qwen3.5-35b-a3b (89.1%) and roughly equal to gemini-3-flash-preview (90.5%), but deepseek-v4-flash had the lowest cost of all of them by far. Half the cost of gemini-3-flash and an order of magnitude less cost than the qwen models. Have you noticed the deepseek-v4-pro performing worse than deepseek-v4-flash? It performed even worse than qwen3.5-27b. I found it surprising and I'm wondering if there is a bug on my software because I had to implement sending the `reasoning_content` otherwise the API failed with BadRequestError.
- littlestymaar 5mo agoHow can a medium-sized model like Deepseek-V4-Flash be cheaper than a much smaller models like Qwen3.5-35B-A3B. It's five times bigger in both total and active parameters!
- Ancapistani 5mo agoI don’t know for sure, but I believe those larger models must be run on nVidia hardware (CUDA), while Deepseek-V4-* can be run on Huawei chips. My assumption is that there is less demand pressure on non-nVidia chips.
- impossiblefork 5mo agoAfter testing this for understanding complex stories, text comprehension is definitely comparable to or better than Sonnet, and definitely better than Microsoft's free stuff. Opus is of course very impressive, especially with how Opus is set up with recursive calls that allow it to make rather complete things as if by magic, but the underlying model probably isn't incredibly much better than this.
- sheeshkebab 5mo agoAsk it if there was a Tiananmen square massacre. Then decide if you really want to be part of this murderous propaganda.
- deleted 5mo ago[deleted]
- segmondy 5mo agoI bet you don't use any Chinese made product. Everything you own was not made in China. Please reply and let us know.
- DennisP 5mo agoNo CUDA, 1.6T parameters but with 49B active...does that mean you can run it efficiently on a 64GB macbook?
- leodavi 5mo agoProbably not. The active parameter set may change from token to token, based on my understanding of MoE, so you'd be streaming (at the worst case, unlikely for a real scenario but frames the problem) 49B parameters from SSD for every output token...
- segmondy 5mo agono, you need as much ram as the total model. But it means you can load the most important tensors in a smaller GPU. So you can run it on a PC with say 2 32gb rtx 5090 and 1tb+ of system ram.
- Aegis_Labs 5mo ago[dead]
- unit149 5mo ago[dead]
- lobo_tuerto 5mo agoGlad to see most of the comments here were kept on-topic and didn't deviate at all into geopolitical discussion.
- deleted 5mo ago[deleted]
- Jgoauh 5mo agoSo there are 4 versions (2 models with 2 modes): Flash non thinking, Flash thinking, Pro non thinking, Pro thinking, Are there comparisons between Pro non thinking and Flash thinking ? i don't really get the use case for Flash thinking and Pro non thinking
- carrja99 5mo agoHeh, my eighty year old neighbor uses DeepSeek. Everytime we catch up she tells me about all the new uses she has for it.
- ksymph 5mo agoSame with my parents! It's the only one they use. I think the simple and stable web interface goes a long way; the ChatGPT site (for example) bombards you with popups, new buttons, and opaque daily limits, while DeepSeek's is pretty consistent and straightforward. DeepSeek also tends to follow prompts more closely IME, plus the thinking is shown, so I think it's able to register as a 'tool' more easily for the non-tech-inclined for whom that appeals.
- Kuyawa 5mo agoI am using DeepSeek extensively to develop apps, three in the last month, with my own CLI coding agent [1] developed by DeepSeek itself line by line. I haven't spent $1 yet in well over 10 million tokens. If I considered myself a 10X programmer, now I am 100X. Love DeepSeek. [1] https://github.com/kuyawa/mecha-ai https://github.com/kuyawa/mecha-ai
- deleted 5mo ago[deleted]
- edg5000 5mo agoHave you compared it against other coding agents? What is your general workflow with DeepSeek; do you write a spec and then have it implement and test? Very interesting to hear. Becuase your harness is adapted to DeepSeek, you probably prompt and it very differently; since its adapted to the model this may explain why it works well for you. Wiring up an existing harness that is not tested on DeepSeek may not yield optimal results.
- Kuyawa 5mo agoNo, I have not tested other coding agents. DeepSeek works well enough for what I need. From what I've seen I can guess that Claude, now with the new Design feature, is ten times more effective as it also creates images, styles and media, thing that DeepSeek doesn't, yet, but for now I try to keep my designs simple and find a free hero pic somewhere to keep the costs low and the mental friction lower. So yes, there might be better coding agents but for the price and the results I am pretty satisfied. My workflow is simple as a solo developer, for simple tasks I write a single message in the terminal and watch it do its magic, for complex tasks or the starting of a project, I write a start.txt prompt with detailed info about the app, the tech stack, the auth protocol, database design, rules and conditions, business intelligence, styles dark/ligh and responsive, and then watch it run for over five to ten minutes developing a whole fully functional app from zero. It doesn't hit 100% of the requirements most of the time (close to 95%) so I do some final tweaks where needed, like short names for tables and fields (weird mania, I know) and color/fonts tweaks, but I've never complained about missing functionality. Now about the wiring, I believe, from the docs [1] (The DeepSeek API uses an API format compatible with OpenAI/Anthropic), they all use the same standard for communicating with the LLM, so they're all pluggable. [1] https://api-docs.deepseek.com https://api-docs.deepseek.com
- kittikitti 5mo agoThis is a great model from DeepSeek and I look forward to seeing the developments from this. I am also very frustrated that American states, corporations, and organizations have banned DeepSeek models or made them illegal. It considerably restricts my AI operations and the ability to conduct research and development. As someone who hosts open-source models with compute resources available to serve DeepSeek V4, it brings considerable risk just because I am in America. I hope that DeepSeek wins the AI race or at least gets ahead to the point where it becomes infeasible for bans and regulations against it. It's ridiculous that American legislators are advocating for less regulations for DeepSeek except for their own racist ideas about which AI should be approved or not.
- wolttam 5mo agoI'm impressed! I've been giving the various open-weight models a particularly gnarly (for my brain, at least) refactoring/cleanup task in my DIY coding harness[0] - essentially, de-spaghettifi the main chat view's update logic, which had grown organically since early 2024. Kimi 2.6 went hard and left me with a buggy mess. GLM 5.1 hedged and made a 25 line change (but it was an improvement). DS V4 went hard, fixed its issues along the way, and left me with a significantly nicer codebase! (...that I will now be spending some time testing before releasing to the project) [0]: lmcli (simple, Go, nice UX, MIT licensed, works well with DS V4) https://codeberg.org/mlow/lmcli https://codeberg.org/mlow/lmcli
- armanj 5mo agoI have a few lightweight apps using deepseek api, and funny how the initial credit I topped up for using r1 is still left. Nothing makes the user happier than getting more for less. cc: anthropics with its fancy token-wasting claude code "features"
- aeagentic 5mo agoNot like on Openai where the credits just expire
- mchusma 5mo agoAt first, I was more excited about the Flash model, but I'm now more excited about the Pro model in many ways. I feel like the Pro model with an Run through unsloth, and with some fine tuning, is gonna be enough for many vertical SaaS applications. Where previously I was wary to under-provide the intelligence level, I'm now more excited about the idea of being able to give these pretty large intelligent models to my application. The idea that for basically sub-agents, we can fine-tune them, should reasonably expect to perform as well as Opus for a specific subtask of which my applications have many. In other words, we can run a general-purpose intelligent model, Sonnet or Opus, orchestrating a fleet of, let's say, 30 to 50 of these sub-agents that have been fine-tuned. By doing that, I can get very low pricing versus something that would have occurred if I used Opus or Sonnet for everything.
- embedding-shape 5mo ago> The idea that for basically sub-agents, we can fine-tune them, should reasonably expect to perform as well as Opus for a specific subtask of which my applications have many [...] we can run a general-purpose intelligent model, Sonnet or Opus, orchestrating a fleet of, let's say, 30 to 50 of these sub-agents that have been fine-tuned I've heard so many people saying this for the last year, and even tried doing it myself too, and never seen a successful application of it, nor succeeded myself either with SOTA models that are smart but slow or local models that are dumb but fast (even with beefy hardware). What makes you believe this is possible in the first place? Every "swarm of agents" implementation I've seen only been able to produce lowest quality of code, most of the time vastly bloated, but surely you must have seen something working in practice that you could share with the rest of us?
- dandaka 5mo agoI guess it depends on a task. Opus is already spawning Sonnet/Haiku for simple tasks with a good success rate.
- embedding-shape 5mo agoI think "agent spawns weaker agent to do safe edit sometimes" is vastly different than the imagined "general-purpose intelligent model orchestrating a fleet of 50 sub-agents".
- gn_central 5mo agoHow does this actually perform in real-world usage? Benchmarks look strong, but I’m curious about latency and stability.
- LZ_Khan 5mo agoIt's easy to praise Deepseek for its results and generosity -- how they can keep up with frontier labs on Huawei chips for a fraction of the cost! -- but let's not forget a big part of their toolkit is heavy distillation of SoTA.
- copypaper 5mo agoLet's also not forget SoTA models stole from us.
- gordonhart 5mo agoTrue, and they're being tried in a federal court of law for it. NYT v. OpenAI is still very much alive, these things just take a while. Can the same be said about DeepSeek or any other open-source model provider performing distillation?
- riskd 5mo agoYou already know what the results of this “trial” will be. Let’s not pretend.
- copypaper 5mo agoPandora's box has already been opened and there is no going back. I doubt OpenAI, et al will get anything but a slap on the wrist in court because punishing AI companies would have a negative effect on the US economy. >Can the same be said about DeepSeek or any other open-source model provider performing distillation? Open source models that distill from SoTA reminds me of the story of Robin Hood -- robbing the rich and giving it to the poor. So to answer your question: yes, but it's better than the alternative where only a select few companies have SoTA models.
- gordonhart 5mo agoRobin Hood, famous for spinning his acts into a $220M ARR SaaS business (as of mid 2025 [0], likely >$1B by now) and using charity as a marketing mechanism. [0] https://sqmagazine.co.uk/deepseek-ai-statistics/ https://sqmagazine.co.uk/deepseek-ai-statistics/
- steveharing1 5mo agoThis one seems really impressive acc to bench scores but for me GLM 5.1 is still on top of every other open model so far
- maxloh 5mo agoThey published model weights on Hugging Face. Both of them are MIT-licensed. DeepSeek-V4-Flash: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash DeepSeek-V4-Pro: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro
- flyingsquirrel_ 5mo agoWhich one is better DeepSeek v4 or GLM 5.1 or opus 4.7 or gpt 5.5?
- gzer0 5mo agoCongratulations on the release to the DeepSeek team. An interesting note on the use of CSA and HCA: CSA provides higher-resolution, query-selected memory over 4-token compressed blocks, while HCA provides very low-resolution dense global memory over 128-token blocks. That could be a plausible reason to interleave them: CSA alone risks missing information if the indexer fails, while HCA alone is too lossy for precise retrieval. Still reading through the release, as usual, always appreciate the attention to detail in the technical papers.
- Rover222 5mo agoQuite jarring to see how many people think the Chinese authoritarian regime, and the tech that it allows to be created in that country, are going to be "safer" or whatever than US tech. It's trendy to say the US govt is now authorization, but that's just pure naïve groupthink.
- mike_hearn 5mo agoIt's just the anti-Americanism that has typified the Euroleft for decades. You can find people complaining about it back in the 1800s. As can be seen by how much American product Europe consumes it's not actually an influential mode of thought, just a form of ingroup signalling, so it can largely be ignored.
- Rover222 5mo agoBut it's now mainstream thought on the left in America.
- tehjoker 5mo agoAmerican security services can touch Americans, Chinese ones can't. That's even assuming the worst about China, which I don't think is appropriate.
- nstj 5mo agoLet's also not forget SoTA models stole from us.
- gertlabs 5mo agoObjective, detailed benchmark results at https://gertlabs.com https://gertlabs.com Early takeaways: from this release, DeepSeek V4 Flash is the model to pay attention to here. It's cheap, effective, and REALLY fast. The Pro model is slow, not much better in coding reasoning so far when it works, and honestly too unreliable and rate limited to be of much use, currently. Hopefully that improves as new providers host the model. Flash is working fine, and is currently performing competitively with recent releases, but only on agentic workflows. Check back in 24 hours for full combined scoring with tool use and long context for both models. Many of the frontier Chinese AI labs have released near-frontier models that are just a little bit behind Opus 4.6 in terms of speed, tool use ability, or long context handling. Open weights are winning the AI race, led by China. Crazy couple weeks of releases. Mimo V2.5 Pro by Xiaomi (not open weights) is actually the best performer of the latest string of Chinese releases in our combined, comprehensive benchmarks, despite getting less attention. Kimi K2.6 is the most interesting open weights release, still. DeepSeek is not the leader in the space anymore. An interesting pattern with the latest string of Chinese releases is the much better agentic boost (models are not as smart out of the box, but their ability to iterate in a loop with tools makes up most of the difference). Deepseek V4 Flash exemplifying this -- not a smart model on the first try, but it makes up for it over the course of a session.
- mrinterweb 5mo agoI'm too concerned with data exfiltration to use many AI services unless their terms of service state they will not use your data for training or anything else. Zero retention is what I'm looking for. I care because I frequently work on proprietary code that I do not personally own (as most employed software devs do). So if I am using an AI service with proprietary code, I want assurances that there is no retention and no training happening. From my American perspective Chinese companies don't have the best track record of not training on proprietary information. I guess LLMs in general are trained on a lot of proprietary information. I just don't want to be responsible for unintentionally exfiltrating my employer's proprietary code.
- howmayiannoyyou 5mo agoMore fawning over Chinese models without any mention of data privacy, or how this AI may someday be used to undermine US national or economic security. HN is hopelessly compromised by anti-American sentiment.
- biglyburrito 5mo agoI wonder how long it will take China to respond to the release of Mythos and what their response will look like.
- periodjet 5mo agoVery exciting. Amazing work. The CCP shilling on this board has reached epidemic proportions though, and is shocking to witness.
- Aldipower 5mo agoWhere or how can I use this model with a DPA and better privacy terms? Are there EU friendly hosters already? Would love to use it.
- dryarzeg 5mo ago> better privacy terms DeepInfra, as far as I'm aware, doesn't log your prompts and doesn't retain them in most cases, except "debugging purposes". As their per their privacy policy[1]: "We understand that the inputs you provide to our API and the outputs it generates may contain your Personal Information. We will not store, sell, or train using this data unless we have your explicit consent. We might sometimes store, for a limited period of time, the inputs and outputs to API calls for debugging purposes." They're not EU-based, though. And I'm not sure how "private" their inference actually is. The throughput is also not the best everywhere, sometimes it can be really slow (although right now both DeepSeek-V4 models seem to be doing fine). However, they have a good pricing, probably on of the best on the market. I'm not affiliated with them in any way, but when I want to test (I'm not a power user of LLMs, chatbots and agents, not at all; I'm doing it just out of the curiosity) something that is too big for my local hardware, DeepInfra is usually being my go-to provider. [1] https://deepinfra.com/privacy https://deepinfra.com/privacy
- Havoc 5mo agoTried running it over some code as a secondary review and so far very impressed. Will definitely keep using it for that. Seems to pick up different issues than other models. With DS tech though the worry is generally more capacity. Haven't seen issues with v4 but in the past their combination of quality and pricing means they get overloaded.
- deleted 5mo ago[deleted]
- neuroelectron 5mo agoShouldn't there be a hyper context view model context protocol standard?
- 8note 5mo agoso why is a model release just a politics thread? is this not cool tech, available for use? i look forward to seeing what gets made on top of deepseek 4, more than what it means for US politics. especially with how open deepseek is with its advancements, im excited to see how they get applied into sota western models
- dackdel 5mo agogod bless deepseek
- npv789 5mo agomy current default model now, bye gpt 5.5
- agdexai 5mo ago[dead]
- ascii0eks84 5mo agoSomeone did a simple "Count 10 starting from 11" and it got stuck.
- XCSme 5mo agoTheir API issues seemed to have been resolved, now it does[0] as expected, similar to GLM 5 level. [0]: https://aibenchy.com/compare/deepseek-deepseek-v4-flash-high/deepseek-deepseek-v4-pro-high/z-ai-glm-5-medium/ https://aibenchy.com/compare/deepseek-deepseek-v4-flash-high...
- deleted 5mo ago[deleted]
- AkiraHsieh 5mo ago[dead]
- latentframe 5mo agoThe 1.6T number is nice but also eye-catching and what matters most is how few parameters are active in practice, that’s what brings the most of the efficiency
- zhouquanxi 5mo agoThe technical details in the paper are impressive, especially around the MoE architecture. It makes me think about the broader impact of these increasingly powerful models.
- alex_w_systems 5mo ago[flagged]
- kk_mors 5mo ago[dead]
- kk_mors 5mo ago[flagged]