17 ms·
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf]
https://huggingface.co/deepseek-ai/DeepSeek-V3.2 https://huggingface.co/deepseek-ai/DeepSeek-V3.2
https://api-docs.deepseek.com/news/news251201 https://api-docs.deepseek.com/news/news251201
- arthurcolle 10mo agoSurely OpenAI will follow up with a gpt-oss-780b
- nimchimpsky 10mo agoPretty amazing that a relatively small Chinese hedge fund can build AI better than almost anyone.
- BoorishBears 10mo ago3.2-Exp came out in September: this is 3.2, along with a special checkpoint (DeepSeek-V3.2-Speciale) for deep reasoning that they're claiming surpasses GPT-5 and matches Gemini 3.0 https://x.com/deepseek_ai/status/1995452641430651132 https://x.com/deepseek_ai/status/1995452641430651132
- deaux 10mo agoThe assumption here is that 3.2 (without suffix) is an evolution of 3.2-Exp rather than being the same model, but they don't seem to be explicitly stating anywhere whether they're actually different or that they just made the same model GA.
- zparky 10mo agoBenchmarks are super impressive, as usual. Interesting to note in table 3 of the paper (p. 15), DS-Speciale is 1st or 2nd in accuracy in all tests, but has much higher token output (50% more, or 3.5x vs gemini 3 in the codeforces test!).
- futureshock 10mo agoThe higher token output is not by accident. Certain kinds of logical reasoning problems are solved by longer thinking output. Thinking chain output is usually kept to a reasonable length to limit latency and cost, but if pure benchmark performance is the goal you can crank that up to the max until the point of diminishing returns. DeepSeek being 30x cheaper than Gemini means there’s little downside to max out the thinking time. It’s been shown that you can further scale this by running many solution attempts in parallel with max thinking then using a model to choose a final answer, so increasing reasoning performance by increasing inference compute has a pretty high ceiling.
- jodleif 10mo agoI genuinely do not understand the evaluations of the US AI industry. The chinese models are so close and far cheaper
- beastman82 10mo agoThen you should short the market
- newyankee 10mo agoYet tbh if the US industry had not moved ahead and created the race with FOMO it would not had been easier for Chinese strategy to work either. The nature of the race may change as yet though, and I am unsure if the devil is in the details, as in very specific edge cases that will work only with frontier models ?
- jazzyjackson 10mo agoValuation is not based on what they have done but what they might do, I agree tho it's investment made with very little insight into Chinese research. I guess it's counting on deepseek being banned and all computers in America refusing to run open software by the year 2030 /snark
- bilbo0s 10mo ago>I guess it's counting on deepseek being banned And the people making the bets are in a position to make sure the banning happens. The US government system being what it is. Not that our leaders need any incentive to ban Chinese tech in this space. Just pointing out that it's not necessarily a "bet". "Bet" imply you don't know the outcome and you have no influence over the outcome. Even "investment" implies you don't know the outcome. I'm not sure that's the case with these people?
- coliveira 10mo agoExactly. "Business investment" these days means that the people involved will have at least some amount of power to determine the winning results.
- TIPSIO 10mo agoIt's awesome that stuff like this is open source, but even if you have a basement rig with 4 NVIDIA GeForce RTX 5090 graphic cards ($15-20k machine), can it even run with any reasonable context window that isn't like a crawling 10/tps? Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further
- bigyabai 10mo agoPeople with basement rigs generally aren't the target audience for these gigantic models. You'd get much better results out of an MoE model like Qwen3's A3B/A22B weights, if you're running a homelab setup.
- Spivak 10mo agoYeah I think the advantage of OSS models is that you can get your pick of providers and aren't locked into just Anthropic or just OpenAI.
- hnfong 10mo agoReproducibility of results are also important in some cases. There are consumer-ish hardware that can run large models like DeepSeek 3.x slowly. If you're using LLMs for a specific purpose that is well-served by a particular model, you don't want to risk AI companies deprecating it in a couple months and push you to a newer model (that may or may not work better in your situation). And even if the AI service providers nominally use the same model, you might have cases where reproducibility requires you use the same inference software or even hardware to maintain high reproducibility of the results. If you're just using OpenAI or Anthropic you just don't get that level of control.
- Aachen 10mo agoWho is the target audience of these free releases? I don't mind free and open information sharing but I have wondered what's in it for the people that spent unholy amounts of energy on scraping, developing, and training
- 10mo ago
- red2awn 10mo agoWorth noting this is not only good on benchmarks, but significantly more efficient at inference https://x.com/_thomasip/status/1995489087386771851 https://x.com/_thomasip/status/1995489087386771851
- deleted 10mo ago[deleted]
- ode 10mo agoDo we know why?
- hammeiam 10mo agoSparse Attention, it's the highlight of this model as per the paper
- pylotlight 10mo agoI'll have to wait for the bycloud video on this one :P
- culi 10mo agoHow did we come to the place that the most transparent and open models are now coming out of China—freely sharing their research and source code—while all the American ones are fully locked down
- pennomi 10mo agoOver reliance on investors who demand profits more than engineering. The best innovation always happens before being tainted by investment.
- victor9000 10mo agoBecause the entire US economy is being propped up by AI hype.
- zug_zug 10mo agoWell props to them for continuing to improve, winning on cost-effectiveness, and continuing to publicly share their improvements. Hard not to root for them as a force to prevent an AI corporate monopoly/duopoly.
- srameshc 10mo agoAs much I agree with your sentiment, but I doubt the intention is singular.
- echelon 10mo agoI don't care if this kills Google and OpenAI. I hope it does, though I'm doubtful because distribution is important. You can't beat "ChatGPT" as a brand in laypeople's minds (unless perhaps you give them a massive "Temu: Shop Like A Billionaire" commercial campaign). Closed source AI is almost by design morphing into an industrial, infrastructure-heavy rocket science that commoners can't keep up with. The companies pushing it are building an industry we can't participate or share in. They're cordoning off areas of tech and staking ground for themselves. It's placing a steep fence around tech. I hope every such closed source AI effort is met with equivalent open source and that the investments made into closed AI go to zero. The most likely outcome is that Google, OpenAI, and Anthropic win and every other "lab"-shaped company dies an expensive death. RunwayML spent hundreds of millions and they're barely noticeable now. These open source models hasten the deaths of the second tier also-ran companies. As much as I hope for dents in the big three, I'm doubtful.
- raw_anon_1111 10mo agoI can’t think of a single company I’ve worked with as a consultant that I could convince to use DeepSeek because of its ties with China even if I explained that it was hosted on AWS and none of the information would go to China. Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political capital. They would feel the same way about using xAI or maybe even Facebook models.
- htrp 10mo agowhat is the ballpark vram / gpu requirement to run this ?
- rhdunn 10mo agoFor just the model itself: 4 x params at F32, 2 x params at F16/BF16, or 1 x params at F8, e.g. 685GB at F8. It will be smaller for quantizations, but I'm not sure how to estimate those. For a Mixture of Experts (MoE) model you only need to have the memory size of a given expert. There will be some swapping out as it figures out which expert to use, or to change expert, but once that expert is loaded it won't be swapping memory to perform the calculations. You'll also need space for the context window; I'm not sure how to calculate that either.
- petu 10mo agoI think your idea of MoE is incorrect. Despite the name they're not "expert" at anything in particular, used experts change more or less on each token -- so swapping them into VRAM is not viable, they just get executed on CPU (llama.cpp).
- jodleif 10mo agoA common pattern is to offload (most of) the expert layers to the CPU. This combination is still quite fast even with slow system ram, though obviously inferior to a pure VRAM loading
- anvuong 10mo agoI think your understanding of MoE is wrong. Depending on the settings, each token can actually be routed to multiple experts, called experts choice architecture. This makes it easier to parallelize the inference (each expert on a different device for example), but it's not simply just keeping one expert in memory.
- deleted 10mo ago[deleted]
- lalassu 10mo agoDisclaimer: I did not test this yet. I don't want to make big generalizations. But one thing I noticed with chinese models, especially Kimi, is that it does very well on benchmarks, but fails on vibe testing. It feels a little bit over-fitting to the benchmark and less to the use cases. I hope it's not the same here.
- vorticalbox 10mo agoThis used to happen with bench marks on phones, manufacturers would tweak android so benchmarks ran faster. I guess that’s kinda how it is for any system that’s trained to do well on benchmarks, it does well but rubbish at everything else.
- make3 10mo agoyes, they turned off all energy economy measures when benchmarking software activity was detected, which completely broke the point of the benchmarks because your phone is useless if it's very fast but the battery lasts one hour
- make3 10mo agoI would assume that huge amount is spent in frontier models just making the models nicer to interact with, as it is likely one of the main things that drives user engagement.
- not_that_d 10mo agoWhat is "Vibe testing"?
- BizarroLand 10mo agoI would assume that it is testing how well and appropriately the LLM responds to prompts.
- deleted 10mo ago[deleted]
- 10mo ago
- spullara 10mo agoI hate that their model ids don't change as they change the underlying model. I'm not sure how you can build on that. % curl https://api.deepseek.com/models \ -H "Authorization: Bearer ${DEEPSEEK_API_KEY}" {"object":"list","data":[{"id":"deepseek-chat","object":"model","owned_by":"deepseek"},{"id":"deepseek-reasoner","object":"model","owned_by":"deepseek"}]}
- KronisLV 10mo agoOh hey, quality improvement without doing anything! (unless/until a new version gets worse for your use case)
- hnfong 10mo agoAgree that having datestamps on model ids is a good idea, but it's open source, you can download the weights and build on those. In the long run, this is better than the alternative of calling API of a proprietary model and hoping it doesn't get deprecated.
- deaux 10mo agoAnthropic has done similar before (changing model behavior on the same dated endpoint).
- cherioo 10mo agoAllegedly deepseek is doing this because they don’t have enough gpu to serve two models concurrently.
- twistedcheeslet 10mo agoHow capable are these models at tool calling?
- potsandpans 10mo agoFrom some very brief experimentation with deepseek about 2 months ago, tool calling is very hot or miss. Claude appears to be the absolute best.
- segmondy 10mo agoDepends on if they are trained for tool calling, this model is experimentation with new architecture, training methods, etc. It's not designed for tool calling. If you want to tool call, then you should look into DeepSeekv3.1-Terminus.
- Foobar8568 10mo agoAt least, there is no doubt where he is from ! which version are you? 我是DeepSeek最新版本模型! 如果你想了解具体的版本号信息,我建议你: 查看官方文档 - DeepSeek官网和文档会有最准确的版本信息 关注官方公告 - 版本更新通常会在官方渠道公布 查看应用商店/网页版 - 使用界面通常会显示当前版本 我具备DeepSeek的所有最新功能特性,包括: 强大的对话和推理能力 128K上下文长度 文件上传处理(图像、文档等) 联网搜索功能(需手动开启) 完全免费使用 如果你需要知道精确的版本号用于技术对接或其他特定用途,最好直接查阅官方技术文档,那里会有最准确和详细的技术规格说明。 有什么其他问题我可以帮你解答吗?
- schlauerfox 10mo agoIt's so strange when it obviously hits a preprogrammed non-answer in these models, how can one ever trust them when there is a babysitter that interferes in an actual answer. I suppose that asking it what version it is isn't a valid question in it's training data so it's programmed to say check the documentation, but still definitely suspicious when it gives a non-answer.
- embedding-shape 10mo ago> DeepSeek-V3.2 introduces significant updates to its chat template compared to prior versions. The primary changes involve a revised format for tool calling and the introduction of a "thinking with tools" capability. At first, I thought they had gone the route of implementing yet another chat format that can handle more dynamic conversations like that, instead of just using Harmony, but looking at the syntax, doesn't it look exactly like Harmony? That's a good thing, don't get me wrong, but why not mention straight up that they've implemented Harmony, so people can already understand up front that it's compatible with whatever parsing we're using for GPT-OSS?
- throwdbaaway 10mo agoThat DSML in the encoding directory looks quite a bit different from the Harmony chat template.
- catigula 10mo agoTo push back on naivety I'm sensing here I think it's a little silly to see Chinese Communist Party backed enterprise as somehow magnanimous and without ulterior, very harmful motive.
- deleted 10mo ago[deleted]
- jascha_eng 10mo agoOh they need control of models to be able to censor and ensure whatever happens inside the country with AI stays under their control. But the open-source part? Idk I think they do it to mess with the US investment and for the typical open source reasons of companies: community, marketing, etc. But tbh especially the messing with the US, as a european with no serious competitor, I can get behind.
- catigula 10mo agoThey're pouring money to disrupt American AI markets and efforts. They do this in countless other fields. It's a model of massive state funding -> give it away for cut-rate -> dominate the market -> reap the rewards. It's a very transparent, consistent strategy. AI is a little different because it has geopolitical implications.
- ForceBru 10mo agoWhen it's a competition among individual producers, we call it "a free market" and praise Hal Varian. When it's a competition among countries, it's suddenly threatening to "disrupt American AI markets and efforts". The obvious solution here is to pour money into LLM research too. Massive state funding -> provide SOTA models for free -> dominate the market -> reap the rewards (from the free models).
- catigula 10mo agoWe don't do that.
- mcbuilder 10mo agoAfter using it a couple hours playing around, it is a very solid entry, and very competitive compared with the big US relaeses. I'd say it's better than GLM4.6 and I'm Kimi K2. Looking forward to v4
- energy123 10mo agoDid you try with 60k+ context? I found previous releases to be lacklustre which I tentatively attributed to the longer context, due to the model being trained on a lot of short context data.
- gradus_ad 10mo agoHow will the Google/Anthropic/OpenAI's of the world make money on AI if open models are competitive with their models? What hurt open source in the past was its inability to keep up with the quality and feature depth of closed source competitors, but models seem to be reaching a performance plateau; the top open weight models are generally indistinguishable from the top private models. Infrastructure owners with access to the cheapest energy will be the long run winners in AI.
- tsunamifury 10mo agoPure models clearly aren’t the monetizing strategy, use of them on existing monetized surfaces are the core value. Google would love a cheap hq model on its surfaces. That just helps Google.
- gradus_ad 10mo agoHmmm but external models can easily operate on any "surface". For instance Claude Code simply reads and edits files and runs in a terminal. Photo editing apps just need a photo supplied to them. I don't think there's much juice to squeeze out of deeply integrated AI as AI by its nature exists above the application layer, in the same way that we exist above the application layer as users.
- tsunamifury 10mo agoGemini is the most used model on the planet per request. All the facts say otherwise to your thoughts here.
- Craighead 10mo ago[dead]
- dotancohen 10mo agoPeople and companies trust OpenAI and Anthropic, rightly or wrongly, with hosting the models and keeping their company data secure. Don't underestimate the value of a scapegoat to point a finger at when things go wrong.
- wosined 10mo agoRemember: If it is not peer-reviewed, then it is an ad.
- vessenes 10mo agoI mean.. true. Also, DeepSeek has good cred so far on delivering roughly what their PR says they are delivering. My prior would be that their papers are generally credible.
- Havoc 10mo agoGood general approach, but deepseek has thus far always delivered. And not just delivered, but under open license too. "Ad" as starting assumption seems overly harsh
- orena 10mo agoAny results on frontier math or arc ?
- Havoc 10mo agoNote combination of big frontier level model and MIT license.
- singularity2001 10mo agoWhy are there so few 32,64,128,256,512 GB models which could run on current consumer hardware? And why is the maximum RAM on Mac studio M4 128 GB??
- jameslk 10mo ago128 GB should be enough for anybody (just kidding). I hope the M5 Max will have higher RAM limits
- aryonoco 10mo agoM5 Max probably won’t, but M5 Ultra probably will
- eldenring 10mo agothe only real benefit is privacy which 99.9% of people dont get about. Almost all serving metrics (cost, throughput, ttft) are better with large gpu clusters. Latency is usually hidden by prefill cost.
- cowpig 10mo agoMore and more people I talk to care about privacy, but not in SF
- mistercheph 10mo agoand sovereignty. I can go into the woods with a fuzzy approximation of all internet text in my backpack
- ainch 10mo agoAs LLMs are productionised/commodified they're incorporating changes which are enthusiast-unfriendly. Small dense models are great for enthusiasts running inference locally, but for parallel batched inference MoE models are much more efficient.
- sidcool 10mo agoCan someone kind please ELI5 this paper?
- 0xedd 10mo ago[dead]
- HarHarVeryFunny 10mo agoThey've developed a sparse attention mechanism (which they document and release source code for) to increase model efficiency with long context, as needed for fast & cost-effective extensive RL training for reasoning and agentic use They've built a "stable & scalable" RL protocol - more capable RL training infrastructure They've built a pipeline/process to generate synthetic data for reasoning and agentic training These all combine to build an efficient model with extensive RL post-training for reasoning and agentic use, although they note work is still needed on both the base model (more knowledge) and post-training to match frontier performance.
- nickandbro 10mo agoFor anyone that is interested "create me a svg of a pelican riding on a bicycle" https://www.svgviewer.dev/s/FhqYdli5 https://www.svgviewer.dev/s/FhqYdli5
- chronogram 10mo agoIt created a whole webpage to showcase the SVG with animation for me: https://output.jsbin.com/qeyubehate https://output.jsbin.com/qeyubehate
- sfdlkj3jk342a 10mo agoWhat version is actually running on chat.deepseek.com? It refuses to tell me when asked, only that it's been train with data up until July 2024, which would make it quite old. I turned off search and asked it for the winner of the US 2024 election, and it said it didn't know, so I guess that confirms it's not a recent model.
- scottyeager 10mo agoYou can read that 3.2 is live on web and app here: https://api-docs.deepseek.com/news/news251201 https://api-docs.deepseek.com/news/news251201 The pdf describes how they did "continued pre-training" and then post training to make 3.2. I guess what's missing is the full pre-training that absorbs most date sensitive knowledge. That's probably also the reason that the versions are 3.x still.
- chistev 10mo agoI've found it better than ChatGPT lately, at least the free version of GPT. I don't know, but GPT seems to have regressed a lot, at least the free version.
- johnnienaked 10mo agoAre we the baddies?
- a96 10mo agoThe AI says shake... "Signs point to yes."
- samir123766 10mo agonice
- EternalFury 10mo agoIt does seem good, but it’s slow.
- nickstinemates 10mo agoI am waiting for the first truly open model without any of the censorship built in. I wonder how long it will take and how quickly it will try to get shut down.
- naeq 10mo agoMost open models have been converted to uncensored versions. Search for the model name with the suffix "abliterated".
- Aldo_MX 10mo agoThat's not a realistic expectation. Classic examples like: User: I'm feeling bad LLM: Have you considered k*****g yourself? Are a good example of what an LLM "without censorship" looks like: Good at predicting the most common sequence of text (ex. the most common sarcastic reply from Reddit), but effectively useless. In order to build a useful LLM (ie. one that actually follows instructions) you need to teach the LLM to prefer the most helpful answer, and that process by itself is already an implicit layer of "censorship" as it requires human supervision, and different humans have different perceptions on what the most helpful answer is, especially when their paycheck is conditioned to a list of "corporate values". You can only pick between a parrot that repeats random text from the Internet, or a parrot lobotomized to follow the orders from their trainers (which occasionally repeats random text from the Internet, because the training isn't perfect). Unsurprisingly, the lobotomized parrot is more useful to get actual work done, even if it won't tell you what the CIA[1] did to Mexican Students on October 2nd, 1968. [1]: https://www.bbc.com/mundo/noticias-america-latina-45662739 https://www.bbc.com/mundo/noticias-america-latina-45662739
- johnxie 10mo agoCool to see open models catching up fast. For builders the real question is simple. Which model gives you the tightest loop and the least surprises in production. Sometimes that is open. Sometimes closed. The rest is noise.
- imbusy111 10mo agoFunny to see tau2-bench on the list of benchmarks, when tau2-bench is flawed and 100% score is impossible, unless you add the tasks to the training set: https://github.com/sierra-research/tau2-bench/issues/89 https://github.com/sierra-research/tau2-bench/issues/89
- mark_l_watson 10mo agoI used DeepSeek-v3.2 to solve two coding problems by pasting code and directions as one large prompt into a chat interface and it performed very well. VERY WELL! I am still happy to pay Google because of their ecosystem or Gemini app, NotebookLM, Colab, gemini-cli, etc. Google’s moat for me is all the tooling and engineering around the models. That said, my one year Google AI subscription ends in four months and I might try an alternative, or at least evaluate options. Alibaba Cloud looks like an interesting low cost alternative to AWS for building systems. I am now a retired ‘gentleman scientist’ now and my personal research is inexpensive no matter who I pay for inference compute, but it is fun to spend a small amount of time evaluating alternatives even though mostly using Google is time efficient.
- cgearhart 10mo agoSo DSA means a lightweight indexing model evaluated over the entire context window + a top-k attention evaluation. There’s no soft max in the indexing model, so it can run blazingly fast in parallel. I’m surprised that a fixed size k doesn’t experience degrading performance in long context windows though. That’s a _lot_ of responsibility to push into that indexing function. How could such a simple model achieve high enough precision and recall in a fixed size k for long context windows?
- swframe2 10mo agoThe AI market is hard to predict due to the constant development of new algorithms that could emerge unexpectedly. Refer to this summary of Ilya's opinions for insights into the necessity of these new algorithms: https://youtu.be/DcrXHTOxi3I https://youtu.be/DcrXHTOxi3I DeepSeek is a valuable product, but its open-source nature makes it difficult to displace larger competitors. Any advancements can be quickly adopted, and in fact, it may inadvertently strengthen these companies by highlighting weaknesses in their current strategies.
- matt-alive 10mo agoIs it open source vs enterprise or China vs US?
- Frannky 10mo agoSmart model—I use it as my main chat. It's interesting that markets were able to predict that it would lower the revenue of the paid ones.