9 ms·
DSpark: Speculative decoding accelerates LLM inference [pdf]
- preetham_rangu 3mo agodo they use their OCR, or someone else?
- Havoc 3mo agoNice. Guessing the timing isn't accidental. Demonstrated openness vs harsh regulation
- declan_roberts 3mo agoNobody forced anthropic to go on a media blitz loudly proclaiming the dangers their new AI model. Serves them right honestly.
- cr125rider 3mo agoChina = Open. US = Harsh Regulation Strange timeline, though this only works because it’s aligned with Xi’s goals.
- Havoc 3mo agoYeah can definitely see a world where china pivots and we're stuck with closed/closed Mistral...don't fumble this
- skeledrew 3mo agoWhat are some things that China has pivoted on in history?
- Ey7NFZ3P0nzAe 3mo agoIIRC Mao was surprisingly quite okay with male homosexuality.
- cindyllm 3mo ago[dead]
- Jackobrien 3mo agoI see a world soon where there’s an extremely wide variety of small models for speculative decoding, unique to use cases, companies, and even individuals.
- nicce 3mo agoHopefully that is the case and hardware does not get impossible to get.
- pydry 3mo agoyes, heavily constrained by sophisticated guardrails. this is definitely where things are going. the enormous "eat the world" models have extreme diminishing returns by comparison.
- Der_Einzige 3mo agoYou clearly didn't read the recent speculative decoding papers because it's been possible to use any model to speculate for any other model for awhile. They solved the tokenization problems that prevented this in the past.
- ricardobeat 3mo agoPresumably this has been in production for a while, and is one of the reasons they were able to dramatically lower prices a month ago?
- _0ffh 3mo agoLookahead Sparse Attention should be playing a big role as well, as it dramatically slashes memory consumption.
- chronogram 3mo agoYes. Section 5 talks about real-world deployment: 5.1: "The DSpark draft models are co-deployed with the preview versions of DeepSeek-V4-Flash and DeepSeek-V4-Pro"; 5.4: "MTP-1 represents the former production setup, having been superseded by DSpark two weeks following the DeepSeek-V4-preview release."
- sourcecodeplz 3mo agogood catch, they reduced the prices 75% seems like exactly in line with the speed/inference optimizations gains?
- piterrro 3mo agoI’ve been using DeepSeek v4 pro for a month now in Kilo Code and its great. Fast, reliable, large context window and cheap as… Did 1,5B tokens this month and cost me 40usd (majority cached, but still).
- spiderfarmer 3mo agoIs there a way to see how many tokes one does with claude code (pro)?
- cptchaos 3mo agohttps://ccusage.com/ https://ccusage.com/
- bpavuk 3mo agothe casino has no clocks, as one HN user put it some time ago. I second ccusage, it's nice
- edg5000 3mo agoIt's in the JSONs in ~/.claude, but last 30 days only I think. You can have the model analyze history. So for correct history you'd need to run history analysis on a cron job or something. Kinda hacky.
- Stagnant 3mo agoThe 30 day limit can be overridden by adding "cleanupPeriodDays": 9999 to .claude/settings.json
- O_H_E 3mo agohttps://github.com/kenn-io/agentsview https://github.com/kenn-io/agentsview > Local-first session search, analytics, insights, and token use statistics for coding agents, supporting Claude Code, Codex, and more than 20 other agents. solid piece of software
- richardlblair 3mo ago
- rvz 3mo agoThis is just one of many papers DeepSeek have released to be able to serve models at extremely cheap prices, unlike the others taking on >$100B+ of debt in building data centers for the same thing. > As with V4-Flash, we treat this point as an indication that DSpark sustains useful throughput under an interactivity target that the baseline cannot efficiently support. At matched system capacities, DSpark delivers 57% to 78% faster per-user generation. Reminds me of the flawed solution in scaling servers in 2017 that use memory-intensive technologies by adding even more servers to solve the problem. (It just increases costs.) Rather than doing that, think about which critical parts of your app can be written in a more performant technology. Fast forward to 2026, now you can see who is just throwing more money at the problem to create even more problems where as DeepSeek is giving us optimized solutions. I know exactly who I would pay attention to, and it is absolutely not Anthropic.
- denverllc 3mo agoFor so long American companies have operated under the assumption that servers are cheaper than developers, and that was used to justify all sorts of inefficient practices. The last year has shown that’s not true anymore (even for web servers).
- simianwords 3mo ago...... are you really suggesting OpenAI and Anthropic don't have access to these techniques?
- sourcecodeplz 3mo agoif they didn't, they do now. as deepseek published the howto
- 2838383838 3mo agoMust be wonderful to be on the board of OpenAi et al & their PE investors whilst China keeps blowing up these mines under their feet lmao. Luckily Korean pension funds will buy all the trash as usual but goddamn you gotta start moving quick or you are gonna need some serious AGI to show you how to offload those bonds
- ForHackernews 3mo ago"We will build the machine-god and pray for it to pay for itself."
- FridgeSeal 3mo agoEvery day, the rate of “could post a picture of 40k tech priests and have it taken unironically” goes up, and it’s starting to get concerning.
- ozgrakkurt 3mo agoDon’t worry they will sell all the hardware and data they acquired with their grift
- throwa356262 3mo agoWhy do you think they have started accusing Chinese labs of stealing and distillation? A&O no longer have the most to justify their high valuation. The only thing they can do now is to get the government forbid the Chinese models.
- kamranjon 3mo agoDeepSeek continues to not only push the boundaries but also publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately. Chinese labs are doing the most interesting work in AI right now.
- herodoturtle 3mo agoPublishing by necessity I wonder? American labs on the cutting edge pioneering the way forward, so Deepseek open sourcing what they’ve got is to help even the playing field. Hopefully the experts here can offer insight. The above is just my hunch and I’m not a specialist in this field.
- jonplackett 3mo agoWouldn’t that just help the American labs anyway though? Or do they assume they’ve actually already figured this stuff out and kept it secret?
- 7speter 3mo agoFrom what I gather, the Chinese are behind, but a lot of their research amounts to scrappy, clever discoveries in how to use more novel technologies (for Qwen and Deepseek, its mixture of expert models, that can do inference using a portion of the model at a time). The chinese also distill information from American models, so there’s that. The American companies, from my impression don’t involve themselves with such lowly “hacks” because they have so much money to just push forward with doing everything on big heavy models that run on the most cutting edge nvidia chips that they can, the moment, kinda sorta get on demand (I say that in some degree of jest).
- idiotsecant 3mo agoThe American companies would love to develop these 'hacks' because it would make them more money, something they are in existential need of right now. They don't develop them because they don't collaborate publicly anymore. Where would the whole industry be if Google never allowed publishing the transformers paper? It's not a coincidence that the American AI industry grew fastest in capability when it was the most open.
- pokot0 3mo agoI am wondering if this is why they can offer their pro model at ~1/4th of the price compared to the other providers offering the same model, and if other providers will be able to do the same in a short timeframe.
- vidarh 3mo agoIt'd presumably help a lot, but also when you use their endpoint they get more training data.
- nicce 3mo agoThis applies to every provider. OpenAI seems to be the worst hoarder.
- pokot0 3mo agoactually you can buy inference on third party providers that serve deepseek v4 pro with zero data retention (ZDR).
- nicce 3mo agoOnly reliable way to have zero data retention is to self-host.
- LeBit 3mo agoTrue. But at some point you got to close your eyes and take a step forward. It’s like with VPN providers. Is Mullvad actually collaborating with law enforcement? They very well could be. It is a calculated risk. Is DeepInfra actually logging and training or selling the logs? They could be.
- flipped 3mo agoMullvad has proved it doesn't collect. It's laughable to even suggest it. They have been raided multiple times, tons of audits, does bleeding edge research on privacy preserving tech, donates to GOS, etc etc. You don't see this kind of VPN company at all because none exists.
- imrozim 3mo ago[dead]
- danielabinav160 3mo agoWould love to see these numbers reproduced on consumer GPUs, not just A100s.
- tommica 3mo agoMaybe somaday an 8gb videocard can be used for coding...
- romanusrome 3mo ago[dead]
- wolttam 3mo agoThis is an efficiency improvement that significantly lowers the amount of RAM you have to look at, on average, during decode. It should improve performance on most hardware because most LLMs are memory bandwidth bound during decode.
- kamranjon 3mo agoThe hugging face models are already up and seem to be the original models with the speculative decoding module built in which is very cool: Flash: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark Pro: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark Excited to see if this makes it into DwarfStar for local inference, have been using the flash model extensively since the 2-bit quants were made available by antirez.
- ilaksh 3mo agoAny chance they will have this for Qwen 27 b also?
- kamranjon 3mo agoThe paper actually references testing their DSpark speculative decoding strategy with Qwen 3 4b, 8b and 14b models so while I doubt they will release builds themselves, they’ve open sourced (DeepSpec) their training pipeline for this so we will likely see folks adopting for other models.
- lelanthran 3mo agoThese companies providing tokens, whether SOTA or not, that want to IPO are so fucked as time goes on. Can't sell their SOTA models, only slightly better than the open source models for the models they can sell, cost 20x to 50x for good models, a TAM that consists almost solely of developers, with no customer of theirs actually boasting increased profits as a result of AI... I fear their time to IPO may have passed.
- deleted 3mo ago[deleted]
- utopiah 3mo agoThe question is even, was there EVER a time for an IPO? If the business model requires hundreds of billions to get the required quality (R&D but also infrastructure to collect data and train, either purchased or rented to 3rd party) while "only" dozens of billions can be earned back (as costs still exist to earn, it's not free once models are trained), then maybe there NEVER was nor till be a good time for an IPO in a rational market.
- 2838383838 3mo agoIPOs with massive bags can be wework or spacex, it all depends on vibes. If they buy a couple more articles doomposting and glazing AI on the financial times right before exit they will def find a bunch of boomers to buy their bags. If the narrative changes before they IPO its over.
- notnullorvoid 3mo ago> in a rational market. Unfortunately the market is often not rational in this way. Hype within retail market means there are suckers willing to buy. Institutional market knows there are suckers when the hype is high. Both would drive the price up, and retail investors the ones left when it falls.
- bflesch 3mo agoAt this point why can't someone produce a fridge or container-sized AI appliance based on legacy chips (12nm)? I imagine this would cover 80% of corporate use cases where you need to "google-in-a-box" functionality. The state-of-the-art nanometer are impossible to achieve but if you have infinite solar energy during business hours does it really matter? Every company has a parking spot so this ASIC-like appliance could be as big as a shipping container. If it could just run recent open models for a handful of users it would be such a nobrainer to buy.
- sixhobbits 3mo agoNvidia is already selling exactly this I think, not sure when it's expected to ship
- scrlk 3mo agoSee "exabox" from George Hotz: https://tinycorp.myshopify.com/products/exabox-preorder https://tinycorp.myshopify.com/products/exabox-preorder
- benjiro29 3mo agoThe issue is that there are only so many fabs in the world that make memory. And if you want the good stuff, your easily going into 400 ~ 750b parameter models. That means at FP4 400 to 750GB memory. Did i mention there are only so many memory makers and they are all busy printing money with HBM memory? Intel is trying with Crescent Island, to make a 160GB GPU that uses LPDDR5X memory. HBM takes multiple times the resources to make vs basic DDR5 memory. So by going this route, you have more memory, with the disadvantage that its only 700GB/s. VS HBM pumping out Terrabyte numbers like its nothing. These cards is reasonably priced, may be good alternative to $10k 96GB Nvidia Blackwells... You give up on token generation (heavily memory dependent), for more memory to run larger models at home/office/company servers. The problem is, again, there are only so many memory makers and its not like the market is flooded with DDR5 memory anymore, as the big 3 moved a lot of production to HBM. Another approach is Sandisk making HBF ... Flash memory, like your typical NVME but designed around maximum speed. So instead of loading the models into expensive HBM memory, you use the benefits of density in Flash memory, to offload models into that. Cheaper, but slower... But it leaves your expensive HBM memory free for things like KV Cache, Active parameters, etc... So your model will be slower, but your hybrid using it. As in, faster then running a model from system memory with normal DDR memory, but not as fast as HBM. So yea, there is a lot in development to reduce the dependance of that resource eating HBM memory. For the wafer cost of 1GB HBM, you normally got 4GB normal memory. That is why the world supply of memory dropped. Not just the insane buying but be HBM is just very inefficient in wafer usage. Can we not use DDR4 production and create some kind of hybrid solution? Sure, but the big 3 moved away from DDR4 in favor of DDR5 a long time ago. We have competition from China with a mix of DDR4/DDR5, but they also need to scale up. Nobody expected to see a large part of the world production vanish into HBM... Even if its about DDR4 and older nodes, ironically, most companies had been moving away from DDR4. There is only so much wafer capability in the world, to the point that companies are moving to using DDR2 ... Yea, not a typo, like 2007 DDR2! for IOT devices etc, stuff that does not need fast memory. Because even DDR3 got too expensive for them. Its not like the old nodes are not used anymore ... Like that capacity was sitting idle. It was still in production making other stuff. The only real solution is that we need more fabs, and those take years to build. And the big 3 delayed investing in new fabs for a long time, unsure about the whole AI bubble stuff. Aka, they did not want to make a ton of fabs to end up with over capacity if the AI growth collapsed.
- StizzurpXDD 3mo agoDeepSeek is, as I feel currently, the sole AI company which is actually trying to innovate rather than top mere benchmarks. Others like OpenAI, Anthropic and Google are mostly just competeing with each rather than keep innovating around the clock.
- Alifatisk 3mo ago> DeepSeek is, as I feel currently, the sole AI company which is actually trying to innovate rather than top mere benchmarks. I'd also include the other Chinese labs like Moonshot (behind Kimi) and Z.ai (behind GLM). They are innovating and continue openly sharing their research to the public. I believe the founder of Moonshot even shared 40 minute video on Twitter where he goes through techniques that powers Kimi.
- alecco 3mo agoIsn't GLM-5.2 mostly DeepSeek V3 architecture? More and more I suspect Z.ai just has deeper pockets and access to the Claude traces while DeepSeek is punching way above their class.
- Alifatisk 3mo agoPerhaps, but Z.ai contributed with techniques such as IndexShare, which helps reduce computation for larger context windows (1M).
- Reubend 3mo agoThat understates how difficult it is to get to the level of performance they attained. The fact that it surpasses DeepSeek v4 in most ways shows that they accomplished some great work in this space.
- smcleod 3mo agoQwen as well.
- 3mo ago
- articlepan 3mo agoTitle is bad, it's the first line of the abstract instead of the paper title. Speculative decoding for LLM inference was published in 2022: https://arxiv.org/abs/2211.17192 https://arxiv.org/abs/2211.17192 This paper seems to be an improvement to speculative decoding but I haven't read it yet.
- swordlucky666 3mo ago[flagged]
- playorizaya 3mo ago[flagged]
- xnx 3mo agoIs this newer/better than the speculative decoding from 2022? https://arxiv.org/abs/2211.17192 https://arxiv.org/abs/2211.17192
- tiahura 3mo agoSeems like they focus on improving the drafter and the verification policy so speculation keeps producing net speedups rather than wasted verification work at deepseek scale.
- alok-g 3mo agoThat paper is cited in the 'introduction' and 'background' sections. This paper is improving by removing some bottlenecks.
- wg0 3mo agoThat's why I pay them. Regularly. Without fail. Despite my token usage isn't that much. But I vote for these heroes with my wallet. Just yesterday did again.
- noIdeaTheSecond 3mo agoCudos to you!If people realized how much power we had we's have a better world
- lightedman 3mo agoAnyone want to bet that much like speculative execution, speculative decoding is going to introduce a whole slew of vulnerabilities in the ways LLMs work?
- skirmish 3mo agoDon't think so because all tokens predicted speculatively are still validated against the main model (which is faster than predicting them from scratch) and only accepted if they match exactly.
- lightedman 3mo agoIf your main model is inherently-busted does validation actually matter?
- skirmish 3mo agoBut then it's not a problem of speculative decoding, fix your dam main model. BTW, in case there is confusion, we are not talking about CPU speculative execution affecting model inference at all, just about this specific technique: predict tokens via a smaller drafter model, then validate them against the main model in batch.
- eddysir 3mo ago[flagged]
- segmondy 3mo agoAs we can see again, this has nothing to do with distillation, yet for every gain Chinese labs make, the US labs will accuse them of theft. Yet they are constantly innovating.
- zftnb666 3mo agoAI making AI faster. Next up: AI writing papers about how AI makes AI faster
- myshapeprotocol 3mo ago[flagged]
- einrealist 3mo agoYet another band aid.
- porphyra 3mo agoI thought this had something to do with the DGX Spark at first from the name haha. (Incidentally, a lot of recent work has gone into making the DGX Spark better at inference, like MTP yielded a 50-100% speedup, so DSpark will likely be very helpful to that end as well)
- dnchdnd 3mo agoInteresting side effect that this seems to place a downward pressure margins of their western competitors: by sharing not only models but serving optimisations, even those third parties serving the DeepSeek models can do so more efficiently multiplying the effect