6 ms·
DeepSeek Open Source FlashMLA – MLA Decoding Kernel for Hopper GPUs
- helloericsf 2y agoX:https://x.com/deepseek_ai/status/1893836827574030466 https://x.com/deepseek_ai/status/1893836827574030466 BF16 support Paged KV cache (block size 64) 3000 GB/s memory-bound & 580 TFLOPS compute-bound on H800
- WithinReason 2y agoThat's 90% bandwidth efficiency and 60% compute efficiency https://www.nvidia.com/en-us/data-center/h100/ https://www.nvidia.com/en-us/data-center/h100/
- helloericsf 2y agoThey don't have h100. wink,wink.
- rfoo 2y agoThey have H800s which have exactly same memory bandwidth and max FLOPS.
- pk-protect-ai 2y agoWhat about NVLink? Does it plays a role here?
- rfoo 2y agoFor FlashMLA? No. The code here runs on one GPU only and do not have a builtin communication part.
- pk-protect-ai 2y agoBut for the training it does. You need to communicate gradient changes between GPUs.
- deleted 2y ago[deleted]
- deyiao 2y agoI heard their inferencing framework is way lower than typical deployment methods. Can this be verified from that open-source project? How does it stack up against vllm or llama.cpp
- helloericsf 2y agoWhat do you mean by "lower"? To my understanding, they will open 5 infra related repos this week. Let's revisit your comparison question on Friday.
- find0x90 2y agoI don't see any use of PTX, might be in one of the other repos they plan to release.
- DesiLurker 2y agoright, I think PTX use is a bigger deal than its getting coverage for. this opens an opening for other vendors to get their foot in with PTX to LLVM-ir translation for existing cuda kernels.
- reissbaker 2y agoBy "lower" you mean cheaper/better? I suspect it's much higher throughput than vLLM, which in turn is much higher throughput than llama.cpp. The MLA kernel they just open-sourced seems to indicate that, although we'll see how it does in third party benchmarks on non-hobbled GPUs vs FlashAttention. They only released the BF16 version — whereas most people, including DeepSeek themselves, serve in FP8 — so it might not be immediately useful to most companies quite yet, although I imagine there'll be FP8 ports soon enough.
- mohsen1 2y agoI'm confused. Wasn't there sanctions against Chinese companies about Hopper GPUs? Are they just admitting that they had access to H100 against the US sanctions?!
- thot_experiment 2y agoJust the H100, the H800 is a region-specific version of the card for china with shitty nvlink bandwidth which makes it rougher for making big clusters, but deepseek was able to mitigate the impact of that by being clever (rumored to have made significant use of PTX assembly instead of just using CUDA, we'll probably find out in the releases this week)
- deleted 2y ago[deleted]
- Tiberium 2y agoH800 is the export variant that they had access to. They directly reference it in the repo: >Achieving up to 3000 GB/s in memory-bound configuration and 580 TFLOPS in computation-bound configuration on H800 SXM5, using CUDA 12.6.
- ahofmann 2y agoIt isn't illegal for chinese companies to buy H100 cards. It is illegal for USA companies to sell them to China. So the "admit" part wouldn't be on Chinas side.
- jofzar 2y agoIt's also totally legal to sell h100 cards to a country that is very close to China. Unrelated, it's always impressed me how Singapore buys 15% of the world's h100's. Really is the AI development capital of the world.
- xbmcuser 2y agoNot really Singapore is a trading hub a lot of multi national companies have regional offices or head offices in Singapore so if the head office buys anything for any where the purchase will show up as Singapore. Despite Nvidia showing such a large revenue from Singapore actual number of gpu shipped to Singapore is not that high. Not that some of the gpus are not going China but their is a valid reason for the Nvidia Singapore revenue numbers. https://www.tomshardware.com/tech-industry/deepseek-gpu-smuggling-probe-shows-nvidias-singapore-gpu-sales-are-28-percent-of-its-revenue-but-only-1-percent-are-delivered-to-the-country-report https://www.tomshardware.com/tech-industry/deepseek-gpu-smug...
- behnamoh 2y agoOpen AI is back!
- echelon 2y agoThe real "Open" AI.
- fsndz 2y agoDeepSeek is just the gift that keeps on giving. I now agree with people who say open source AI will win: https://open.substack.com/pub/transitions/p/deepseek-is-coming-for-openais-neck?r=56ql7&utm_campaign=post&utm_medium=web&showWelcomeOnShare=false https://open.substack.com/pub/transitions/p/deepseek-is-comi...
- baq 2y agoOpen sourcing is the runner-up’s way to ensure the current best player doesn’t steal the whole market. The elephant in the room is obviously the cluster size required, it hardly matters for normal people that the weights are free. We needed more efficiency breakthroughs.
- PeterStuer 2y agoIt matters a lot, even if you never intend to run it yourself or look at the code. It means that people can and will provide this service, and 1000's will build on this and make offers that you can use in either a commodity base market, or with a specific niche target. It means regulatory capture and control will be much, much harder to execute. It means AI might continue to be a benefit also to you rather than just a way to control, propagandize and exploit you.
- fsndz 2y agoabsolutely on point!
- helsinkiandrew 2y ago
- rvz 2y agoThis is the minimum bar that I expect very elite programmers should be striving for in the age of AI and DeepSeek should be studied as an example and this is the only just the first of many projects from them. There is an extremely high chance (in fact a 99.9% chance) that an AI did not build this and the ones who are able to build or adapt projects like this which are deep into hardware systems will be the most sort after. Not the horrendous JS or even TS slop across GitHub that is extremely easy for an AI to generate correctly. You've got until 2030 to decide. And my advice is to study the codebases of pytorch (backends), DeepSeek, tinygrad and ggml.
- beernet 2y agoLLM generated comments are so 2024
- BoorishBears 2y agoNothing about that comment implies it's LLM generated, and it's bizzare how it's being received since it's a pretty reasonable take.
- deleted 2y ago[deleted]
- rnewme 2y agoI don't find it a reasonable take, it's like saying stackoverflow.com is taking developer jobs by making it easy to code, we better develop new stackoverflow.com
- jbm 2y agoIt's an interesting opinion, but I read the exact same opinions about JS developers in 2008 too. I do agree that if you are "only" a developer, you will have to be in some sort of tightly defined niche, and how long those niches survive is anyone's guess.
- KeplerBoy 2y ago
- nokun7 2y ago[flagged]
- m3kw9 2y agoMHGA making hopper great again
- eigenvalue 2y agoNice, probably saved a bunch of FANG devs a lot of hours of work trying to knock this off.
- nicce 2y agoThere were likely some startups that tried to sell the same thing…
- anon389r58r58 2y agoYou mean like Modular?
- nicce 2y agoOr Silo AI (as an example of why) : https://www.silo.ai/blog/amd-to-acquire-silo-ai-to-expand-enterprise-ai-solutions-globally https://www.silo.ai/blog/amd-to-acquire-silo-ai-to-expand-en...
- refibrillator 2y agovLLM supports MLA for Deepseek models as of 3 weeks ago. 3x higher generation throughput and 10x token memory capacity. https://github.com/vllm-project/vllm/releases/tag/v0.7.1 https://github.com/vllm-project/vllm/releases/tag/v0.7.1 MHA is still faster in low QPS regime apparently. https://neuralmagic.com/blog/enhancing-deepseek-models-with-mla-and-fp8-optimizations-in-vllm/ https://neuralmagic.com/blog/enhancing-deepseek-models-with-... Also published this month was theoretical proof showing that for the same KV Cache overhead, MLA consistently offers greater expressive power than GQA. Furthermore, widely used GQA-based pre-trained models (e.g. LLaMA, Qwen, Mixtral) can be converted into MLA-based models. https://arxiv.org/pdf/2502.07864 https://arxiv.org/pdf/2502.07864
- albertzeyer 2y agoI also just read that paper. But I wonder, even though MLA is strictly more powerful, do you really gain by that in experiments? This paper doesn't really do too much experimental comparisons. GQA on the other side should still be faster (no need to an extra linear transformation).
- menaerus 2y agoPretty significant improvements. However, my back on the napkin math suggests that MLA, FlashAttention and similar optimizations will provide the benefits only when memory access time dominates the compute in attention implementation? Those would be the prefill-phase (or TTFT) and training (when batch_size >> 1) but not the decode phase (inference)?
- rfoo 2y agoYou've got it backwards. After FlashAttention, it's the decoding part being bound mainly by memory access. With FA as long as you have enough batch size you can push training/prefill to be compute-bound.
- menaerus 2y agoI don't think I got it backwards, I believe what I said is correct - FA does not improve inference time. From the authors of FlashAttention: > This [decoding] operation has been optimized with FlashAttention (v1 and v2 recently) in the training case, where the bottleneck is the memory bandwidth to read and write the intermediate results And then they continue with: > However, these optimizations don’t apply directly to the inference case, because the bottlenecks are different. For training, FlashAttention parallelizes across the batch size and query length dimensions. During inference, the query length is typically 1 ... With a batch size of 1, FlashAttention will use less than 1% of the GPU! And then they come up with a different proposal, FlashDecoding, that optimizes for inference time: > Our new approach Flash-Decoding is based on FlashAttention, and adds a new parallelization dimension: the keys/values sequence length. It combines the benefits of the 2 approaches from above. Like FlashAttention, it stores very little extra data to global memory, however it fully utilizes the GPU even when the batch size is small, as long as the context length is large enough. Link: https://crfm.stanford.edu/2023/10/12/flashdecoding.html https://crfm.stanford.edu/2023/10/12/flashdecoding.html
- FL33TW00D 2y agoIt seems to me that MLA will become the standard from here on out. If Deepseek R1 had used standard MHA, they would need 1749KB per token for KV cache storage. This means that once the conversation reaches ~46,000 tokens, the KV cache will have exceeded the entire storage capacity of a single H100. Using MLA, each token now consumes 125KB. This means you can hit ~640,000 tokens (2x Ulysses) before overflowing.
- antonmarsh 2y ago[dead]
- ur-whale 2y agoFor those who wonder ... it's somewhat likely that MLA mean Multi-head latent attention https://verticalserve.medium.com/group-query-attention-58283b337c65 https://verticalserve.medium.com/group-query-attention-58283... https://paperswithcode.com/method/multi-head-attention https://paperswithcode.com/method/multi-head-attention
- rob_c 2y agoGreat work any plans to integrate with pyT or TF I wonder? (Showing my lack of breadth of knowledge in the ecosystem (s))
- mclau156 2y agoWas really hoping we could get flash games back with AI
- kridsdale1 2y agoAsk an LLM to write you some ActionScript3
- imranq 2y agoDang only forward passes. The real secret was in the backward pass! I was also curious to learn how they implemented the dualpipe scheduler
- rfoo 2y agoDo they even have an optimized backward? It looks like optimizations like this aren't needed during training. Their V2 paper also suggests so.
- syntex 2y agoWhat i can do with that?
- rfoo 2y agoProbably nothing. Inference providers like Fireworks, or major clouds, can use this to reduce their cost, if they don't already have a replication with similar perf. vLLM and SGLang may integrate this to be faster at serving DeepSeek-V2/V2.5/V3/R1 on H100/H800s. I believe that's why they didn't release this back then, this is part of their "moat" (pretty weak tho) and it only benefits competitors. Open sourcing this after being very popular may indicate that they don't want all the users to use their API/Chat and now want the world to serve it instead? Idk.