5 ms·
vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention
- jokoon 3y agoNow do the same for image classifiers. I tried a few of them, they're just horribly slow. This is pretty outrageous considering the first robust image image classifiers appeared around 2007.
- wskwon 3y agovLLM has been adopted by LMSYS for serving Vicuna and Chatbot Arena.
- thewataccount 3y agoThis is really cool to see. > Large: Takes up to 1.7GB for a single sequence in LLaMA-13B. > Dynamic: Its size depends on the sequence length, which is highly variable and unpredictable. As a result, efficiently managing the KV cache presents a significant challenge. We find that existing systems waste 60% – 80% of memory due to fragmentation and over-reservation. This mentions improvements for throughput which is great, and it mentions memory savings. I'm a bit confused how 80% of the memory could be wasted by the KV cache when the vast majority of the memory is usually holding the model itself? How much memory savings does this translate to effectively for say a 30B 4bit model?
- zhisbug 3y agoThis really depends on what GPUs you use. If you GPUs has very small amount of memory, vLLM will help more. vLLM addresses the memory bottleneck for saving KV caches and hence increases the throughput.
- Solvency 3y agoSemi-related question: this page is full of little charts and diagrams. There are thousands of similar projects/sites/experiment sites with their own charts and diagrams. But it seems like there are always subtle-to-large differences in them that indicate they're made with totally different libraries. Are there just thousands of homebrewn non-standard chart & diagram builders out there? How does one even begin to pick a standard to whip out quickies like these? Google SEO makes it virtually impossible to get to substance.
- daedbe 3y agoI often see charts produced using matplotlib or plotly - often you can tell based on the colour schemes used. For example, the bar chart at the bottom of this paper looks like it was made with plotly. I think the reason for such variance in the style of charts is largely due to the flexibility frameworks such as matplotlib provide: you can control basically every aspect of a chart and use any number of predefined or custom stylesheets to change the look and feel.
- kristjansson 3y agoThe color scheme on these implies Google Drawing, but I don't know how they made them into animations - maybe just manually?
- mattnewton 3y agoGoogle slides I think.
- wskwon 3y agoWe used matplotlib for the performance charts, and used a free website to convert google slides to the animation gifs.
- e12e 3y agoWhich "free website"?
- marcopicentini 3y agoIs it available an hosted demo? What are use cases for which open source models are equivalent of GPT 3.5?
- wskwon 3y agoYou can think of LMSYS Vicuna: https://chat.lmsys.org https://chat.lmsys.org as our hosted demo, as it actually uses vLLM as the backend.
- gwph 3y agoIon Stoica's lab continues to be a powerhouse of innovation. Previous successes of Stoica and his students include (but are certainly not limited to) Apache Spark, Ray, Apache Mesos and Alluxio.
- kossTKR 3y agoDoes this mean that GPT-4/65b level performance is closer to running on a say a m1/m2 with only 24+ gigabytes of ram?
- wskwon 3y agoNot really. vLLM optimizes the throughput of your LLM, but does not reduce the minimum required amount of resource to run your model.
- e12e 3y agoBut (in theory) - llama.cpp could implement similar approach to paging/memory and see a speedup for 4bit models on cpu?
- aftbit 3y agoNope. You will still need a proper GPU. You can't yet run large language models on tiny hardware like an m1/m2. Even the llama.cpp magic is only possible with very small models at beam size 1, which really limits the "creativity" of these models.
- kkielhofner 3y agoWe run into this constantly with Willow[0] and the Willow Inference Server[1]. There seems to be a large gap in understanding with many users. They seem to find it difficult to understand a fundamental reality: GPUs are so physically different and better suited to many/most ML tasks all the CPU tricks in the world cannot bring CPU even close to the performance of GPUs (while maintaining quality/functionality) for many tasks. I find this interesting because everyone seems to take it as obvious that integrated graphics vs discrete graphics for gaming aren't even close. Ditto for these tasks. With Willow Inference Server I'm constantly telling people: a six year old $100 Tesla P4/GTX 1070 walks all over even the best CPUs in the world for our primary task of speech to text/ASR - at dramatically lower cost and power usage. Seriously - a GTX 1070 is at least 5x faster than a Threadripper 5955WX. Our goal is to provide an open-source commercial voice assistant equivalent user experience and that is and will be fundamentally impossible for the foreseeable future on CPU. Slight tangent but there are users in the space who seem to be under the impression that they can use their Raspberry Pi for voice assistant/speech recognition. It's not even close to a fair fight but with the same implementation and settings a GTX 1070 is roughly 90x (nearly two orders of magnitude) faster[2] than a Raspberry Pi... Yes, all-in a machine with a GTX 1070 uses and order of magnitude more power (3w vs 30x) than a Raspberry Pi but the power cost in even countries with the most expensive power in the world results in a $2-$3/mo cost difference - which I feel, at least, is a reasonable trade-off considering the dramatic difference in usability (Raspberry Pi is essentially useless - waiting 10-30 seconds for a response makes pulling your phone out faster). [0] - https://github.com/toverainc/willow https://github.com/toverainc/willow [1] - https://github.com/toverainc/willow-inference-server https://github.com/toverainc/willow-inference-server [2] - https://github.com/toverainc/willow-inference-server/tree/wisng#benchmarks https://github.com/toverainc/willow-inference-server/tree/wi...
- bioemerl 3y agoI'm spoiled by 4 bit and unfortunately it doesn't appear to be supposed here so this isn't of much use to me, but it's awesome to see people working on the inference speed side of things regardless.
- george_123 3y agothis approach to managing KV cache can work with 4bit. imagine the speedup of pagedattention with quantization..
- zhisbug 3y agoyep, it is agonistic to 4-bit. You can deploy a 4-bit model and still use vllm + pagedattention to double or even triple your serving throughput.
- ynniv 3y agoIf this were submitted as a new comment it would be at the top of the page.
- ipsum2 3y agoprobably mean agnostic, agonistic implies the opposite.
- zhisbug 3y agooops typo
- baobabKoodaa 3y agoYou mean like, theoretically, in the future? Or you mean today?
- scv119 3y agoPretty cool stuff and the results are amazing. Hoping we will see virtual memory get standardized in pytorch or cuda.
- two_in_one 3y agoI wonder if this sort of memory management can be made for Pytorch transformers as under the hood optimization.
- brucethemoose2 3y agoReading between the lines, it sounds like some of the speedup comes from VRAM savings on an otherwise close to full GPU? This is definitely cool and needed, but it might not be so dramatic running 3-5 but quant on a less full GPU.
- wskwon 3y agoYes, vLLM focuses on maximizing throughput when the VRAM is fully utilized. Nevertheless, I believe users can still benefit from vLLM even if they don't utilize the memory to its full capacity, because vLLM also includes other optimizations orthogonal to the PagedAttention (e.g., optimized CUDA kernels).
- thomasahle 3y agoI wonder how this compares to Flash Attention (https://github.com/HazyResearch/flash-attention https://github.com/HazyResearch/flash-attention), which is the other "memory aware" Attention project I'm aware of. I guess Flash Attention is more about utilizing memory GPU SRam correctly, where this is more about using the OS/CPU memory better?
- ipsum2 3y agoThe ideas are orthogonal, and can be used (theoretically) at the same time.
- scv119 3y agoI believe you can slightly change the flash attention kernel to implement the same kernel of this page attention, since both of them work on the key/value cache at block level.
- karmasimida 3y agoI think they are orthogonal. Flash attention is just another way to compute exact attention. This work mainly concerns how to resolve memory fragmentation across different sequences You still need to compute attention as is once you retrieve the needed key values
- wskwon 3y agoThanks for the explanation! I believe the two ideas are basically orthogonal. FlashAttention reduces memory read/writes, while PagedAttention reduces memory waste.
- yieldcrv 3y agoParallelization, Paging! What will those AI/ML PhD’s think of next!
- quickthrower2 3y agoNow just waiting for a timing attack paper where you can see or guess someone else's conversation that is hosted in the same data center :-). Or maybe you get typically get dedicated machine time during inference?
- SimFG 3y agoCool, I prefer the OpenAI-Compatible api. Although this is not very technically difficult, it is really intimate, because it make me feel free to use all ChatGPT applications.
- wskwon 3y agoThanks! Please try it out and share any feedback you might have.