3 ms·
Almost everyone is using stuff like tensorrt which is far from platform agnostic. But if AMD can build an API-compatible library that performs well it should in
by progbits 2y ago
Almost everyone is using stuff like tensorrt which is far from platform agnostic. But if AMD can build an API-compatible library that performs well it should indeed be easy enough to switch.
- qeternity 2y agovLLM is widely used (and what is used in this benchmark) which is far easier to setup and maintain than TensorRT. vLLM is not quite as performant but it's pretty close in production environments.
- KaoruAoiShiho 2y agoIs it really that close in production? The benchmarks I've seen show it's almost 3x slower compared to the SOTA. https://bentoml.com/blog/benchmarking-llm-inference-backends https://bentoml.com/blog/benchmarking-llm-inference-backends
- qeternity 2y agoYou say benchmarks (plural) but I see just one here, and this benchmark does not add up in my experiences. But let's take L3 8b in fp16 (the first graph) which shows the performance to be comparable on throughput (vLLM faster on TTFT across the board). On the 4bit L3 70b test, they are using AWQ. GPTQ with Marlin kernels (which now gets repacked on the fly in vLLM) is much faster in our tests, and has much better optimizations amongst vLLM contributors. So yeah, this is sample size of n=1, as are my own tests. But my experiences have been much closer to the L3 8b fp16 charts, so I'm going to presume the bigger delta comes down to that. EDIT: here is one of our own recent internal tests where we observed Marlin performing 25-30% better than AWQ in vLLM - https://miro.medium.com/v2/resize:fit:1400/format:webp/1*F9i8w_ytuKO3_BrThfGvqg.png https://miro.medium.com/v2/resize:fit:1400/format:webp/1*F9i...
- KaoruAoiShiho 2y agoThat's cool but a 25-30% improvement wouldn't make up for the 300% advantage LMDeploy and TRT has over vllm on the 70b benchmark. Where are your results for vllm vs lmdeploy and trt?
- qeternity 2y agoI'm not suggesting it's down to that. I'm suggesting that anyone running AWQ on vLLM is probably not running it optimally somewhere else. I don't have a chart handy, but in our tests, TRT was ca. 10% faster (but it's much more difficult to setup, you have to convert models, etc). LMDeploy we tested months ago, maybe it's improved now, but they were making fairly wild claims that we didn't observe. Basically, you shouldn't publish a benchmark without providing all of the details. How long did test run? How were requests staggered? What were the input/output token sizes? A lot of these 256 in / 256 out benchmarks are useless. If your system prompt is 4k tokens, prefill becomes a serious issue. How performant is your continuous batching? Can you do chunked prefill? Are you running prompt prefix caching (and if so, how performant is that)? It's a lot more complex than simply generating some tokens at various batch sizes.
- KaoruAoiShiho 2y agoYeah, that makes sense. Here's another benchmark that shows trt being about 2-3x faster, 2x on fp16 and 3x on int4. https://blog.premai.io/prem-benchmarks/ https://blog.premai.io/prem-benchmarks/ Though to be fair it's 512 tokens and not 4k tokens. Would love to see more public testing done for all the different variations.
- qeternity 2y agoYep, agreed on the more testing. I don't actually have any allegiance to vLLM. We've just been happy with it. A few things in the linked article that make my eyebrows raise. Notably, they claim to achieve 206 t/s at bs=1 on a single A100 80GB (2 TB/s bandwidth) in fp16. That model is 14.5GB in size, and best case would require aggregate memory bandwidth of almost 3 TB/s to achieve. The BentML you linked to, also ran on an A100 80GB, and achieved ~650 t/s at bs=10 so roughly 65 t/s per stream. Granted this is for an 8B param model, not 7B, and it will of course be faster at bs=1 but absolutely not by this margin. In this same benchmark, they achieved roughly 45 t/s per stream at bs=50 so you can see the sort of scaling we're dealing with. At bs=10 you should achieve relatively similar per-stream throughput to bs=10. Baseten did a benchmark with TensorRT-LLM and an A100 80GB on Mistral 7B and at bs=1 they achieved 75 t/s and at bs=8 they achieved 66 t/s per stream (https://www.baseten.co/blog/unlocking-the-full-power-of-nvidia-h100-gpus-for-ml-inference-with-tensorrt/ https://www.baseten.co/blog/unlocking-the-full-power-of-nvid...) I strongly suspect there is something very wrong with the Premai (never heard of them before) or there has been some huge inference breakthrough that I am unaware of. The BentoML benchmarks look pretty good, and I suspect there is some vLLM performance left on the table, however not enough to close the gap. In our testing, TensorRT-LLM was definitely faster, but not enough to warrant all of the other headaches.