3 ms·
Throughput Is Not All You Need: Maxing Goodput in LLM Serving via Disaggregation
- zhisbug 3y agoNew work from the vLLM team that disaggregates prefill and decoding to maximize goodput (throughput subject to latency constraints) in LLM serving
3 ms·