2 ms·
New work from the vLLM team that disaggregates prefill and decoding to maximize goodput (throughput subject to latency constraints) in LLM serving
by zhisbug 3y ago
New work from the vLLM team that disaggregates prefill and decoding to maximize goodput (throughput subject to latency constraints) in LLM serving