4 ms·Cutting LLM Batch Inference Time by Half with Dynamic Prefix Bucketing2 points by DISCURSIVE 11mo ago