3 ms·
Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
- zhisbug 3y agoNew parallel decoding algorithm that trades flops for latency reduction
- atlas_hugged 3y agoOh sweet, this looks like a nice boost. Pretty simple too. Surprised it hasn’t been tried before that I’m aware of at least.