3 ms·
Hi! one of the contributors to the paper — we have kernels not released yet that can shave down decoding latency by >20%. Also when we ran experiments for stre
by ow5 1y ago
Hi! one of the contributors to the paper — we have kernels not released yet that can shave down decoding latency by >20%.
Also when we ran experiments for streaming with the current kernels, we were median ~1.3x slower at inference
- ein0p 1y agoThanks for chiming in! How do you explain the top-most graph in Figure 5? Am I misreading it?