4 ms·
You can amortize memory loading with large continuous batching. I imagine more compute would help the problem for certain workloads like speculative decoding
by dramlord 3y ago
You can amortize memory loading with large continuous batching. I imagine more compute would help the problem for certain workloads like speculative decoding
- qeternity 3y agoBatching helps throughput and anyone running in production will be doing batching. But it's not free, and still comes at a cost of per-stream latency. Speculative decoding seems less effective in practice than in theory.