3 ms·
This is interesting stuff. I wonder if these sorts of tricks are already in use at the big labs. Incidentally, I would recommend trying implementing speculativ
by libraryofbabel 7mo ago
This is interesting stuff. I wonder if these sorts of tricks are already in use at the big labs.
Incidentally, I would recommend trying implementing speculative decoding yourself if you really want to understand LLM inference internals (that, and KV caching of course). I tried it over the Christmas holidays and it was a wonderful learning experience. (And hard work, especially because I forced myself to do it by hand without coding agent assistance.)
- born-jre 7mo agoi think this matters more for lower batch sizes (local llm and private enterprise deployment where there wont be big user at specific time for big batch size) going from mem Io bottleneck to compute.