23 ms·
You need the relaxed acceptance to get those cost savings. Every time you determine your small model to be "good enough", it allows the large model to skip one
by fxtentacle 1y ago
You need the relaxed acceptance to get those cost savings. Every time you determine your small model to be "good enough", it allows the large model to skip one iteration in its recursive decoding loop. You are correct in the sense that you can estimate the probability that the large model would have had for predicting the token(s) that your small model chose, but you don't know which tokens the large model might have predicted based on tokens that the small model did not predict. Unless, of course, your "small" model becomes as precise as the large model, at which point it's not small anymore.
In other words: The speculative decoding causes "holes" in your beam search data. You can fill them by sampling more, increasing hosting costs. Or you fill them with approximations, but that'll skew the results to be more "safe" => more generic, less reasoning.