3 ms·
My guess would be that they tried to save money with speculative decoding and they had too loose thresholds for the verification stage. As someone who has impl
by fxtentacle 1y ago
My guess would be that they tried to save money with speculative decoding and they had too loose thresholds for the verification stage.
As someone who has implemented this myself, I know that it’s pretty easy to make innocent mistakes there. And the only visible result is a tiny distortion of the output distribution which only really becomes visible after analysing thousands of tokens. And I would assume that all providers are using speculative decoding by now because it’s the only way to have good inference speed at scale.
As a quick recap, you train a small model to quickly predict the easy tokens, like filler words, so that you can jump over them in the recurrent decoding loop. That way, a serial model can predict multiple tokens per invocation, thereby easily doubling throughput.
And the fact that they need lots of user tokens to verify that it works correctly would nicely explain why it took them a while to find and fix the issue.
- metadat 1y agoSpeculative Decoding, for the uninitiated (like me..): https://research.google/blog/looking-back-at-speculative-decoding/ https://research.google/blog/looking-back-at-speculative-dec...
- buildbot 1y agoStandard speculative decoding without relaxed acceptance has no accuracy impact as far as I understand things. If you always run the verification; you always have the true target model output.
- fxtentacle 1y agoYou need the relaxed acceptance to get those cost savings. Every time you determine your small model to be "good enough", it allows the large model to skip one iteration in its recursive decoding loop. You are correct in the sense that you can estimate the probability that the large model would have had for predicting the token(s) that your small model chose, but you don't know which tokens the large model might have predicted based on tokens that the small model did not predict. Unless, of course, your "small" model becomes as precise as the large model, at which point it's not small anymore. In other words: The speculative decoding causes "holes" in your beam search data. You can fill them by sampling more, increasing hosting costs. Or you fill them with approximations, but that'll skew the results to be more "safe" => more generic, less reasoning.
- mickdarling 1y agoWould that show up more with heavily cached data or less? Because I've been using very heavily cached data on the order of 30 to 1 cache versus not. And I haven't had much trouble with Claude Code at all in the last month. A few problems here and there, but they were sporadic. While I know other people have had massive issues, I wonder if the issues that occurred didn't effect cache as much and thus prevented me from seeing some of the worst aspects of it.