4 ms·Llama.cpp speculative sampling: 2x faster inference for large models4 points by bobivl 3y agoxchip 3y agoSmall caveat: this is only true for generating text with simple grammar, like code. For human text this doesn't work so well.