4 ms·
If this is "plug and play" can it be added to say llama.cpp and give ~3.67 speedup to existing models or is there some complication?
by MyFirstSass 2y ago
If this is "plug and play" can it be added to say llama.cpp and give ~3.67 speedup to existing models or is there some complication?
- jncraton 2y agoThe speedup would not be that high in practice for folks already using speculative decoding[1]. ANPD is similar but uses a simpler and faster drafting approach. These two enhancements can't be meaningfully stacked. Here's how the paper describes it: > ANPD dynamically generates draft outputs via an adaptive N-gram module using real-time statistics, after which the drafts are verified by the LLM. This characteristic is exactly the difference between ANPD and the previous speculative decoding methods. ANPD does provide a more general-purpose solution to drafting that does not require training, loading, and running draft LLMs. [1] https://github.com/ggerganov/llama.cpp/pull/2926 https://github.com/ggerganov/llama.cpp/pull/2926
- MacsHeadroom 2y agoWho is already using speculative decoding? I haven't seen anything about it in the llama.cpp or ollama docs.
- eshoyuan 2y agohttps://github.com/ggerganov/llama.cpp/tree/master/examples/speculative https://github.com/ggerganov/llama.cpp/tree/master/examples/...
- tripplyons 2y agoThe HuggingFace transformers library already has support for a similar method called prompt lookup decoding that uses the existing context to generate an ngram model: https://github.com/huggingface/transformers/issues/27722 https://github.com/huggingface/transformers/issues/27722 I don't think it would be that hard to switch it out for a pretrained ngram model.