7 ms·
This seems to be a wrapper around llama.cpp with several tuned parameter settings. Most of the speedup is gained by enabling speculative decoding [1]. No amaz
by smokel 25d ago
This seems to be a wrapper around llama.cpp with several tuned parameter settings. Most of the speedup is gained by enabling speculative decoding [1]. No amazing breakthroughs, just good tuning!
[1] https://en.wikipedia.org/wiki/Speculative_decoding https://en.wikipedia.org/wiki/Speculative_decoding
- wowitsbase 25d agoSpeculative decoding was on in all of my tests, the gain came from moving the round's inputs off host memory.