14 ms·
Instead of creating your own engine, would it really be that hard to add paged attention to llama.cpp?
by naasking 1mo ago
Instead of creating your own engine, would it really be that hard to add paged attention to llama.cpp?
- naasking 1mo agoThere's actually already a fork that implements the preliminaries: https://github.com/ggml-org/llama.cpp/discussions/21961 https://github.com/ggml-org/llama.cpp/discussions/21961