2 ms·
For the impatient, I merged llama.cpp tentative branches to get it running here https://github.com/alainnothere/llama.cpp/tree/disk-cache-eviction https://githu
by xlayn 1mo ago
For the impatient, I merged llama.cpp tentative branches to get it running here https://github.com/alainnothere/llama.cpp/tree/disk-cache-eviction https://github.com/alainnothere/llama.cpp/tree/disk-cache-ev..., thing runs at 23.54 token/sec and my setup runs at high 30 the 3.8 dense 27B.
and this is the pelican from the iq4_xs model
https://github.com/alainnothere/llama.cpp/blob/disk-cache-eviction/media/pelicanRidingBycicle.jpg https://github.com/alainnothere/llama.cpp/blob/disk-cache-ev...
- cmrdporcupine 1mo agoIf you've got a DGX Spark try my little engine: https://github.com/rdaum/eider/ https://github.com/rdaum/eider/ NVFP4 quant