4 ms·
Is that just because nobody has made an effort yet to port them upstream, or is there something inherently difficult about making those changes work in llama.cp
by breakingcups 2y ago
Is that just because nobody has made an effort yet to port them upstream, or is there something inherently difficult about making those changes work in llama.cpp?
- leblancfg 2y agoI get the impression most llama.cpp users are interested in running models on GPU. AFAICT this optimization is CPU-only. Don't get me wrong – a huge one! – and opens the door to running llama.cpp on more and more edge devices.