3 ms·
The router mode and matrix routing in llama.cpp is still early days and it can't easily juggle multiple models as easily as llama-swap so there's still benefits
by cptskippy 2mo ago
The router mode and matrix routing in llama.cpp is still early days and it can't easily juggle multiple models as easily as llama-swap so there's still benefits if you're using 24-32GB cards that can run multiple models simultaneously.
The gulf is quickly shrinking though and it seems like the need for llama-swap will disappear soon.