2 ms·
I use llama.cpp w/ llama-swap https://github.com/ggml-org/llama.cpp https://github.com/ggml-org/llama.cpp https://github.com/mostlygeek/llama-swap https://gith
by computershit 2mo ago
I use llama.cpp w/ llama-swap
https://github.com/ggml-org/llama.cpp https://github.com/ggml-org/llama.cpp
https://github.com/mostlygeek/llama-swap https://github.com/mostlygeek/llama-swap
- LeBit 2mo agoThanks for the link to llama-swap. Didn’t know about it and will definitely install it.
- rancor 2mo agoFYI, llama-server can now be run in router mode so llama-swap is probably only needed for more exotic scenarios.
- kingo55 2mo agoI'm running it in router mode, but people on Reddit were recommending people use llama-swap instead. Am I missing something by using router mode?
- mudkipdev 2mo agoI believe it's useful for running multiple llama.cpp forks at the same time (e.g. a model you want requires special patching)
- computershit 2mo agoI’m using llama-swap because it can manage arbitrary backends, not just llama-server instances. I have llama.cpp chat and embedding models running alongside whisper-server all behind a single endpoint with per-model TTLs so they don't fight over the limited vram I have available on this box. Native routing could replace the llama.cpp part but not whisper so I guess I'm exotic ;)