4 ms·
you can pull directly from huggingface with llama.cpp, and it also has a decent web chat included
by pheggs 6mo ago
you can pull directly from huggingface with llama.cpp, and it also has a decent web chat included
- speedgoose 6mo agoDoes it have a model registry with an API and hot swapping or you still have to use sometime like llama swap as suggested in the article ? Or is it CLI?
- dminik 6mo agoYou can have multiple models served now with loading/unloading with just the server binary. https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md#model-presets https://github.com/ggml-org/llama.cpp/blob/master/tools/serv...
- speedgoose 6mo agoIt only lacks the automatic FIFO loading/unloading then. Maybe it will be there in a few weeks.