Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
llmtosser
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
llmtosser
1y ago
Interesting - it does indeed seem like llama-server has the needed endpoints to do the model swapping and llama.cpp as of recently also has a new flag for the dynamic CPU offload now. However the approach to model swapping is not 'olla
2.
▲
by
llmtosser
1y ago
This is not true. No inference engine does all of: - Model switching - Unload after idle - Dynamic layer offload to CPU to avoid OOM
3.
▲
by
llmtosser
1y ago
Distractions like this probably the reason they still, over a year now, do not support sharded GGUF. https://github.com/ollama/ollama/issues/5245 If any of the major inference engines - vLLM, Sglang, llama.cp