5 ms·
This is not true. No inference engine does all of: - Model switching - Unload after idle - Dynamic layer offload to CPU to avoid OOM
by llmtosser 1y ago
This is not true.
No inference engine does all of:
- Model switching
- Unload after idle
- Dynamic layer offload to CPU to avoid OOM
- ekianjo 1y agothis can be added to llama.cpp with llama.swap currently so even without Ollama you are not far off