2 ms·
If you pull the llama.cpp repo and use their convert/quantize tools on the pytorch version of the models uploaded to huggingface, they will load just fine into
by okwhateverdude 3y ago
If you pull the llama.cpp repo and use their convert/quantize tools on the pytorch version of the models uploaded to huggingface, they will load just fine into ollama:
https://old.reddit.com/r/LocalLLaMA/comments/18av9aw/quick_start_guide_to_converting_your_own_ggufs/ https://old.reddit.com/r/LocalLLaMA/comments/18av9aw/quick_s...
https://github.com/ggerganov/llama.cpp/discussions/2948 https://github.com/ggerganov/llama.cpp/discussions/2948
You can run ollama (and a web UI) pretty trivially via docker:
docker run -d --gpus=all -v /some/dir/for/ollama/data:/root/.ollama -p 11434:11434 --name ollama ollama/ollama:latest
docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway --name ollama-webui ghcr.io/ollama-webui/ollama-webui:main
That particular webui will let you upload models (with configuration). Other wise, you can use the api directly (you'll need to POST a `blob` first):
https://github.com/ollama/ollama/blob/main/docs/api.md#create-a-model https://github.com/ollama/ollama/blob/main/docs/api.md#creat...