5 ms·
Look at how RingAttention is implemented: it's blockwise attention distributed among many GPUs, in other words bruteforce parallelization. For inference they us
by benob 3y ago
Look at how RingAttention is implemented: it's blockwise attention distributed among many GPUs, in other words bruteforce parallelization. For inference they use TPUs v4-128, not running this at home any time soon.
- DougMerritt 3y agoWhat is this month's best choice to run at home?
- benob 3y agoollama
- CuriouslyC 3y agoDepends on what you're looking for. Check reddit.com/r/localllama, but I'm guessing it's still Mixtral for most things, with Yi-34 fine tunes being useful in some cases since Mixtral fine tuning is still not there yet.
- okwhateverdude 3y agoIf you pull the llama.cpp repo and use their convert/quantize tools on the pytorch version of the models uploaded to huggingface, they will load just fine into ollama: https://old.reddit.com/r/LocalLLaMA/comments/18av9aw/quick_start_guide_to_converting_your_own_ggufs/ https://old.reddit.com/r/LocalLLaMA/comments/18av9aw/quick_s... https://github.com/ggerganov/llama.cpp/discussions/2948 https://github.com/ggerganov/llama.cpp/discussions/2948 You can run ollama (and a web UI) pretty trivially via docker: docker run -d --gpus=all -v /some/dir/for/ollama/data:/root/.ollama -p 11434:11434 --name ollama ollama/ollama:latest docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway --name ollama-webui ghcr.io/ollama-webui/ollama-webui:main That particular webui will let you upload models (with configuration). Other wise, you can use the api directly (you'll need to POST a `blob` first): https://github.com/ollama/ollama/blob/main/docs/api.md#create-a-model https://github.com/ollama/ollama/blob/main/docs/api.md#creat...
- dsrtslnd23 3y agoIs there any information on a suggested inference setup? I guess they had something different in mind than TPU v4-128 when they put it on HuggingFace?
- dsrtslnd23 3y agolooking at https://github.com/LargeWorldModel/LWM https://github.com/LargeWorldModel/LWM - they seem to indeed suggest to use a TPU vm
- ZeroCool2u 3y agoI suppose you could try with a Google Colab notebook attached to a free TPU instance? Probably would be quite limited if it worked at all.
- brucethemoose2 3y agoIts llama 7B, so anything that runs that. You can quantize the cache and fit quite a bit on GPUs. At least 75k on my mere 24GB 3090, maybe 200K with a fancy quantization repo.