4 ms·
> We’re releasing Llama 3.1 405B Is it possible to run this with ollama?
by throwaway_2494 2y ago
> We’re releasing Llama 3.1 405B
Is it possible to run this with ollama?
- jessechin 2y agoSure, if you have a H100 cluster. If you quant it to int4 you might get away with using only 4 H100 GPUs!
- sheepscreek 2y agoAssuming $25k a pop, that’s at least $100k in just the GPUs alone. Throw in their linking technology (NVLink) and cost for the remaining parts, won’t be surprised if you’re looking at $150k for such a cluster. Which is not bad to be honest, for something at this scale. Can anyone share the cost of their pre-built clusters, they’ve recently started selling? (sorry feeling lazy to research atm, I might do that later when I have more time).
- vorticalbox 2y agoIf you have the ram for it. Ollama will offload as many layers as it can to the gpu then the rest will run on the cpu/ram.
- deleted 2y ago[deleted]
- Havoc 2y agoIf you want your first token around tomorrow lunch sure