4 ms·
I've been running some of the larger models (like Llama 405B) via CPU on a Dell R820. It's got 32 Xeon's (4 chips), and 256 GB RAM. I bought it used for aroun
by thijson 2y ago
I've been running some of the larger models (like Llama 405B) via CPU on a Dell R820. It's got 32 Xeon's (4 chips), and 256 GB RAM. I bought it used for around $400. The memory is NUMA, so it makes sense if the computing is done on local data, not sure if Ollama supports that.
The tokens per second is very slow though, but at least it can execute it.
I think the future will be increasingly more powerful NPU's built into future CPU's. That will need to paired with higher bandwidth memory, maybe HBM, or silicon photonics for off chip memory.