3 ms·
What's the recommended way to run LLMs these days? Ollama seems to work with DeepSeek R1 with enough memory using an older CPU but it's around 1 token/second o
by jfim 2y ago
What's the recommended way to run LLMs these days?
Ollama seems to work with DeepSeek R1 with enough memory using an older CPU but it's around 1 token/second on my desktop.
- CamperBob2 2y agoI've looked into it and the only sane answer right now is still "If it flies, floats, or infers, rent it." You need crazy high memory bandwidth for good inference speed, and that means GPUs which are subject to steep monopoly pricing. That doesn't look to be changing anytime soon. Second place is said to be the latest Macs with lots of unified memory, but it's a distant second place. The recently announced hardware from nvidia is either underpowered, overpriced, or both, so there's not much point waiting for it.
- jfim 2y agoMakes sense, thanks for sharing! I'll take it your recommendation to not buy also includes the upcoming project digits box from Nvidia?
- CamperBob2 2y agoFrom what I've seen and read (273 GB/s, wtf?), DGX Spark nee DIGITS is a non-starter. If this pricing turns out to be true: https://www.reddit.com/r/LocalLLaMA/comments/1jgnye9/rtx_pro_blackwell_pricing_listed/ https://www.reddit.com/r/LocalLLaMA/comments/1jgnye9/rtx_pro... ... then this generation of RTX Pro hardware sounds better to me. At the end of the day I don't know anything I didn't see on YouTube or /r/LocalLLaMA, though.