3 ms·
What device do I need and how much will it cost to install one at home so that it works as quickly as the Claude Code answer (and it answers quite slowly)?
by mrwaip 2mo ago
What device do I need and how much will it cost to install one at home so that it works as quickly as the Claude Code answer (and it answers quite slowly)?
- SwellJoe 2mo agoIf you want it to respond in under a minute, you need more hardware than this application is intended for. This thing's response is measured in seconds per token, not tokens per second. To get Claude Code responsiveness from even a pretty small (but still usable) model, you need, maybe two 32GB GPUs? I run Gemma 4 31B and Qwen 3.6 27B on my dual 32GB GPU setup (cheap old Radeon Pro V620 GPUs) at about 20 t/s, which is not fast enough for comfortable interactive agentic use. A couple of new GPUs, like Radeon AI Pro 9700 at $1400 each, probably gets you fast enough for comfortable interactive use with small models like those. Those small models are not competitive with Claude models, however (maybe they beat Haiku sometimes). They can write a little Python or make a web page, they can't architect a real application. To run a near-frontier model like Kimi K3 or GLM 5.2 at comfortable speeds, you need serious hardware with 768GB VRAM, minimum. I think Asus is releasing something like that for about $150k soon. You can run DeepSeek V4 Flash at almost comfortable speeds and in a decently capable quantization on two DGX Sparks or Asus GX10s (about $10,000). Or, you could use DeepSeek V4 Flash from DeepSeek.com, at blistering speeds and with huge contexts, for something like a decade or two for that same $10,000.
- f6v 2mo ago> two DGX Sparks or Asus GX10s (about $10,000) > for something like a decade or two for that same $10,000 I've been using DS V4 Flash through OpenCode and it's mind-blowing how I get near-SOTA AI model for the price of two coffees. But let's put the broken economy of AI APIs aside for a second. Running such a model with under 6 figures of fixed costs is equally mind-blowing. And we know that hardware costs are real, Nvidia is selling those at a profit. I'm not talking about consumers, but businesses. I remember paying much more for something more trivial, like Datadog.
- SwellJoe 2mo agoDeepSeek feels like the most...trying to think of the word, honest isn't quite right but close, of the AI companies. They make very good models that are phenomenally efficient and they sell them at what seems to be a modest margin. They're just running a business. Doesn't feel like a Ponzi scheme. It doesn't feel like they're trying to take over the world and be the only AI company through regulatory capture or cornering the market on RAM and compute, they're not adding "memory" and personality tuned to maximize psychosis or addiction. They're not benchmaxxing the way Kimi and Qwen are benchmaxxed; you can find the holes and weirdness in those models capabilities pretty quickly. DeepSeek models feel like good all-rounders. Not the best at anything, but you won't be shocked by it doing stupid shit, either. Qwen keeps stunning me with really dumb bugs it creates even while being able to work on really big really hard problems, it's surprisingly sloppy. I dunno. I don't have good feelings about most of the AI companies; I feel like they're trying to be predatory and monopolistic. I don't get that feeling from DeepSeek, they seem like a decent company making a good product at a fair price. So, they're consistently my choice for API usage, even if I still use the best American models for agentic use, via a subscription.