3 ms·
Does anyone else feel like the writing is on the wall for a future of local models? Spamming data centres everywhere, powering them, having to commit insane cap
by mindwok 2mo ago
Does anyone else feel like the writing is on the wall for a future of local models? Spamming data centres everywhere, powering them, having to commit insane capital to hardware, all the effort to serve inference over a network reliably - when here we are with a frontier model nearly running on a laptop.
Local AI on your device seems like a much more likely future to me than datacenters in space. For inference at least, training is another story.
- nomel 2mo ago0.01 tk/s on an M1 Max is not "nearly". This is completely unusable, and in no way cost effective. 0.01 tokens per second means 1 million tokens ($3 worth of API usage [1]) takes 3.2 YEARS. [1] https://www.kimi.com/resources/kimi-k3-pricing https://www.kimi.com/resources/kimi-k3-pricing
- mindwok 2mo agoOk in terms of running a 2.8T parameter model, that's true. Looking more broadly though, a model I can run on my laptop (Gemma 4) is ~4 points away from GPT-5.3 codex or Sonnet 4.5 on arena.ai LLM leaderboard. Those models were SOTA less than a year ago.
- nomel 2mo agoAnd the leaderboard has Opus 4.8 thinking +2 points away from Gemini 3.6 flash. Try to use Gemini flash in an agentic coding harness and see what comes out.