4 ms·
By early 2028, major players like Intel, AMD, QC will ship accelerators in consumer laptops capable of running ~1T MoE models at ~100 tok/s
by typon 3mo ago
By early 2028, major players like Intel, AMD, QC will ship accelerators in consumer laptops capable of running ~1T MoE models at ~100 tok/s
- kennywinker 3mo agoUnless there are major improvements to how much hardware it takes to run a 1T model, this is deeply unrealistic. First because why release hardware that puts your biggest customers (data centers) out of business. Second because as I understand it the data centers have bought up all the high end chip production capacity for at least the next year and unless the bubble pops that'll continue for a while.
- iwontberude 3mo agoBecause for the company that will actually do it, their biggest customers aren’t data centers they are iPhone owners.
- kennywinker 3mo agoFirst off the math doesn’t math. Datacenters are willing to pay $50k for a single high end GPU. If you have unlimited capacity, yeah sell millions for $100 a pop or $10 a pop or whatever the bom cost of a phone GPU would be - but if you have limited capacity, you’re gonna sell all of that to the customer who is willing to pay the most PER UNIT. Second off, this doesn’t work from a power consumption standpoint. When I run qwen3.6-35b, a far smaller model than op is suggesting, power usage spikes to 150-200W during inference. To fit a 1T model in the palm of my hand, the amount of processing required doesn’t fit the amount of power available. Now I’m not saying this will never happen - there are some great leads, e.g. burning models directly on to a chip - but op’s scenario is definitely not happening in two years. Maybe 5, a lot more likely 10, unless of course local ai is made illegal
- andrekandre 3mo ago> Datacenters are willing to pay $50k for a single high end GPU. its true for now, because capital is flowing like a torrent, but how long will that last if returns start to be expected (aka the bubble pops)?
- kennywinker 3mo agoEven if the bubble pops and anthropic and openai et al implode - genie doesn’t go back in the bottle. The usefulness of LLMs for coding is proven, and a chip in a datacenter running 24/7 is always going to be more valuable than in a personal device running occasionally. That doesn’t change until production capacity exceeds the datacenter demand. When that happens, they’ll start selling them down the market until it eventually reaches phones and toasters and whatever. But not in two years.
- iwontberude 3mo agoLLMs for coding is too small of a benefit to justify this investment, the bubble is indeed going to burst. Genie is already on its way back into the bottle.
- kennywinker 3mo agoI agree it’s too small a benefit to justify the investment, and I agree the bubble will pop. I just don’t think that means hardware prices become sane again for quite a while. I think if you half the price of a server GPU because demand from the big ai companies drops out, we’ll still have a shortage - it’ll just being going into commodity data centers to run open weight models.
- typon 3mo agoThere is ton of room for improvement "down there". * Software inference optimizations * Heavy quantization * Chips with hardcoded transformer architecture * Much cheaper HBM * Much sparser models - 1T total with ~1-10B active params e.g. * Not to mention - 2 years of today's frontier models writing RTL and kernels at superhuman levels.
- veber-alex 3mo agoLiterally the only way this is going to happen is if aliens come to earth and gift us some amazing technology.
- typon 3mo agoYes that's called Mythos 2 or GPT 6