3 ms·
Nothing would really change IMO? 99% of users don't have anything like a RTX5070 (mobile especially). Even if it did, it still doesn't make much economic sense
by martinald 2mo ago
Nothing would really change IMO? 99% of users don't have anything like a RTX5070 (mobile especially).
Even if it did, it still doesn't make much economic sense running a model locally vs on a datacentre.
For example, I managed to just about squeeze a Q2 quant of Qwen 3.7 27b on my 9070XT. I get around 60tps decode (slightly faster prefill). _but_ it uses 300W of power to do so. At UK electricity rates of 30c/kWh this works out at something like 42c/MTok. I can get far far better models on openrouter cheaper than that, plus I'm not horrendously constrained on context length.
- protocolture 2mo agoI dunno a lot of things said about AI economics sound like an IBM executive making reassuring statements about their terminal/mainframe business before the personal computer took off. Like even if you run it in a datacenter in this scenario, you could do it on a cheap GPU instance in Azure, you still wouldnt need OpenAI or Anthropic specific clouds. >uses 300W of power to do so. There are plenty of people with phat electricity pipes in their on prem server rooms that have been vacated for cloud. Companies who want the benefits of AI but dont want the risk of sending their data to foreign API endpoints.
- JacobAsmuth 2mo agoYou're suggesting that if a very good and cheap AI model came out tomorrow everyone would rush out to rent Azure instances to run batch size 1 inference on their model?
- protocolture 2mo agoI am suggesting that Azure and AWS would change course and push corporate customers towards more expensive, but more private options.
- Jlagreen 2mo agoThe analogy with IBM mainframe completely ignores Murphy's law which came up and lead to the small and fast chips we have today. But Murphy's law is dead. No future chip will leapfrog easily current chips because we have reached hard phyical limits in chip density and downsizing. Huang's law by Jensen Huang focuses on something else and that is token performance per Watt at scale. Blackwell needs double TDP than Hopper and Rubin again needs almost double TDP on a rack but in the end Rubin will be like 100x token performance per watt on a scaled data center. This means you have more energy need but you get multiples of token performance because you start scaling in the data center. The local chip will never be able to keep up with the data center scaling economics. This is why everyone is so crazy about building data centers because they can see the economocs behind it. What people don't seem to understand if tokens become more available and cheaper then not only more people can use them but a single person can use more as well. Why should you be limited to 1 AI agent? Why can't have you have multiple agents running on multiple devices daily for you? This is why demand will grow exponentially with the growth of token economics. We have seen it for the last few years and much more is yet to come.
- protocolture 2mo agoDo... do you mean Moores Law? >The local chip will never be able to keep up with the data center scaling economics. Assumes the software has been completely solved. >Why should you be limited to 1 AI agent? Why can't have you have multiple agents running on multiple devices daily for you? At some point we cap out the bandwidth of the human to keep up with their mistakes.