3 ms·
When efficiency reaches the point where local models on consumer hardware are good enough, demand for cloud tokens could rapidly shrink.
by adrianN 2mo ago
When efficiency reaches the point where local models on consumer hardware are good enough, demand for cloud tokens could rapidly shrink.
- pseudosavant 2mo agoIt could be quite a while before we reach that point though. 5+ years easily. I've been keenly interested in the ability to run local models, but the hardware is just not there. Consumer RAM speeds and capacity will have to significantly increase before local models will be able to perform as well as even the lowest end GPT-5.6 Luna model. This is on the backdrop of RAM becoming prohibitively expensive. And without the speed and quantity of RAM, it becomes impossible to generate tokens at interactive speeds, regardless of model. There is a fundamental dependency between calculating all of the active params with the given RAM speed. Even with a model that has been quantized all the way down to Q4, the DGX/RTX Spark chip with 128GB of RAM can only generate ~18 tokens/sec for a MoE model with only 30B active parameters. There haven't been any broadly useful models below 30B active parameters. And that is for a $5000+ piece of hardware that will be one of the best for running on-device models. I really want to buy instead of rent my AI, but the economics are truly terrible.
- bryanlarsen 2mo agoVery few consumers are going to spend multiple thousands of dollars to save $10 per month. Companies absolutely will to save hundreds per month per employee, but that's not consumer hardware.
- HDBaseT 2mo agoMany gamers already spend $1000+ on a GPU. If you can integrate AI accelerators into consumer cards (you can), you can have local AI for "reasonably" cheap. This is Nvidia's long term goal if you listen to what Jensen has to say. The limitation is entirely on memory right now. Just a few years ago we could of been strapping 80-100GB to cards for under $200 (BoM).
- adrianN 2mo agoWell if the trend that the comment further up in this thread claimed continues and compute requirements keep dropping exponentially then perhaps in a few years you can have today’s frontier performance on the normal laptop you already have on your desk anyway.