3 ms·
Very cool paper and some interesting things for local serving. Though its a little apples-to-oranges we do publish live energy stats for all models on our serv
by scottcha 16d ago
Very cool paper and some interesting things for local serving. Though its a little apples-to-oranges we do publish live energy stats for all models on our service here https://portal.neuralwatt.com/energy-pricing https://portal.neuralwatt.com/energy-pricing in case you are interested in what this looks like on the cloud side. FWIW DSV4.1 flash is really getting popular due to its IPW.
Some of the items like model routing, if you do it per request instead of per session, can break down on the cloud from an energy and cost POV since one of the best things you can do for both is to maintain the KV cache which both reduces time component of energy and the quite expensive prefill energy.
I am keen on the future where we have local/cloud hybrid serving which is cache aware. I do think that could be the best use of energy resources for AI.