3 ms·
Yep, I completely agree with your first paragraph, and adding quantization & distributed inference is one of the bigger next steps. Adding numbers from consumer
by jaywonchung 3y ago
Yep, I completely agree with your first paragraph, and adding quantization & distributed inference is one of the bigger next steps. Adding numbers from consumer GPUs would be nice for non-data center use cases, too.
Decoder-only models with the same number of parameters will roughly consume the same amount of energy per token. But what makes the response length meaningful is that we actually ran all the models on the same set of prompts from ShareGPT, meaning that some models are innately more verbose/terse than others. These properties are also not completely aligned with model quality benchmarks, which makes an interesting tradeoff.