4 ms·
An energy-efficiency leaderboard for LLMs
- brucethemoose2 3y agoThis is... Kinda not useful? Users aren't running HF Transformers LLaMA 7B on an A100. They are running higher parameter count finetunes highly quantized on exLLaMA, GPTQ or llama.cpp, usually with clamped (and sometimes forced) response lengths, usually on cheaper GPUs unless they are running LLaMA 65B (or Falcon 40B). All the LLaMA finetunes consume the same amount of energy at the same parameter count unless they are doing something exotic like an 8K context with superhot, and response length from those tests is kinda a questionable metric.
- jaywonchung 3y agoYep, I completely agree with your first paragraph, and adding quantization & distributed inference is one of the bigger next steps. Adding numbers from consumer GPUs would be nice for non-data center use cases, too. Decoder-only models with the same number of parameters will roughly consume the same amount of energy per token. But what makes the response length meaningful is that we actually ran all the models on the same set of prompts from ShareGPT, meaning that some models are innately more verbose/terse than others. These properties are also not completely aligned with model quality benchmarks, which makes an interesting tradeoff.