4 ms·
Larger models are more expensive to run (ceteris paribus). But we're seeing we can squeeze more performance from smaller models. You need to compare like-for-l
by silveraxe93 2y ago
Larger models are more expensive to run (ceteris paribus). But we're seeing we can squeeze more performance from smaller models.
You need to compare like-for-like. You can't say that the cost of building a 5-story apartment is increasing by pointing at the burj khalifa.
- menaerus 2y agoNow remind us what HW did we need to run local inference of llama2-69B (July, 2023)? And then contrast it to the HW we need to run llama3.1-70B (July, 2024)? In particular, which optimizations and in what way did they dramatically cut down the cost of the inference? I seriously don't get this argument and I see it being repeated all over and over again. Although model possibilities are increasing, no doubt in that, HW costs for inference remained the same and they're mostly driven by the amount of (V)RAM you need.