2 ms·
Quite. The big question at the time was "how much data do we need to train GPT-3 equivalent models". Open models had failed to live up to GPT performance, even
by ijk 2y ago
Quite. The big question at the time was "how much data do we need to train GPT-3 equivalent models". Open models had failed to live up to GPT performance, even ones with a massive number of parameters. So getting results that suggested a reason why other models were massively undertrained was important.
Meanwhile, people noticed that for deployed models, inference cost often outweighs the initial training costs. It's sometimes better to train a smaller, faster model longer on more data, because it has lower overall cost (including environmental impact) if you're expecting to run the model a few million or billion times (e.g., [1]). So training past the Chinchilla optimum point became a lot more common, particularly after Llama.
[1] https://arxiv.org/abs/2401.00448 https://arxiv.org/abs/2401.00448