3 ms·
This isn't overfitting. Chinchilla optimises training compute. LLaMa optimises inference compute. Overtraining (according to Chinchilla), not overfitting.
by tomp 2y ago
This isn't overfitting.
Chinchilla optimises training compute.
LLaMa optimises inference compute. Overtraining (according to Chinchilla), not overfitting.