3 ms·
Llama is trained with _more_ data than is chinchilla optimal in order to make it better and cheaper at inference time, instead of just getting the highest quali
by ntonozzi 4y ago
Llama is trained with _more_ data than is chinchilla optimal in order to make it better and cheaper at inference time, instead of just getting the highest quality of model that you can based on a given training budget. Llama has fewer parameters and was trained on more data specifically so that it would get high quality results on cheaper hardware and be easier and faster to run at inference time.