3 ms·
It's cheaper and faster to train a small model, which is better for a research team to iterate on, right? If Google decides that a particular small model is rea
by coder543 2y ago
It's cheaper and faster to train a small model, which is better for a research team to iterate on, right? If Google decides that a particular small model is really good, why wouldn't they go ahead and release it while they work on scaling up that work to train the larger versions of the model?
- summerlight 2y agoI have no knowledge of Google specific cases, but in many teams smaller models are trained upon bigger frontier models through distillation. So the frontier models come first then smaller models later.
- coder543 2y agoTraining a "frontier model" without testing the architecture is very risky. Meta trained the smaller Llama 3 models first, and then trained the 405B model on the same architecture once it had been validated on the smaller ones. Later, they went back and used that 405B model to improve the smaller models for the Llama 3.1 release. Mistral started with a number of small models before scaling up to larger models. I feel like this is a fairly common pattern. If Google had a bigger version of Gemini 2.0 ready to go, I feel confident they would have mentioned it, and it would be difficult to distill it down to a small model if it wasn't ready to go.
- int_19h 2y agoOn the other hand, once you have a large model, you can use various distillation techniques to train smaller models faster and with better results. Meta seems to be very successful doing this with Llama, in particular.
- coder543 2y ago> Meta seems to be very successful doing this with Llama, in particular. Kind of sort of: https://news.ycombinator.com/item?id=42391096 https://news.ycombinator.com/item?id=42391096