5 ms·
> Trained using the Chinchilla formula, these models provide the highest accuracy for a given compute budget. I'm confused as to why 111 million parameter mode
by eldenring 4y ago
> Trained using the Chinchilla formula, these models provide the highest accuracy for a given compute budget.
I'm confused as to why 111 million parameter models are trained with the Chinchilla formula. Why not scale up the training data? If you're training smaller models, surely optimizing performance is better than optimizing total compute.
Seems like a silly misunderstanding of the Chinchilla paper, but I'm sure I'm missing something
- gamegoblin 4y agoTrue. There was a good blog post published about this a few weeks ago: https://finbarr.ca/llms-not-trained-enough/ https://finbarr.ca/llms-not-trained-enough/ Money quote for those who don't want to read the whole thing: ''' When people talk about training a Chinchilla-optimal model, this is what they mean: training a model that matches their estimates for optimality. They estimated the optimal model size for a given compute budget, and the optimal number of training tokens for a given compute budget. However, when we talk about “optimal” here, what is meant is “what is the cheapest way to obtain a given loss level, in FLOPS.” In practice though, we don’t care about the answer! This is exactly the answer you care about if you’re a researcher at DeepMind/FAIR/AWS who is training a model with the goal of reaching the new SOTA so you can publish a paper and get promoted. If you’re training a model with the goal of actually deploying it, the training cost is going to be dominated by the inference cost. This has two implications: 1) there is a strong incentive to train smaller models which fit on single GPUs 2) we’re fine trading off training time efficiency for inference time efficiency (probably to a ridiculous extent). Chinchilla implicitly assumes that the majority of the total cost of ownership (TCO) for a LLM is the training cost. In practice, this is only the case if you’re a researcher at a research lab who doesn’t support products (e.g. FAIR/Google Brain/DeepMind/MSR). For almost everyone else, the amount of resources spent on inference will dwarf the amount of resources spent during training. '''
- haldujai 4y agoWhile true I think this also misses that “for almost everyone else” you’re probably not (or at least should not) be trying to optimize zero-shot performance if you have an intended high inference use case so I don’t think Chinchilla would be all that relevant.
- vintermann 4y agoI have a suspicion that good zero-shot performance is a good starting point for fine-tuning. If you have more than one intended high inference use case, or can imagine a couple of new ones on the horizon, it might still be best to not target the first use case directly.
- haldujai 4y agoWell yeah that’s kind of intuitive, my point is that if you just optimize for zero-shot you end up with something like GPT4 when an enterprise could probably be using finetuned LLaMA-7B with similar performance.
- sebzim4500 4y ago>Chinchilla implicitly assumes that the majority of the total cost of ownership (TCO) for a LLM is the training cost. In practice, this is only the case if you’re a researcher at a research lab who doesn’t support products (e.g. FAIR/Google Brain/DeepMind/MSR). For almost everyone else, the amount of resources spent on inference will dwarf the amount of resources spent during training. I'm not so convinced, especially if people are doing multiple training runs for hyperparameter tuning, cleaning data, fixing bugs, etc. I would be very interested in knowing what portion of OpenAI's compute budget is training. I would not be surprised if it was a significant minority.
- aiappreciator 4y ago"the training cost is going to be dominated by the inference cost." That's only true for general-mass-consumer models. Companies may want to fine-tune/train their own models, which don't have that many users for their narrow use cases (possibly only internal staff), will find that training cost is a substantial chunk of the TCO
- haldujai 4y agoYou’re not wrong, the Chinchilla rationale is that it may be more compute efficient to obtain a given loss using larger model sizes if the budget allows. As another commenter states this ignore the inference part of the equation. As an example the BERT/RoBERTa family were trained for much longer than Chinchilla, you do get diminishing returns though. There is a point of overtraining where downstream performance is impacted but that’s pretty high. I think part of the answer to this is also that xxx million parameter decoder-only models don’t seem to be that useful so it may not be worthwhile to optimize them for performance?
- ftxbro 4y agoThe point of those smaller models is for the "Cerebras Scaling Law for Compute-Optimal Training" which is the straight line plot in the image at the top of their webpage when you click the link. They want you to think it's reasonable that because the line is so straight (on a flops log scale) for so long, it could be tempting to extrapolate the pile-loss consequences of continuing compute-optimal training for larger models beyond their largest 13B one, with the obvious caveat that the extrapolation can't continue linearly much further if for no other reason than the test loss isn't going to go below zero (it will flatten out sooner than that). If you trained beyond compute-optimality on smaller models, it would mess up their straight line and make it look like we are sooner hitting diminishing returns on test loss.
- bjornsing 4y ago> the extrapolation can't continue linearly much further if for no other reason than the test loss isn't going to go below zero Isn’t the test loss logarithmic? If so it sure can go below zero.
- deleted 4y ago[deleted]
- ftxbro 4y agoAccording to https://pile.eleuther.ai/paper.pdf https://pile.eleuther.ai/paper.pdf the test loss on the pile is the log of the perplexity, and the perplexity is 2^H where H is an entropy which is non-negative. So the perplexity is always at least one, so its log is always at least zero. So yes the test loss can be seen as a log, but no it's not allowed to go below zero. The intuition is that the test loss is the number of bits that the model would need on average to encode each next token in the test part of the pile, given that you have seen the preceding parts.
- bjornsing 4y agoGood point!