3 ms·
> our 770M T5 model outperforms the 540B PaLM model using only 80% of available data on a benchmark task. That sounds too good, can someone more knowledgeable
by davidkunz 3y ago
> our 770M T5 model outperforms the 540B PaLM model using only 80% of available data on a benchmark task.
That sounds too good, can someone more knowledgeable comment?
- verdverm 3y agoLarge model trains small model, I suspect you end up with a better training session than purely unsupervised. You can think of it like transfer learning > Knowledge distillation has been successfully used to transfer knowledge from larger, more competent teacher models into smaller student models affordable for practical applications It sounds like a bit of Chain of Thought going on too > We propose a new paradigm, Distilling step-bystep, that leverages the ability of LLMs to reason about their predictions to train smaller models in a data-efficient way. If we anthropomorphize a bit, you can think of this as a person reading all the things and learning on their own vs learning with a mentor, which results in much better results for humans. So maybe this is an interesting, expected result. It will also be interesting to see how this can be combined with DeepMind's RETRO ideas. https://web.mit.edu/5.95/readings/bloom-two-sigma.pdf https://web.mit.edu/5.95/readings/bloom-two-sigma.pdf
- marcyb5st 3y agoGoogler here, but not in the research team that authored this paper. It seems that using a "teacher" LLM to train a smaller model in this step-by-step fashion you get much more out of your parameters. Specifically, if you look at section 3 in the paper, they mention that they use LLMs generated rationales as additional guidance for the smaller model. This approach has already been tried, but it had some limitations (section 3.2) which they circumvent by doing things differently: "In this work, instead of using rationales as additional model inputs, we frame learning with rationales as a multi-task problem. ... (this) enables the model to learn to generate the intermediate reasoning steps for the prediction, and could therefore guide the model in better predicting the resultant label." It seems that this is the key for the success they are showing. I am still digesting the paper and I am not 100% sure that is the key factor.
- svantana 3y agoThe wording makes it seem that the small one outperforms the large one, all else equal. But PaLM is tested in the few-shot setting, specifically chain-of-thought prompting, which means: 770M model uses 80% of dataset 540B model uses 0.01% of dataset
- verdverm 3y agoPart of what is going on here is that this paradigm is to train a smaller, task-specific model from a large, generalized model. When comparing to other methods for doing this, it outperforms. So one of these smaller, specialized models can outperform PaLM on some tasks, but not all. The other benefit comes at inference time, where you can get much faster and cheaper outputs.
- riku_iki 3y agofrom abstract, it looks like it was one task. It is usually much easier to build much more performant specialized model for single task, while it will be weak in many other tasks compared to larger but generalized model.