3 ms·
> You prompt the LLM to generate a bunch of arithmetic questions, prompting it to show its working. Then you take that output, remove the intermediate steps, an
by time_to_smile 3y ago
> You prompt the LLM to generate a bunch of arithmetic questions, prompting it to show its working. Then you take that output, remove the intermediate steps, and train on the results.
Removing or editing the output of the model is providing new information to the model, what's improving the performance of the model in this scenario is that you are explicitly adding new information and fine tuning it on these new cases.
> information theory is not especially relevant here
It's extremely relevant because people seem to be arguing with about mathematical facts as though they were somehow opinions.
You cannot improve the performance of a model without adding new information to that model.
- sebzim4500 3y agoWhat information is being provided? I'm not suggesting that you would do any kind of check to see if the new training data is correct. Just take the output and remove every line between the first and last. Deterministically removing data is not adding information, unless you are defining information in an extremely unusual way.
- deleted 3y ago[deleted]
- YeGoblynQueenne 3y ago(Not the OP) What I wonder about is why would performance increase with what you suggest. In my mind, you're proposing getting a language model to generate its own training data. But data on its own is not enough to train a good model. The data has to have some kind of information content that guides the learning algorithm to select a good model. If you feed the training algorithm data that doesn't increase its information about what a good model is, you're not helping it train a better model. For the record, I tried something like what you suggest when I was doing my Master's. That was back in 2014, and I had to train a classifier for a machine learning class. I was given a training set while a separate validation set and test set were kept private (it was all set up in Kaggle as a private competition). To clarify, the idea was that you trained your classifier of choice on the training set, then labelled the validation set with your trained model and submitted the labelling to get a score that you could use to improve your model. The last day of the competition you had to make a choice and submit a final model, that would be evaluated on the test set, for which you had no information. The problem was that the training data was not very much. There was more data in the validation set, but the data in the validation set wasn't labelled. So I tried to label the validation set with a model I trained on the training set. And, what would you know. My classifier scored 100% accuracy on the validation set. But when I submitted my trained model on the test set it did much worse, I think close to 60% or so. Empirically demonstrated then: you can't dogfood a classifier to a better version of itself. When you train a classifier on some data, the classifier learns the underlying distribution of the data, with some amount of error. If you then label new data with the trained classifier and retrain the classifier on its own labelling, you end up multiplying the error. Btw, that doesn't change with language models, large or small, and it doesn't make a difference whether the model has an unsupervised training step or not. As long as your model is, well, modelling, some unknown true distribution and incurring some error, reusing the trained model to generate new data will generate data with error. So I don't think what you say can work and I'm curious to understand why you think it will. What are you saying will happen, exactly, if you dogfood an LLM's generations, like you suggest, that will make it improve?