3 ms·
The theory behind models is that they learn the distribution used to “generate” the data. A good model has a good approximation of the distribution. But then an
by tempusalaria 3y ago
The theory behind models is that they learn the distribution used to “generate” the data. A good model has a good approximation of the distribution. But then anything generated from it follows the approximation of the distribution not the distribution itself I.e. the generated data is not going to be good training data for the generating model.
You can for example train a very large and good model and use that to generate more data to train a smaller model e.g. for a faster inference use case. That data follows a closer approximation of the underlying distribution than the small model has so it can still be used for convergence to the underlying distribution.
- cubefox 3y agoAlphaGo Zero was trained on synthetic data which it created itself, it learned from self-play. So why wouldn't synthetic data work for GPTs? I guess a difference is that AlphaGo Zero uses reinforcement learning with objective reward signals (such as winning the game), while GPTs perform unsupervised learning (imitating the training text) without such a signal. There is no signal which tells the model whether the text it generated was or wasn't close to the real distribution. For some types of language model this signal could be generated though. If we have a language model which is able to generate pairs of the form (mathematical conjecture, proof attempt) in a formal language, then an automatic proof checker could generate the reward signal, i.e. proof correct / incorrect. The difficulty is probably to get the process bootstrapped since you need a certain amount of base proving capability to get the ball rolling.
- tempusalaria 3y agoYou’re confusing two different concepts. AlphaGo learns a distribution where given a game state it generates a move that maximises its internal probability of victory. There is a second “distribution” namely that any terminated sequence of go moves has an objective result. AlphaGo samples from the latter distribution in a guided way (as the space of all go games is computationally intractable). It uses its learned distribution to do that guided sampling and uses the objective outcomes of the known distribution to inform its own learned distribution. One way to think about this in the context of language modelling. Suppose I want to build a language model that says the word “goal” at least once every 2000 tokens generated. I could then repeatedly generate from the model and objectively score whether it has generated that word or not in each occurrence (the analogy of the finished go game). I then can use this objective scoring function to compete models against each other and do the alpha go style training. You can see here how the new training data is sampled from a different distribution than just regular language.
- cubefox 3y agoYour "goal" example sounds like a more useless version of my theorem prover AI?