6 ms·
There's no free lunch when it comes to learning signal. The right comparison here is between training on synthetic data and training on the data used to create
by williamtrask 3y ago
There's no free lunch when it comes to learning signal. The right comparison here is between training on synthetic data and training on the data used to create the synthetic data generator.
Synthetic data doesn't create more learning signal, it just creates more data repeating the same learning signal with some bias assumptions that might as well be in your final model.
The difference is moreso in training times for retraining new models. If you've taken highly disparate data and compressed it into a particularly "juicy" dataset which is smaller, that could bring down training times.
That is to say, it's related to curriculum learning moreso than breaking information theory.
- bheadmaster 3y agoIf a model truly generalizes, why wouldn't it be possible for it to generate more real signal than was used for training it? Isn't that the whole point of using models instead of, I don't know, huge lookup tables?
- tempusalaria 3y agoThe theory behind models is that they learn the distribution used to “generate” the data. A good model has a good approximation of the distribution. But then anything generated from it follows the approximation of the distribution not the distribution itself I.e. the generated data is not going to be good training data for the generating model. You can for example train a very large and good model and use that to generate more data to train a smaller model e.g. for a faster inference use case. That data follows a closer approximation of the underlying distribution than the small model has so it can still be used for convergence to the underlying distribution.
- cubefox 3y agoAlphaGo Zero was trained on synthetic data which it created itself, it learned from self-play. So why wouldn't synthetic data work for GPTs? I guess a difference is that AlphaGo Zero uses reinforcement learning with objective reward signals (such as winning the game), while GPTs perform unsupervised learning (imitating the training text) without such a signal. There is no signal which tells the model whether the text it generated was or wasn't close to the real distribution. For some types of language model this signal could be generated though. If we have a language model which is able to generate pairs of the form (mathematical conjecture, proof attempt) in a formal language, then an automatic proof checker could generate the reward signal, i.e. proof correct / incorrect. The difficulty is probably to get the process bootstrapped since you need a certain amount of base proving capability to get the ball rolling.
- tempusalaria 3y agoYou’re confusing two different concepts. AlphaGo learns a distribution where given a game state it generates a move that maximises its internal probability of victory. There is a second “distribution” namely that any terminated sequence of go moves has an objective result. AlphaGo samples from the latter distribution in a guided way (as the space of all go games is computationally intractable). It uses its learned distribution to do that guided sampling and uses the objective outcomes of the known distribution to inform its own learned distribution. One way to think about this in the context of language modelling. Suppose I want to build a language model that says the word “goal” at least once every 2000 tokens generated. I could then repeatedly generate from the model and objectively score whether it has generated that word or not in each occurrence (the analogy of the finished go game). I then can use this objective scoring function to compete models against each other and do the alpha go style training. You can see here how the new training data is sampled from a different distribution than just regular language.
- cubefox 3y agoYour "goal" example sounds like a more useless version of my theorem prover AI?
- time_to_smile 3y agoThis comment reminds me of a time, many years ago, when someone wanted an the internal design team to make a big poster out of a low res image. The design team explained that they couldn't blow the image up to the size of a poster because it would look too grainy as it was extremely low resolution. The individual demanding the poster replied with a 'brilliant' solution: "Why don't you just take a picture of the low-res image, and then use that hi-resolution picture to blow up the image!" The same applies here. You need more information to learn the signal better, by definition there is no new information in the model.
- bheadmaster 3y ago> The same applies here. Does it, really? Analogies are great for explaining ideas, but aren't that great as logical arguments. > You need more information to learn the signal better You need more information than there is in the training set - which a well-generalizing model is supposed to be able to generate. Yes, there may be some bias if model doesn't generalize well, but that's why I put it in the assumption...
- version_five 3y ago> Does it, really? Yes. It's information theory. You can add information, by making rules like the class doesn't change if the thing is a different color or size or position, or on a different background. But you can't automatically create new images from a distribution that has been learned and expect them to add information.
- flangola7 3y agoThen how do humans do it? Our knowledge relies heavily on sending the output of humans to the input of humans (verbal communication, books, self-thought, etc). At one point total human knowledge was little more than "some berries are bad" but through this recursive process arrived at the rich garden of information we have today.
- 3y ago
- YeGoblynQueenne 3y ago>> If a model truly generalizes, why wouldn't it be possible for it to generate more real signal than was used for training it? If by "a model [that] truly generalizes" you mean one that generalises outside its training distribution, then neural nets can't train such models. I don't reckon any machine learning approach can do that. In any case this is ImageNet we're talking about that's been done to death many times over already. Far from learning more general models, after a certain point in time improvements in accuracy have meant that models are getting better at overfitting. Same with MNIST, where that happened a long time ago. Obligatory reference to defend against this comment being knee-jerked to oblivion: >> This stands in sharp contrast with what deep nets do, which I would call "local generalization": the mapping from inputs to outputs performed by deep nets quickly stops making sense if new inputs differ even slightly from what they saw at training time. Francois Chollet, The limitations of deep learning https://blog.keras.io/the-limitations-of-deep-learning.html https://blog.keras.io/the-limitations-of-deep-learning.html
- moyix 3y agoI genuinely don't see how this viewpoint is compatible with the grokking results.
- YeGoblynQueenne 3y agoI'm sorry, I'm not familiar with the grokking results. What are you referring to?
- MrScruff 3y agoFrom the linked blog post (2017) > Even with this data, you could not train a deep learning model to simply read a product description and generate the appropriate codebase. Would be make the same claim now I wonder?
- YeGoblynQueenne 3y agoI wonder too. I've communicated with Chollet in the past (about his ARC dataset) and he's responsive. You could email him and ask, or you could ask him on his twitter (I don't use twitter).
- technocratius 3y agoTrue, with the connotation that architectur and preprocessing pipeline are kept equal. E.g. if you're training convnets, data augmentation like rotating and flipping increases generalization.
- naveen99 3y agochecking correctness of things like factorization is easier than factoring in the first place.