3 ms·
Data is an excellent place to look at to get a sense of where the model is likely to work or not (what kinds of images), and for prompt design ideas because, ro
by karpathy 4y ago
Data is an excellent place to look at to get a sense of where the model is likely to work or not (what kinds of images), and for prompt design ideas because, roughly speaking, the probability of something working well is proportional to its frequency (or of things very similar to it) in the data.
The story is more complex though because the data can often be quite far away from actual neural net training due to preprocessing steps, data augmentations, oversampling settings (it's not uncommon to not sample data uniformly during training), etc. So my favorite place to scrutinize is to build a "batch explorer": During training of the network one dumps batches into pickles immediately before the forward pass of the neural net, then writes a separate explorer that loads the pickles and visualizes them to "see exactly what the neural net sees" during training. Ideally one then spends some quality time (~hours) looking through batches to get a qualitiative sense of what is likely to work or not work and how well. Of course this is also very useful for debugging, as many bugs can be present in the data preprocessing pipeline. But a batch explorer is harder to obtain here because you'd need the full training data/code/settings.
- jimsimmons 4y agoWith big generative models, seeing data even once is more than sufficient to memorize it. So your claim that performance relates to frequency is not exactly correct. The whole point of this model class is that one can learn one word from one sample, another pixel from another one and so on to master the domain. The emergent, non-trivial generalization is what makes them so fascinating. There is no simple, linear/first order relationship with data and behaviour. Case in point: GPT3 can do few-shot learning despite not having used any explicit few-shot formatted data during training. Not saying you are wrong, but the story is not as simple as simple supervised learning with small datasets
- galangalalgol 4y agoWhat does that say about how these models will behave as an increasingly large portion of their training data is outputs from similar models? Our curation of the outputs will hopefully help. And if one image really is enough, perhaps the smaller number of human created images will be sufficient to inject new stuff rather than stagnating?