3 ms·
I don't really see how it's different from, say, a large convolutional neural network learning progressively higher-order features as you progress through the l
by throwaway1851 4y ago
I don't really see how it's different from, say, a large convolutional neural network learning progressively higher-order features as you progress through the layers of the network. At the lowest layers it's learning simple edge filters, which get combined into shapes, which get combined into filters that activate on faces, which get combined in ways that can be recognized as "family portrait", and so on. Of course, transformers have some unique advantages in terms of having very large context windows, being very parallelizable, etc.
When ChatGPT or any generative model produces output from a prompt, it's sampling from the (frozen) statistical structure it has learned. It makes sense that as you increase model capacity and the volume of training data, you can capture statistical patterns that occur at a very high level of abstraction. So, instead of just predicting the next token based on token-to-token patterns, it can predict the next token based on something resembling concept-to-concept patterns. (This somewhat demystifies the poetry generation / style transfer stuff that it can do. If you ask for a breach of contract complaint written as a sonnet, it can sample from both the patterns it's learned from legal documents and the patterns it's learned from poetry.)
What I wonder is how far this track of scaling up can take us. At the end of the day, interacting with ChatGPT is not that different from "interacting" with y=3x+7 by plugging in a value for x. It's just a much, much larger function.
- krackers 4y agoThanks, this is a nice explanation! Also found a paper related to emergence in LLMs: https://arxiv.org/pdf/2206.07682.pdf https://arxiv.org/pdf/2206.07682.pdf