4 ms·
Hold on, every predicted token is only a function of the previous token? I must have something wrong. This would mean that within the embedding of "was", which
by rollinDyno 2y ago
Hold on, every predicted token is only a function of the previous token? I must have something wrong. This would mean that within the embedding of "was", which is of length 12,228 in this example. Is it really possible that this space is so rich as to have a single point in it encapsulate a whole novel?
- evolvingstuff 2y agoYou are correct, that is an error in an otherwise great video. The k+1 token is not merely a function of the kth vector, but rather all prior vectors (combined using attention). There is nothing "special" about the kth vector.
- causal 2y agoI read this comment yesterday and keep thinking about it. That final token really must "comprehend" everything leading up to it, right? In which case longer context lengths are just trying to pack more meaning into that embedding state. Which means the embedding model must do a lot of the lifting to be able to accurately represent meaning across long contexts so well. Now I want to know more about how those models are derived.
- diedyesterday 2y ago> "Is it really possible that this space is so rich as to have a single point in it encapsulate a whole novel?" Not with this GPT. The context size would not allow keeping attention to the total meaning of more than 2048 tokens (as reflected in the transformed embedding of that context's last token). For a substantial part of a novel, it would require a much larger context size with then presumably will need a higher dimensional embedding/semantic space.
- vanjajaja1 2y agoat that point what it has is not a representation of the input, its a representation of what the next output could be. ie. its a lossy process and you can't extract what came in the past, only the details relevant to next word prediction (is my understanding)
- rollinDyno 2y agoIf the point was the presentation of only the next token, and predicted tokens were a function of only the preceding token, then the vector of the new token wouldn’t have the information to produce new tokens that kept telling the novel.
- faramarz 2y agoit's not about a single point encapsulating a novel, but how sequences of such embeddings can represent complex ideas when processed by the model's layers. each prediction is based on a weighted context of all previous tokens, not just the immediately preceding one.
- rollinDyno 2y agoThat weighted context is the 12228 dimensional vector, no? I suppose that when you each element in the vector weighs 16 bits then the space is immense and capable to have a novel in a point.
- jgehring 2y agoThat's what happens in the very last layer. But at that point the embedding for "was" got enriched multiple times, i.e., in each attention pass, with information from the whole context (which is the whole novel here). So for the example, it would contain the information to predict, let's say, the first token of the first name of the murderer. Expanding on that, you could imagine that the intent of the sentence to complete (figuring out the murderer) would have to be captured in the first attention passes so that other layers would then be able to integrate more and more context in order to extract that information from the whole context. Also, it means that the forward passes for previous tokens need to have extracted enough salient high-level information already since you don't re-compute all attention passes for all tokens for each next token to predict.
- causal 2y ago> you don't re-compute all attention passes for all tokens for each next token to predict. You don't? I imagine the attention maps could be pretty different between n and n+1 tokens. Edit: Or maybe you just meant you don't compute attention Σ(n) times for each new token?