4 ms·
I'd disagree as the K, Q and V have distinct functions within the attention calculation. In particular when you're considering decode (next token calculation du
by gchadwick 1y ago
I'd disagree as the K, Q and V have distinct functions within the attention calculation. In particular when you're considering decode (next token calculation during inference which follows the initial prefill stage that processes the prompt). For decode you have a single Q vector (relating to the in progress token) and multiple K and V vectors (your context, i.e. all tokens that have already been computed).
> you can represent all of the transformer as wide single layer perceptron sequences
This isn't correct, again because of attention. The classic perceptron has static weights, they are not an input. The same mathematical function can be used to compute attention however there are no static weights. You've got your attention scores on one side and the V matrix on the other side.
Indeed I wonder if it's actually possible for a bunch of perceptrons to even 'discover' the attention mechanism given they inherently have static weights and they can't directly multiply two inputs (or directly multiply two internal activations). Given an MLP is a general function approximater I guess a sufficiently large number of them could get close enough?
- ActorNightly 1y ago>For decode you have a single Q vector (relating to the in progress token) and multiple K and V vectors (your context, i.e. all tokens that have already been computed). Sure, but the K/V matrices are pretty much arbitrary weights, and so is the Q vector since its derived from the multiplication of the input vector by a learned matrix. The thing im trying to convey is that the nomeclature of Key/Query/Value doesn't mean anything, so when people learn about transformers, they don't need to understand that those matricies correspond to some predefined structure that maps to the data in a specific way. You can have 2 identical model initialized with random values and trained on same datasets, and end up with different KQV matricies for the same input. >This isn't correct, again because of attention. The classic perceptron has static weights, they are not an input. K/Q/V are all derived by multiplying the input vector by static learned weights, then those are all multiplied together in the attention calculation. Basically just a whole bunch of dot products. You would just have flattened matricies with intermediate layers acting as accumulators. >Indeed I wonder if it's actually possible for a bunch of perceptrons to even 'discover' the attention mechanism It is. It won't be attention in a classical sense, it would just be extra connections on a wider dimension single layer stacks. The learning process would put the right values in the correct place.