4 ms·
Transformers are about converting some input data (usually text) to numeric representations, then modifying those representations through several layers to gene
by vikp 3y ago
Transformers are about converting some input data (usually text) to numeric representations, then modifying those representations through several layers to generate a target representation.
In LLMs, this means go from prompt to answer. I'll cover inference only, not training.
I can't quite ELI5, but process is roughly:
- Write a prompt
- Convert each token in the prompt (roughly a word) into numbers. So "the" might map to the number 45.
- Get a vector representation of each word - go from 45 to [.1, -1, -2, ...]. These vector representations are how a transformer understands words.
- Combine vectors into a matrix, so the transformer can "see" the whole prompt at once.
- Repeat the following several times (once for each layer):
- Multiply the vectors by the other vectors. This is attention - it's the magic of transformers, that enables combining information from multiple tokens together. This generates a new matrix.
- Feed the matrix into a linear regression. Basically multiply each number in each vector by another number, then add them all together. This will generate a new matrix, but with "projected" values.
- Apply a nonlinear transformation like relu. This helps model more complex functions (like text input -> output!)
Note that I really oversimplified the last few steps, and the ordering.
At the end, you'll have a matrix. You then convert this back into numbers, then into text.
- throwawaymaths 3y agoI don't think this description of attention is correct.
- vikp 3y agoYou mean "Multiply the vectors by the other vectors. This is attention - it's the magic of transformers, that enables combining information from multiple tokens together. This generates a new matrix."? It's really oversimplified, as I mentioned. A more granular look is: - Project the vectors with a linear regression. In decoder-only attention (what we usually use), we project the same vectors twice with different coefficients. We call the first projection queries, and the second keys. This transforms the vectors linearly. - Find the dot product of each query vector against the key vectors (multiply them) - (training only) Mask out future vectors, so a token can't look at tokens that come after it - At this point, you will have a matrix indicating how important each query vector considers each other vector (how important each token considers the other tokens) - Take the softmax, which both ensures all of the attention values for a vector sum to 1, and penalizes small attention values - Use the softmax values to get a weighted sum of tokens according to the attention calc. - This will turn one vector into the weighted sum of the other vectors it considers important. The goal of this is to incorporate information from multiple tokens into a single representation.