4 ms·
Ok, how do transformers fit into this understanding of deep learning?
by polotics 1y ago
Ok, how do transformers fit into this understanding of deep learning?
- motoboi 1y agoTransformers don't feel differentiable (because of the attention mechanism), but they actually are (as being back-propagation based forces it to be). The attention mechanism is not a stretching of the manifold, but is trained to be able to measure distances in the manifold surface, which is stretched and deformed (or transformed?) in the feed-forward layers.
- theahura 1y agoTransformers learn embedding representations of tokens, which are easily mapped into a space. Similar tokens are mapped to similar places on the space. The fully connected layer at the end of each transformer block defines a transformation of a set of points in a space to another point in that space, not unlike the example of adding colors together to get a new color
- jebarker 1y agoTransformers (with self-attention being the key operation) are kernel smoothers which fits easily into this view of the world. See here: http://bactra.org/notebooks/nn-attention-and-transformers.html http://bactra.org/notebooks/nn-attention-and-transformers.ht...