3 ms·
Transformer networks have deeper connections to dense associative memory. For example, the update rule to minimize the energy functional of these Hopfield netw
by macrolocal 4y ago
Transformer networks have deeper connections to dense associative memory. For example, the update rule to minimize the energy functional of these Hopfield networks converges in a single iteration and coincides with the attention mechanism [1].
[1] https://arxiv.org/abs/1702.01929 https://arxiv.org/abs/1702.01929
- macrolocal 4y agoMore accessible references: https://mcbal.github.io/post/an-energy-based-perspective-on-attention-mechanisms-in-transformers/ https://mcbal.github.io/post/an-energy-based-perspective-on-... (Modern continuous Hopfield networks section) https://arxiv.org/abs/2008.02217 https://arxiv.org/abs/2008.02217 Note that the connection to Hebbian learning hinges on the softmax function, in particular its exponential!