4 ms·
The self-attention mechanism is explained very well in this blog post. Because of this it is very much worth a read for anybody interested in the state of the a
by rerx 8y ago
The self-attention mechanism is explained very well in this blog post. Because of this it is very much worth a read for anybody interested in the state of the art of deep learning models for machine translation. Other parts of the Transformer model are glossed over more, though.