3 ms·
Yes this seems like an early work in progress, compared to Jay's previous Transformer articles. In addition to your link, I've found a really good Transformer
by ypcx 6y ago
Yes this seems like an early work in progress, compared to Jay's previous Transformer articles.
In addition to your link, I've found a really good Transformer explanation here (backed by a Github repo w/ lively Issues talk): http://www.peterbloem.nl/blog/transformers http://www.peterbloem.nl/blog/transformers
Additionally, there's a paper on visualizing self-attention: https://arxiv.org/pdf/1904.02679.pdf https://arxiv.org/pdf/1904.02679.pdf
- m3at 6y agoThat's a good complement, thank you for the links
- deleted 6y ago[deleted]
- ypcx 6y agoCan't edit the post anymore so adding it here - further research reading on improving the current attention model: https://www.reddit.com/r/MachineLearning/comments/hxvts0/d_breaking_the_quadratic_attention_bottleneck_in/ https://www.reddit.com/r/MachineLearning/comments/hxvts0/d_b...