4 ms·
This explanation walks you through the math and the corresponding code, but (at least in my case, maybe I'm dumb) it failed to help me understand why these step
by hackyhacky 4y ago
This explanation walks you through the math and the corresponding code, but (at least in my case, maybe I'm dumb) it failed to help me understand why these steps are necessary or to relate the math to the intended outcome. As a result, I don't feel that I'm any closer to really understanding the heart of self-attention.
- GaggiX 4y agoThis article explains well why we need attention: https://jalammar.github.io/visualizing-neural-machine-translation-mechanics-of-seq2seq-models-with-attention/ https://jalammar.github.io/visualizing-neural-machine-transl..., and also how it was developed.
- zmmmmm 4y agothose animations are beautiful!
- whoateallthepy 4y agoAt the end of last year I put together a repository to try and show what is achieved by self-attention on a toy example: detect whether a sequence of characters contains both "a" and "b". The toy problem is useful because the model dimensionality is low enough to make visualization straightforward. The walkthrough also goes through how things can go wrong, and how it can be improved, etc. The walkthrough and code is all available here: https://github.com/rstebbing/workshop/tree/main/experiments/transformer_sequence_classification https://github.com/rstebbing/workshop/tree/main/experiments/.... It's not terse like nanoGPT or similar because the goal is a bit different. In particular, to gain more intuition about the intermediate attention computations, the intermediate tensors are named and persisted so they can be compared and visualized after the fact. Everything should be exactly reproducible locally too!
- doug_durham 4y agoI agree. It seems like the target audience is the experienced Deep learning practitioner. Which makes me wonder why such an audience would need this treatment. Why not just read the original paper?