3 ms·
First question: why should the attention mechanism output and residual stream match?
by macrolocal 4y ago
First question: why should the attention mechanism output and residual stream match?
- adamnemecek 4y agoMatch is a bad word, the don’t match, they are duals. The residual stream aka identity mapping needs to be the identity of the attention mechanism as the attention mechanism learns. But this is the same for all residual streams, not just those in transformers. Join my discord to discuss this further https://discord.gg/mr9TAhpyBW https://discord.gg/mr9TAhpyBW
- macrolocal 4y agoWait-- the residual stream makes the attention mechanism learn the difference from the identity! Are you sure you're not thinking about auto-encoders? Edit: ok, Discord it is.
- adamnemecek 4y agoDo you see a similarity between residual stream and Dirac function?
- crosen99 4y agoI don’t believe autodiff is finding the difference in that sense. It’s finding derivatives.
- macrolocal 4y agoWell, the paper uses gradient descent to minimize that difference, like auto-encoders do.
- crosen99 4y agoGradient descent is just how neural networks (including auto-encoders) optimize parameters to minimize the loss function. They do this using derivatives to descend down the slope of the function. Autodiff is one way to compute the derivatives. Maybe we’re saying the same thing.
- macrolocal 4y agoYep, I was just asking Adam* to justify his loss function. *pun intended :)