3 ms·
Is someone aware of transform architecture or implementation where every non linearity is made out of ReLu from linearities (max which can be expressed by ReLus
by freemint 4y ago
Is someone aware of transform architecture or implementation where every non linearity is made out of ReLu from linearities (max which can be expressed by ReLus instead of softmax).
In particular an Attention module which satisfies this property?
I did some googling however i only find papers like this one https://arxiv.org/abs/2204.07731 https://arxiv.org/abs/2204.07731 which deals with the O() behaviour with the window size. Not with the activation of the systems.