3 ms·
from the abstract By incorporating DyT, Transformers without normalization can match or exceed the performance of their normalized counterparts, mostly witho
by gricardo99 2y ago
from the abstract
By incorporating DyT, Transformers without normalization can match or exceed the performance of their normalized counterparts, mostly without hyperparameter tuning.
- gdiamos 2y agoSure, but why would one prefer tanh instead of normalization layers if they have the same accuracy? I suppose normalization kernels have reductions in them, but how hard are reductions in 2025?