4 ms·Gated Linear Attention Transformers with Hardware-Efficient Training2 points by dataminer 3y ago