4 ms·RWKV-LM: RNN with transformer-level performance, without using attention4 points by vletal 4y ago