4 ms·
I’m not deeply familiar with all these papers, but two things stand out to me The model architectures are different, and in the very latest paper they scale th
by rsfern 4y ago
I’m not deeply familiar with all these papers, but two things stand out to me
The model architectures are different, and in the very latest paper they scale these not-transformer models to sequence length of 64k, where the paper you linked only considers up to 8k