3 ms·
I don't understand the paper. It doesn't seem to be written very well. Are they replacing a low-rank adaptation with a banded matrix? And what is the point of t
by programjames 2y ago
I don't understand the paper. It doesn't seem to be written very well. Are they replacing a low-rank adaptation with a banded matrix? And what is the point of the rotations?
- ein0p 2y agoMy understanding is they reshape the input vector into groups and then pipe it through their approximator matrix M and reconstruct on the other end via concatenation. That, however, presents an issue in that M treats dimensions in different groups interchangeably, which they are not. In order to preserve the location information within the vector, they add rotations to each transformed input chunk like in RoPE. That just happens to work. I agree that the paper is not very good. In particular, their “for simplicity” comment actually only confuses things. It seems that you would have to pad the input or the output then, depending on whether you’re up- or down-projecting, and the case they are elaborating upon is the approximation of a square matrix. It’s also not 100% clear why M has to be square. Be that as it may, though, the idea has some merit IMO.