4 ms·
Oddly enough, I was reading their paper just last night: https://arxiv.org/pdf/2302.10866.pdf https://arxiv.org/pdf/2302.10866.pdf I think we're going to see a
by jmole 4y ago
Oddly enough, I was reading their paper just last night: https://arxiv.org/pdf/2302.10866.pdf https://arxiv.org/pdf/2302.10866.pdf
I think we're going to see a lot more in the wavelet/convolution/fft space when thinking about how to increase context length.
I think there's also a lot of room for innovation in the positional encoding and how it's represented in transformer models, it seems like people have been trying lots of things and going with what works, but most of it is like: "look, a new orthonormal basis!".
Hyena sort of seems like the first step in moving to positional embeddings (or joint positional/attentional embeddings).
Very cool work.
- cs702 4y agoI agree this sort of approach looks promising. Maybe using FFTs recurrently to approximate convolutions with input-length filters is the way forward. It's a clever idea. I'm making my way through the paper. Don't fully understand it yet. The main issue I've seen with other wannabe-sub-quadratic-replacements for self-attention is that they all rely on some kind of low-rank/sparse approximation that in practice renders LLMs incapable of modeling enough pairwise relationships between tokens to achieve state-of-the-art performance. I'm curious to see if this kind of approach solves the issue.
- inciampati 4y agoDoes FFT not suggest an implicit kind of quantization to the dimension of the time space model? It's something that I have not understood about S4, Hippo, Hyena and other FFT-trick models.
- cs702 4y agoFFT's inputs are discrete to begin with (evenly spaced samples in the time domain), so I kind of see what you're trying to say... but I don't fully understand the work yet. I'm slowly making my way, starting with getting acquainted with classic state-space models. That said, my sense is there's a significant difference between (a) learning to approximate the matrix of n×n interactions with more run-of-the-mill methods, such as some kind of linear decomposition with learned fixed coefficients, versus (b) learning to approximate the matrix of n×n interactions with dynamic methods that "find the most suitable decomposition for each sample on-the-fly," which is what these class of models appears to be doing. Apologies if all this sounds very hand-wavy; it's the best I can do at the moment.