3 ms·
Still chewing but my takeaways thus far: SSM models strengths are with continuous data like audio and video, they struggle with discrete data like text/ DNA. T
by bananaheel 3y ago
Still chewing but my takeaways thus far:
SSM models strengths are with continuous data like audio and video, they struggle with discrete data like text/ DNA. This newest architecture uses selective attention to try to address the weaknesses around discrete data with some loss in performance in continuous tasks, empirically shown here with audio. The empirical exploration was limited to smaller size models, the performance as larger scales is yet to be explored in practice.
I found my deepest understanding of the selection mechanism came from struggling with the discretization in section 2, followed by the deeper explanations of the variables involved in 3.5.2. This video gives excellent background to SSMs, along with a detailed walk through of the paper itself.[a]
I am still coming to understand S4, SSMs in general but the video suggested this annotated explainer that has been helping a lot[b].
I would also point out section 3.1 and it’s discussion of the tradeoffs between compression and effectiveness as particularly interesting.
I do wonder how many different GPUs / hardware architectures will be able to execute the optimizations that are described as critical. I think the nature of the optimizations is the part of the paper I understand least well.
The paper taken at face value looks very exciting. The promise of a very large context window with 5x throughput for inference would be huge if it proves to scale well. I do wonder if it will make sense to train SSMs without this selection mechanism for specific continuous use cases where it seems to perform better or if other architectures will prove to better serve those cases.
----
[a] https://www.youtube.com/watch?v=ouF-H35atOY https://www.youtube.com/watch?v=ouF-H35atOY
[b] https://srush.github.io/annotated-s4/ https://srush.github.io/annotated-s4/
- Nimitz14 3y agoYeah I also was confused by the discretization step. The video is great thank you for sharing!
- mcint 3y agoThe distinction I remember from the paper, is that discrete data (text, DNA) could be modeled effectively with only real-valued components, while continuous data (audio, video) benefited from complex numbers -- in their state. Evidently summarized from the authors statements characterizing existing work, preexisting wisdom on SSM/S4 models. https://youtubetranscript.com/?v=ouF-H35atOY https://youtubetranscript.com/?v=ouF-H35atOY (same video, I'd watched previous to seeing this hn post) other model details the authors note that most prior State space models use complex numbers in their state but it has been empirically observed that completely real valued State space models seem to work fine and possibly even better in some settings so they use real values as the default which work well for all but one of their tasks next just following this, another impressive snippet it succeeds on test sequence lengths of up to a million tokens which is 4,000 times longer than it saw during training while none of the other methods compared to generalize to Beyond twice their training length