4 ms·
The distinction I remember from the paper, is that discrete data (text, DNA) could be modeled effectively with only real-valued components, while continuous dat
by mcint 3y ago
The distinction I remember from the paper, is that discrete data (text, DNA) could be modeled effectively with only real-valued components, while continuous data (audio, video) benefited from complex numbers -- in their state. Evidently summarized from the authors statements characterizing existing work, preexisting wisdom on SSM/S4 models.
https://youtubetranscript.com/?v=ouF-H35atOY https://youtubetranscript.com/?v=ouF-H35atOY (same video, I'd watched previous to seeing this hn post)
other model details the authors note
that most prior State space models use
complex numbers in their state but it
has been empirically observed that
completely real valued State space
models seem to work fine and possibly
even better in some settings so they use
real values as the default which work
well for all but one of their tasks next
just following this, another impressive snippet
it succeeds on test sequence lengths of
up to a million tokens which is 4,000
times longer than it saw during training
while none of the other methods compared
to generalize to Beyond twice their
training length