3 ms·
restricts freedom in one of the parameters (A) to make training substantially more efficient (easier for a GPU to churn through). the actual flops involved are
by brrrrrm 2y ago
restricts freedom in one of the parameters (A) to make training substantially more efficient (easier for a GPU to churn through).
the actual flops involved are similar to the original SSM-based version, but that's harder to formulate as strictly matrix multiplications