4 ms·
Transformers do have a fixed input/output size though - that's what a context window is. It's just that, via scaling and algorithmic improvements, the length of
by ainch 6mo ago
Transformers do have a fixed input/output size though - that's what a context window is. It's just that, via scaling and algorithmic improvements, the length of usable context windows has increased to the point that they're much less of a bottleneck.
I think your points around parallelisation and the flexibility of quadratic attention are spot-on though.