4 ms·
> more efficient than standard attention whenever the model dimension is greater than or equal to the context length all practical models have context length s
by novaRom 2y ago
> more efficient than standard attention whenever the model dimension is greater than or equal to the context length
all practical models have context length significantly larger than model dimension