4 ms·
“it can detect patterns among an O(n) input size without requiring an O(n^2) size neural net” This might be misleading, the amount of computation for processin
by tehsauce 6y ago
“it can detect patterns among an O(n) input size without requiring an O(n^2) size neural net”
This might be misleading, the amount of computation for processing a sequence size N with a vanilla transformer is still N^2. There has been recent work however which has tried to make them scale better.
- thomasahle 6y agon^2 time yes, but not n^2 weights in the net.
- m3at 6y agoYou raise an important point. The proposed solutions are too many to enumerate, but if I had to pick just one currently I would go for "Rethinking Attention with Performers" [1]. The research into making transformer better for higher dimensional inputs is also moving fast and is worth following. [1] https://arxiv.org/abs/2009.14794 https://arxiv.org/abs/2009.14794