Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
phowon
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
phowon
8y ago
Like I said - pooling. You can take the mean over 3 elements or 10 elements just the same. Pooling is lossy, but it seems that if you have the right architecture the model can still learn what it needs to. It's worth noting that the at
32.
▲
by
phowon
8y ago
>Even transformers use RNN structures right. Nope. >How do you handle variable length input without something like an RNN? Any form of pooling, really. Max, Avg, Sum. The tricky part is how to do the pooling while still taking advanta
33.
▲
by
phowon
8y ago
>Look inside a SAGAN or something and you'll see the conv2d calls. ...Yes, because SAGANs operate on images, so the foundational operation is a convolution. >You're reading that in an overly narrow way and imputing to me som
34.
▲
by
phowon
8y ago
A 1x1 convolution is such an edge case of convolution that it's really not worth discussing its inclusion as related to the success of the Transformer. Calling the Transformer "convolutions with attention" demonstrates a new-
35.
▲
by
phowon
8y ago
But no part of that Transformer section makes any reference to convolutions.
36.
▲
by
phowon
8y ago
Can you elaborate on how Transformers are "convolutions with attention"?