Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pfekin_2nd
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
pfekin_2nd
1y ago
This is brilliant, thanks.
2.
▲
by
pfekin_2nd
1y ago
Author here — a few clarifications up front: How is this different from Performer / linear attention? Performer and related methods approximate the softmax kernel with random features or low-rank projections. Summation is not an approx
3.
▲
Summation-based aggregation: a simpler alternative to self-attention
(techrxiv.org)
2 points
by
pfekin_2nd
1y ago
|
2 comments
4.
▲
by
pfekin_2nd
1y ago
Summation-based aggregation replaces pairwise similarity with position-modulated projections and direct summation, reducing per-layer cost from quadratic to near-linear. On its own, summation is competitive for classification and multimodal