Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
alexlitz
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Emulating ALiBi with Rope
(alexlitzenberger.com)
4 points
by
alexlitz
2mo ago
|
0 comments
2.
▲
by
alexlitz
7mo ago
The buried lede is this, if you have two dimensions and use rope, and hard-max attention you could simply store addresses as a given theta. With RoPE and sufficient precision that pretty easily gets you relative addressing with just one hea
3.
▲
by
alexlitz
7mo ago
> It baffles me that somebody capable of this kind of work would find this surprising. I should be clear I was not surprised that: 1) It struggled particularly hard with this sort of novel task 2) It tried to "correct" funky th
4.
▲
by
alexlitz
7mo ago
Yeah that is plausible enough.
5.
▲
by
alexlitz
8mo ago
Yeah basically it is an implementation detail but most of them are zero, there is an equivalent 14 parameter sparse matrix for that.
6.
▲
by
alexlitz
8mo ago
I imagine getting things to be polysemantic in a way that does not interfere would lead to sublinear scaling. Also there are smaller ones that were trained so would still be more like 311/36 ~= 8.6.
7.
▲
by
alexlitz
8mo ago
For one the specific 36 parameter version is impossible without float64 so you might guess the corollary that it is not exactly amenable to being found by gradient descent. I think the question of how you can structure transformers and neur
8.
▲
Building a Minimal Transformer for 10-digit Addition
(alexlitzenberger.com)
2 points
by
alexlitz
8mo ago
|
0 comments
9.
▲
by
alexlitz
8mo ago
I made a blogpost on my submission (currently the top handwritten one at 36 parameters) https://alexlitzenberger.com/blog/building_a_minimal_transfo...
10.
▲
Monte Carlo Geometry Processing
(youtube.com)
2 points
by
alexlitz
4y ago
|
0 comments