Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
szcs
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
szcs
2y ago
That's just the same distribution laid out along a line instead of a circle
2.
▲
by
szcs
2y ago
Attention is a 3 matrix product, s(QK)V where s is softmax. Each matrix has as many rows (Q and V) or columns (K) as many tokens you have in your context. The plot looks at the processing of a single row of Q (predicting a single token from
3.
▲
by
szcs
2y ago
There is a particularly nice geometric interpretation of attention I just realised recently in a flash of enlightenment, best explained with an interactive Desmos plot (black dot is draggable): https://www.desmos.com/calcula
4.
▲
by
szcs
2y ago
Yes, LLMs are planned too.
5.
▲
by
szcs
2y ago
Author here, I just noticed this. If you have any questions I can try answering them.
6.
▲
by
szcs
2y ago
Here is a quick graph I just generated from data I had. Horizontal axis is average bit depth, vertical is accuracy. PTQ is Post Training Quantisation, QAT is Quantisation Aware Training. It's a simple ResNet trained on CIFAR-10, I don&
7.
▲
by
szcs
2y ago
Author here. I actually came up with the idea a long time ago, I first experimented with variants of this in Caffe (before Tensorflow was a thing).