Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
thw20
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Simple, zero overhead way to compress model, KV cache via Low-Rank Decomposition
(jeffreywong20.github.io)
1 points
by
thw20
5mo ago
|
0 comments
2.
▲
by
thw20
5mo ago
Good work! This is very interesting. Here's a related work that construct low-rank approximation for attention: https://arxiv.org/abs/2505.12942 . Maybe the idea of Query calibration matrix Rxx is of interest to th
3.
▲
by
thw20
7mo ago
The up to date paper documenting and analysing the observation is now available on arxiv!
4.
▲
by
thw20
7mo ago
This project reveals an interesting phenomena, where LLM converts semantic non-informative tokens to attention sinks through middle layer MLP. The converted sinks are termed secondary attention sinks as they are weaker then BOS attention si
5.
▲
Towards understanding multiple attention sinks in LLMs
(github.com)
1 points
by
thw20
7mo ago
|
2 comments
6.
▲
by
thw20
7mo ago
This is so amazing. What a masterpiece for intro to reinforcement learning in llm.
7.
▲
The Existence and Behavior of Secondary Attention Sinks
(arxiv.org)
1 points
by
thw20
7mo ago
|
0 comments