Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fblgit
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Single-layer transformer model "HarEmb" showcasing PII SOTA performance
(huggingface.co)
1 points
by
fblgit
5mo ago
|
1 comments
2.
▲
by
fblgit
5mo ago
one of a kind single-transformer block layer, high throughput. The new generation of transformer-based lightweight models for common NLP tasks?
3.
▲
by
fblgit
3y ago
doesn't require much data, in a 7B can take a couple hours ~
4.
▲
by
fblgit
3y ago
Correct. UNA can align the MoE at multiple layers, experts, nearly any part of the neural network I would say. Xaberius 34B v1 "BETA".. is the king, and its just that.. the beta. I'll be focusing on the Mixtral, its a christm
5.
▲
by
fblgit
3y ago
UNA: Uniform Neural Alignment. Haven't u noticed yet? Each model that I uniform, behaves like a pre-trained.. and you likely can fine-tune it again without damaging it. If you chatted with them, you know .. that strange sensation, you