Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
haoxiaoru
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
haoxiaoru
11mo ago
I've waited so long— four months
2.
▲
Muon is scalable for LLM training
(github.com)
5 points
by
haoxiaoru
2y ago
|
1 comments
3.
▲
by
haoxiaoru
2y ago
Recently, the Muon optimizer based on matrix orthogonalization has demonstrated strong results in training small-scale language models, but the scalability to larger models has not been proven. We identify two crucial techniques for scaling