4 ms·
Architecture thread! Afaict they continue to use gated attention + delta net, which was also adopted+adapted by K3, but im surprised theres no improvements to t
by minimaltom 2mo ago
Architecture thread! Afaict they continue to use gated attention + delta net, which was also adopted+adapted by K3, but im surprised theres no improvements to the residual stream (deepseek are using manifold hyper-connections, kimi have attention residuals) ?
Perf improvements seem to all come from training?
- anana_ 2mo agoAs was the case with GLM 5.3, it seems that there is still much juice to be squeezed from post-training