4 ms·
Reminds me of "Universal pre-training by iterated random computation" https://arxiv.org/pdf/2506.20057 https://arxiv.org/pdf/2506.20057, with bit less formal ap
by benob 7mo ago
Reminds me of "Universal pre-training by iterated random computation" https://arxiv.org/pdf/2506.20057 https://arxiv.org/pdf/2506.20057, with bit less formal approach.
I wonder if there is a closed-form solution for those kinds of initialization methods (call them pre-training if you wish). A solution that would allow attention heads to detect a variety of diverse patterns, yet more structured than random init.
- ACCount37 7mo agoI'm partial to "pre-pre-training" myself.