2 ms·
Since gradient descent converges on a local minima, would we expect different emergent properties with different initialization of the weights?
by waynecochran 1y ago
Since gradient descent converges on a local minima, would we expect different emergent properties with different initialization of the weights?
- jebarker 1y agoNot significantly, as I understand it. There's certainly variation in LLM abilities with different initializations but the volume and content of the data is a far bigger determinant of what an LLM will learn.
- waynecochran 1y agoSo there is an "attractor" that different initializations end up converging on?
- andy99 1y agoDifferent initialization converge to different places, e.g https://arxiv.org/abs/1912.02757 https://arxiv.org/abs/1912.02757 For LLMs (as with other models), many local optima appear to support roughly the same behavior. This is the idea of the problem being under-specified ie many more equations than unknowns so there are many ways to get the same result.
- recursivecaveat 1y agoYou end up with different weights when using different random initialization, but with modern techniques the behavior of the resulting model is not really distinct. Back in the image-recognition days it was like +/- 0.5% accuracy. If you imagine you're descending in a billion-parameter space, you will always have a negative gradient to follow in some dimension: local minima frequency goes down rapidly with (independent) dimension count.