3 ms·
Generally speaking, a phase transition occurs when the symmetries of a system are forced to change, either becoming restricted or liberated, as the result of so
by boerseth 3y ago
Generally speaking, a phase transition occurs when the symmetries of a system are forced to change, either becoming restricted or liberated, as the result of some parameter changing, perhaps past some critical value. The parameter can be a number of things, like density (pushing more particles into a box until the gas becomes more of a liquid) or temperature (like removing the total kinetic energy of the particles in a liquid until they can no longer move around eachother).
In the examples of the article, however, the phase transitions occur as the result of duration of training. Time is not a thermodynamic state variable, so these are not phase transitions in the strict sense of the word. However, I think it would be interesting to see the same experiment with something like stochastic gradient descent or the Metropolis algorithm, using some control parameter T to gradually reduce the randomness. This might let us find critical values of T for specific problems, which is useful because that is the value at which you want to spend the most time.
It may also be that for more involved problems, grokking is not achievable by naive gradient descent, as the landscape may be fraught with local minima. Perhaps, in order to find the way to the global minimum, to anneal the system into the lowest free-energy state, it might even be necessary to use a temperature-like control parameter, and mimic nature's own optimisation algorithm.
- boerseth 3y agoThinking about it more thoroughly, I am pretty convinced that the author's conclusion is correct. At the start, the neural net does not encode anything like an FFT or clever modulo arithmetic. The phase space of similarly-behaved weight-values is quite huge, when the neural net is not very sophisticated. But then eventually gradient descent finds its way into an implementation of FFT and all the rest, which solves the problem exactly. The phase space of weight-values that exhibit this behaviour is far, far smaller than the phase space I described in the previous paragraph. This means that the set of symmetries has been changed, which is the definition of a phase transition. The reason we see a first-order (sudden) phase transition is perhaps because of the size of the neural net being trained. Perhaps there aren't enough parameters available to encode a half-assed solution, and so there is no continuous transition from clueless to grokked. It either "gets it" or it doesn't.