3 ms·
Thinking about it more thoroughly, I am pretty convinced that the author's conclusion is correct. At the start, the neural net does not encode anything like an
by boerseth 3y ago
Thinking about it more thoroughly, I am pretty convinced that the author's conclusion is correct. At the start, the neural net does not encode anything like an FFT or clever modulo arithmetic. The phase space of similarly-behaved weight-values is quite huge, when the neural net is not very sophisticated.
But then eventually gradient descent finds its way into an implementation of FFT and all the rest, which solves the problem exactly. The phase space of weight-values that exhibit this behaviour is far, far smaller than the phase space I described in the previous paragraph. This means that the set of symmetries has been changed, which is the definition of a phase transition.
The reason we see a first-order (sudden) phase transition is perhaps because of the size of the neural net being trained. Perhaps there aren't enough parameters available to encode a half-assed solution, and so there is no continuous transition from clueless to grokked. It either "gets it" or it doesn't.