3 ms·
Hmm, just my intuition: training this model was very sensitive to the initial seed and training hyperparameters. It struggles to actually get to the 3x3 conv so
by montebicyclelo 1y ago
Hmm, just my intuition: training this model was very sensitive to the initial seed and training hyperparameters. It struggles to actually get to the 3x3 conv solution; but once it gets close to that things move much more quickly. This can kind of be seen in the animation of the attention matrix over time, which starts off random / spread out, but then once it starts to get more parts of the attention matrix in place it moves quicker. (Assuming all the experimentation wasn't in some bad part of the hyperparameter space.)
Also, it may just be the nature of the task. Some tasks you might have more to learn, all the time, with each training sample potentially giving information that's different from all the others. But with this, once it gets close to the solution of Life, it's quick.