3 ms·
I see similar loss curves when training ViTs (from scratch), which has always bothered me but I had bigger concerns so never delved too deep into it. The only
by fpgaminer 3y ago
I see similar loss curves when training ViTs (from scratch), which has always bothered me but I had bigger concerns so never delved too deep into it. The only difference is that I see the training loss go _up_ during each epoch. The cliffs between epochs are large enough that training loss goes down overall and validation loss keeps going down the whole time as well. The model gets close-ish to SoTA so I guess it's "normal".
I haven't trained convnets at this scale so I'm not sure if similar behavior has been seen there, but you'd think someone would have mentioned it at some point. So perhaps these strange loss curves are a feature of Transformer based models in particular?
- jph00 3y agoOh wow yeah - I've also seen other people's training loss curves like that, going up during each epoch and then jumping down at the end of the epoch. I've never experienced that myself, and have no idea what's causing it!
- lIIllIIllIIllII 3y agoThe original article mentioned LLMs needing powerful abstractions this is basically the case with transformer networks, which is apparent when learning from scratch. The model seems to be going basically nowhere and totally useless until suddenly, at some random point after a bunch of learning cycles the weights find some minimum on the error surface and bam, suddenly the model can do things properly. And it's because the transformer has learned an abstraction that works for all of the input data in an attentional sense (think how you scan a sentence when reading). Not the best explanation but its from memory from a post I saw on HN a while back
- whimsicalism 3y agoeven in the first epoch the loss goes up? that seems.. odd
- t-vi 3y agoAfter the first epoch, the average time since the present data item was last used for during training is small at the beginning of an epoch grows during the epoch. I'd expect that to positively relate to loss on the present iteration.