4 ms·
I never did as much thinking or testing of dropout on transformers as the author, but it didn't seem to help with my "baby" (~10 million param) transformer mode
by Scene_Cast2 2y ago
I never did as much thinking or testing of dropout on transformers as the author, but it didn't seem to help with my "baby" (~10 million param) transformer models. IIRC the latest Llama models don't use dropout either.
- mattnewton 2y agoSame, I was never able to debug why dropout > 5% really hurt convergence speed for my toy LLMs. I chalked it up to the models not having enough parameters to fit fineweb and just stop using it.
- gpjt 2y agoMy intuition is very undeveloped on this, but it makes some kind of sense to me that dropout would make convergence slower, because you're ignoring a bunch of parameters in every batch. The goal seems to be to get a better, more general model by trading off some training time. The Llama thing is interesting, though!