3 ms·
This was discussed in Gopher paper. The added zero weights don’t integrate well into LLMs during training unfortunately. They actually found that if you duplic
by tempusalaria 3y ago
This was discussed in Gopher paper. The added zero weights don’t integrate well into LLMs during training unfortunately.
They actually found that if you duplicated layers when adding it worked better than zero weights. Which matches some of the commutating layer studies that have been done
- drdeca 3y agoI notice that in this paper, some of the new weights can be initialized arbitrarily, and only some of them have to be zero. In the Gopher paper, were all the new weights zero, or only the ones that this paper calls for being zero? I would guess the latter? (I don't expect you to either have or fetch an answer to this question. If you happen to already have the answer, or simply a more educated guess than I have, I'd like to hear it, but I mean this more as "this question comes to mind" than anything where I expect an answer.)