3 ms·
It’s weird a DeepMind author is on this as they experimentally studies these methods in Gopher and found it was useless. Even the layer duplication trick they
by tempusalaria 3y ago
It’s weird a DeepMind author is on this as they experimentally studies these methods in Gopher and found it was useless.
Even the layer duplication trick they were using doesn’t work. I tried a refinement of that where you fine tuned the activations of the added layers to match the original activations which helped a bit relative to what DeepMind found but it still wasn’t material enough to be worth bothering with
- two_in_one 3y agoI tried adding layer. The conclusion is, it's useful if you cannot train the whole model at once. Then you can split it in 2 and train the first half, with embedding/deembedding layers. After it stops improving add the second half of the layers, and train only it, freezing everything else. This worked in my case, when model params number < 10% of the number of training tokens. Splitting model in 3 is significantly less effective. Don't know why, but the third part adds almost nothing in accuracy and a lot in training time. One thing where additive training works is fine tuning, like LoRA. But that's different. It's like poor man's way. For those who cannot train the whole model. Most likely uptrainig it whole is more efficient, just my guess.