5 ms·
Anything that can reduce the number of parameters required for a transformer is good news. Wonder if this approach can be easily plugged into the recent retriev
by rdedev 5y ago
Anything that can reduce the number of parameters required for a transformer is good news. Wonder if this approach can be easily plugged into the recent retrieval transformer.
- lucidrains 5y agoyea definitely, the ideas are orthogonal
- albertzeyer 5y agoWhy are the number of parameters relevant? The training and decoding runtime and memory consumption are what is relevant. And the number of parameters is not really connected to that. E.g. I assume this DeepNet with 3.2B params is slower and requires more memory than the M2M-100 with 12B params. The number of layers are more connected to both runtime and memory consumption.
- l33tman 5y agoFirstly the number of parameters is important in itself due to the finite size of VRAM on the GPU cards commonly available. But secondly the bulk of the ops you do are scalar products (arranged as matrix muls) where one parameter means one multiply+accumulate. So regardless of the actual net topology, number of layers etc you should see a rough scaling in the number of parameters, in an ideal case with no other overhead (you can write code that scales badly or weirdly). I didn't look at this DeepNet, maybe it will run slower like you say due to some other choice of architecture.. Also yes an architecture that requires 1000x more training iterations will of course be worse at the training stage even if it has 100x fewer parameters, but maybe the resulting inference model is much faster. Lots of stuff to consider :)
- albertzeyer 5y agoIn a Transformer, the most expensive op is the self attention. So when you have more layers, this will be more expensive. You could also share e.g. all the parameters in every layer (like Universal Transformer), and thus reduce the number of params drastically, but without any change in computing time, and only negligible difference in memory consumption because the hidden activations take most of the memory, not the params.
- l33tman 5y agoAh ok I see what you mean, I didn't consider shared parameters (out of habit, in my own AI work I have no shared parameters). But isn't the self attention also trainable and thus is reflected in the number of parameters (with or without sharing)? I naively see the transformer architecture as a generalization of other common topologies like convnets, i.e. you can arrange the token stream and the attention heads (if you have enough of them) to emulate a 2D convnet if you want even though the first uses where for 1D language token streams, and so the transformer arch with proper training can converge to a convnet arch if that is the best solution. Or is this a completely backward generalization? :)
- rdedev 5y agoHadn't come across parameter sharing in transformers. That's something I'll definitely look into. Why is it not wide spread though ? Most famous and powerful models have a huge number of parameters