3 ms·
Part of the scaling benefit comes from the efficiency of training. LSTM and other recurrent architectures take more bits per parameter and the attention mechani
by robbedpeter 4y ago
Part of the scaling benefit comes from the efficiency of training. LSTM and other recurrent architectures take more bits per parameter and the attention mechanism doesn't optimize for context - each lstm neuron uses gating individually, where transformers attention mechanisms can use the context of an entire layer.
LSTMs with attention are super interesting, and there are some neat self-growing hopfield networks with attention that could push into transformer style efficacy and efficiency gains.