4 ms·
A 7k param LSTM is very tiny. Not sure if LSTMs would even work at that scale although someone with more theoretical knowledge can correct me on this. As an a
by rakejake 2y ago
A 7k param LSTM is very tiny. Not sure if LSTMs would even work at that scale although someone with more theoretical knowledge can correct me on this.
As an aside, I'm trying to train transformers for some classification tasks on audio data. The models are "small" (like 1M-15M params at most) and I find they are very finicky to train. Below 1M parameters I find them hard to train at all. I have thrown all sorts of learning rate schedules at them and the best I can get is the network learns for a bit and then plateaus, after which I can't do anything to get them out of that minima. Training an LSTM/GRU on the same data gives me a much better loss value.
I couldn't find many papers on training transformers at that scale. The only one I was able to find was MS's TinyStories [0], but that paper didn't delve much into how they trained the models and whether they trained from scratch or distilled from a larger model.
At those scales, I find LSTMs and CNNs are a lot more stable. The few online threads I've found comparing LSTMs and Transformers had the same thing to say - Transformers need a lot more data and model size to achieve parity and exceed LSTMs/GRUs/CNNs, maybe because the inductive bias provided is hard to beat at those scales. Others can comment on what they've seen.
[0] - https://arxiv.org/abs/2305.07759 https://arxiv.org/abs/2305.07759
- Al-Khwarizmi 2y agoI don't have much help to offer, but just to echo your experience... at my group we have tried to train Transformers from scratch for various NLP tasks and we always have been hit with them being extremely brittle, and BiLSTMs working better. We only succeeded by following a pre-established recipe (e.g. training a BERT model from scratch for a new language, where the architecture, parameters and tasks are as in BERT), or of course by fine-tuning existing models, but just throwing some layers at a problem and training them from scratch... nope, won't work without arcane knowledge that doesn't seem to be written anywhere accessible. This is one of the reasons why I dislike Transformers and I root for the likes of RWKV to take the throne.
- rakejake 2y agoI think the "arcane knowledge" is true for LLMs (billions). But there are lots of people who train models in the open in the hundreds of millions realm, but never below. Maybe transformers simply don't work as well below a size and data threshold.