2 ms·
On my team, there was not intially acceptance from folks that were used to traditional nlp techniques. The models built with transformers were viewed as over p
by bradfox2 3y ago
On my team, there was not intially acceptance from folks that were used to traditional nlp techniques. The models built with transformers were viewed as over parameterized.
It really wasn't until the original BERT paper came out and topped the superglue leaderboards that the move from lstm-cnn based architectures started. I remember feeling at the time that a ~300M parameter model was absolutely huge.