3 ms·
It's relatively small modification of BERT with multi-task fine-tuning and slightly different output heads. It should be easy for any NLP researcher to replicat
by phowon 8y ago
It's relatively small modification of BERT with multi-task fine-tuning and slightly different output heads. It should be easy for any NLP researcher to replicate.
- riku_iki 8y agoexcept you need significant GPU/TPU resources to pretrain language model.
- bitL 8y agoYou can't even train BERT_large on a 12/16GB GPU, and on a single 15TFlops GPU it might take a year to train. GPUs are too slow :-(
- riku_iki 8y agoTPU is also slow, they used pod with 64 TPUs for training BERT. You probably can achieve similar result using distributed training on multiple GPU machines.
- solomatov 8y agoYou can but it will be really slow. You can load just parts of the model, and store them on the disk/in memory :-)
- solomatov 8y agoThe authors of the paper didn't pretrain the language model. They used an existing BERT and fine tuned it in a novel way.
- riku_iki 8y agoCould you provide citation? I tried to find this but couldn't.
- solomatov 8y ago>The training procedure of MT-DNN consists of two stages: pretraining and multi-task fine-tuning. The pretraining stage follows that of the BERT model (Devlin et al., 2018). The parameters of the lexicon encoder and Transformer encoder are learned using two unsupervised prediction tasks: masked language modeling and next sentence pre- diction.3 and this: >Our implementation of MT-DNN is based on the PyTorch implementation of BERT4. We used Adamax (Kingma and Ba, 2014) as our optimizer with a learning rate of 5e-5 and a batch size of 32. The maximum number of epochs was set to 5. A linear learning rate decay schedule with warm-up over 0.1 was used, unless stated otherwise. Fol- lowing (Liu et al., 2018a), we set the number of steps to 5 with a dropout rate of 0.1. To avoid the exploding gradient problem, we clipped the gradi- ent norm within 1. All the texts were tokenized using wordpieces, and were chopped to spans no longer than 512 tokens. You won't be able to train BERT in 3 epochs.
- solomatov 8y agoHere's the quote from BERT: >We train with batch size of 256 sequences (256 sequences * 512 tokens = 128,000 tokens/batch) for 1,000,000 steps, which is approximately 40 epochs over the 3.3 billion word corpus.
- danielcampos93 8y agoCan confirm from conversations I had with the authors.