3 ms·
It's tricky to give you an exact answer. For translation, the minimal size we have been using is about 1 million aligned sentences, although people often report
by srush 10y ago
It's tricky to give you an exact answer. For translation, the minimal size we have been using is about 1 million aligned sentences, although people often report on smaller data. There are also lots of tricks to get around small dataset problem.
(1) You can pre-initialize your model with monolingual word embeddings. We recommend using the Polyglot embeddings which exist for many different languages. See http://opennmt.net/Advanced http://opennmt.net/Advanced for details.
(2) You can train your model with a nearby language. For instance we have a model that uses data from all the Romance languages simultaneously.
(3) You can use monolingual data in other ways. For instance, if you are translating into English, you can combine with a standard language model or pretrain on the English data.
There are a bunch of other approaches in the literature, but these are some of the more common tricks.