4 ms·
I'm glad you mentioned this. There is a lot of interest these days in character-based machine translation, including several papers in review at ICLR. The curre
by srush 10y ago
I'm glad you mentioned this. There is a lot of interest these days in character-based machine translation, including several papers in review at ICLR. The current practical consensus (at least in OpenNMT) is that character-only models are not really worth the efficiency loss. A simple compromise is to use Byte-Pair Encoding as a preprocessing step in morphologically rich languages and allow the model to produce sub-word chunks. This is implemented in OpenNMT as a preprocessing option (see http://opennmt.net/Advanced http://opennmt.net/Advanced).