5 ms·
Access to parallel corpora is a limiting factor in general. A good way to train a language translator is to use an open source dataset (several here http://opus
by saip 8y ago
Access to parallel corpora is a limiting factor in general. A good way to train a language translator is to use an open source dataset (several here http://opus.nlpl.eu/ http://opus.nlpl.eu/) to train a base model, and then fine-tune it with a smaller dataset specific to your domain.
In this case, the author claims pretty good accuracy, almost on par with Google Brain's!
On my test set of 3,000 sentences, the translator obtained a BLEU score of 0.39. This score is the benchmark scoring system used in machine translation, and the current best I could find in English to French is around 0.42 (set by some smart folks as Google Brain). So, not bad.
- jeffreyrogers 8y agoWow, missed that part when I read it. Pretty incredible that using open source data you can outperform the state-of-the-art machine translators of a few years ago.
- zawerf 8y agoFor a historical perspective check out stanford's nlp course: https://youtu.be/IxQtK2SjWWM?t=1267 https://youtu.be/IxQtK2SjWWM?t=1267 Deep learning only started beating tradition methods in 2016!
- samlynnevans 8y agoYes, it shouldn't have been far off Google as the model in the article is Google's itself. In fact you see in some of the examples how Google Translate's output and this model are almost exactly the same. Pretraining is a very cool idea, I have also seen some good results with pretrained embeddings after doing language modelling. Fast ai discusses this and I think even has some pretrained embeddings available in their library!