3 ms·
Just gave that a read. 1. It is pretty out of date (ML is a fast moving field) - I doubt Google is using LSTMs for translation in 2020. > Since the intermedi
by MiroF 6y ago
Just gave that a read.
1. It is pretty out of date (ML is a fast moving field) - I doubt Google is using LSTMs for translation in 2020.
> Since the intermediate language only ever sees English as a source or target language, ambiguities like "spring" literally get lost in translation.
2. This article is trying to dumb down what Google is doing - but you're right that this is why ambiguities get lost in translation, due to pre-training on a different language pair.
That said, there isn't literally a process of "translate into english" and then "translate english into german". These models are trained on Russian-German corpora, but because there is little resources for that, they are supplementing with Russian-English and English-German.
- qayxc 6y ago> It is pretty out of date (ML is a fast moving field) - I doubt Google is using LSTMs for translation in 2020. Sure, but that doesn't change the fact the training data is focused on English-to-X and X-to-English corpora. The underlying architecture of the model is just an implementation detail that doesn't really affect this as demonstrated by my example. > These models are trained on Russian-German corpora, but because there is little resources for that, they are supplementing with Russian-English and English-German. This is exactly what I'd argue isn't the case at all. Otherwise words that have a direct 1:1 translation wouldn't be mistranslated and companies like Yandex wouldn't be able to deliver so much better results. German-Russian isn't low-resource at all, given 95M and 150M native speakers respectively and a close history for the past 150 years. [edit]The rich cultural history of both countries resulting in a vast library of literature, theatre plays, news publications, films and the general cultural relevance of both languages is even more important.[/edit] It's simply (quite comprehensible) bias towards English for research taking place in the USA and the fact that it's much easier to compile English-to-X and X-to-English corpora in a predominantly English-speaking country. There are tons of translated books, films, news paper articles, scientific papers, etc. available for Russian-German and Yandex, being a Russian company, naturally has no problem compiling a Russian-German corpus (since they're not biased towards English).