3 ms·
> This is due to the fact that internally "пружи́на" is translated to English first, resulting in "spring" Do you have a source that that is actually what is h
by MiroF 6y ago
> This is due to the fact that internally "пружи́на" is translated to English first, resulting in "spring"
Do you have a source that that is actually what is happening behind the scenes? I work in this field and would be surprised if that is what Google is doing. My guess is this phenomenon is more due to a model that was trained on an Russian->English corpora first and then trained on Russian->German due to low resource, or something like that.
Someone like Yandex is going to have better translations because (shocker) they have larger russian corpora than google.
- qayxc 6y ago> Do you have a source that that is actually what is happening behind the scenes? Indirectly yes, I do: https://codesachin.wordpress.com/2017/01/18/understanding-the-new-google-translate/ https://codesachin.wordpress.com/2017/01/18/understanding-th... > What this basically means is: If during training you provide it examples of English->Japanese & English->Korean translations, GNMT automatically does Japanese->Korean reasonably well! In fact, this is the biggest achievement of GNMT as a project. It's exactly what Google have been doing for years now: train on X-to-English and English-to-Y only to get X-to-Y for free. Since the intermediate language only ever sees English as a source or target language, ambiguities like "spring" literally get lost in translation.
- MiroF 6y agoJust gave that a read. 1. It is pretty out of date (ML is a fast moving field) - I doubt Google is using LSTMs for translation in 2020. > Since the intermediate language only ever sees English as a source or target language, ambiguities like "spring" literally get lost in translation. 2. This article is trying to dumb down what Google is doing - but you're right that this is why ambiguities get lost in translation, due to pre-training on a different language pair. That said, there isn't literally a process of "translate into english" and then "translate english into german". These models are trained on Russian-German corpora, but because there is little resources for that, they are supplementing with Russian-English and English-German.
- qayxc 6y ago> It is pretty out of date (ML is a fast moving field) - I doubt Google is using LSTMs for translation in 2020. Sure, but that doesn't change the fact the training data is focused on English-to-X and X-to-English corpora. The underlying architecture of the model is just an implementation detail that doesn't really affect this as demonstrated by my example. > These models are trained on Russian-German corpora, but because there is little resources for that, they are supplementing with Russian-English and English-German. This is exactly what I'd argue isn't the case at all. Otherwise words that have a direct 1:1 translation wouldn't be mistranslated and companies like Yandex wouldn't be able to deliver so much better results. German-Russian isn't low-resource at all, given 95M and 150M native speakers respectively and a close history for the past 150 years. [edit]The rich cultural history of both countries resulting in a vast library of literature, theatre plays, news publications, films and the general cultural relevance of both languages is even more important.[/edit] It's simply (quite comprehensible) bias towards English for research taking place in the USA and the fact that it's much easier to compile English-to-X and X-to-English corpora in a predominantly English-speaking country. There are tons of translated books, films, news paper articles, scientific papers, etc. available for Russian-German and Yandex, being a Russian company, naturally has no problem compiling a Russian-German corpus (since they're not biased towards English).
- Shorel 6y agoWhatever the actual reason is, the results feel like English is used as an intermediate. This also happens between e.g. Spanish and Bulgarian.