3 ms·
Mm... I think you'd be hard pressed at this point to find any advantage that traditional methods have over deep methods. Maybe this was true 4-6 years ago, but
by MiroF 6y ago
Mm... I think you'd be hard pressed at this point to find any advantage that traditional methods have over deep methods.
Maybe this was true 4-6 years ago, but no longer.
- _-___________-_ 6y agoDepends what you optimise for. Neural translation fairly consistently produces incorrect results that read like they could be correct, for me. If your goal is "text that makes sense" they seem to be doing great. If your goal is "text that conveys the same meaning as the source" they're universally rubbish, at least for the language pairs I need (mostly english <-> east Asian languages).
- MiroF 6y ago> english <-> east Asian languages I'm curious what the domain is and what languages you're talking about, it's hard for me to evaluate your claims otherwise. I would buy it for English -> Tibetan for instance.
- _-___________-_ 6y agoMostly Chinese, Japanese, and Vietnamese, and mostly business and lightly-technical text. My observation of the failure rate (if you define failure as "meaning has been lost") is well above 80% with Google Translate for these languages and this type of text. It's literally unusable because you have no confidence that the translation has not completely reversed or garbled the meaning of the text, and the failure mode is particularly unpleasant because you may be unaware that it has failed to translate correctly, since it produces text that appears to have meaning - just the wrong meaning.
- numpad0 6y agoKinda funny when people say "East Asia" it mostly means Japan and its isolationism, with China mixed in just to change the tone a bit. I don't find it offensive, just kind of funny/interesting.
- _-___________-_ 6y agoWell, in the context of language it's a useful term because of the close relationship many east Asian languages have.
- numpad0 6y agoDo those languages really have that relationship? When I scratched surface of Chinese language in university it felt almost English written in Kanji. No resemblance to Japanese, and it even made the classical Chinese interpretation methods taught in Japanese education rather dubious to me.
- _-___________-_ 6y agoAlthough they belong to historically disparate language families, modern Vietnamese and Japanese vocabulary both have significant influence from Chinese languages. I know little about Korean but I imagine it's the case there too. Your comment about Chinese feeling like English written in Kanji is really interesting to me, perhaps it's a case of different perspectives - I tend to focus on grammar when learning a language, and so Chinese feels very different to English from my perspective.
- thaumasiotes 6y agoJapanese and Chinese languages are not related by descent. They differ in many very basic ways. But Asian languages are nevertheless related in more mysterious ways. Compare the intro to https://en.wikipedia.org/wiki/Mainland_Southeast_Asia_linguistic_area https://en.wikipedia.org/wiki/Mainland_Southeast_Asia_lingui... : > The Mainland Southeast Asia linguistic area is a sprachbund including languages of the Sino-Tibetan, Hmong–Mien (or Miao–Yao), Kra–Dai, Austronesian and Austroasiatic families spoken in an area stretching from Thailand to China. Neighbouring languages across these families, though presumed unrelated, often have similar typological features, which are believed to have spread by diffusion. In a European context, this is similar to the unusual closeness between the "unrelated" English and French, known to be due to intensive contact between the languages post-William-the-Conqueror. (In a world context, English and French are closely related regardless.)
- MiroF 6y ago> It's literally unusable It's not "literally unusable." Even for language pairs with poor translation, it considerably saves on manual translation time if you can have a translator look over sets of already translated (but lower-quality) sentences and correct those with errors. Plus, there are actual studies out there comparing human translations to machine translations in a double-blinded evaluation by translators and the gap is not as high as your comment suggests. For chinese-english (the other direction), bilingual evaluators have had difficulty differentiating between the machine translation and the human translation.
- _-___________-_ 6y ago> For chinese-english (the other direction), bilingual evaluators have had difficulty differentiating between the machine translation and the human translation. Having attempted to use CN->EN machine translation many times, I find this very surprising. Have you got a reference I can have a look at?
- 082349872349872 6y agoI strongly prefer the disfluent but more likely accurate results from the old days to the fluent-appearing but of unknown accuracy of today. https://news.ycombinator.com/item?id=24302564 https://news.ycombinator.com/item?id=24302564 Bonus clip: https://www.youtube.com/watch?v=0E4aYgKmIko https://www.youtube.com/watch?v=0E4aYgKmIko
- MiroF 6y ago"the disfluent but more likely accurate results" -> Human evaluation does not agree that PBSMT is more accurate than NMT. Sure, the fluency/"sounds good" gains from NMT are much larger than the accuracy gains, but evaluations I've seen put NMT ahead on both metrics.
- 082349872349872 6y agoFair enough, this is just my sentiment, where accurate was not a term of art, but meant "honest where it is weak and where it is confident." Before, I could skim outputs, and get a fair idea of which areas were poorly translated by where it was garbage. Now, I get the impression that (like with GPT or style transfer) something always comes out, whether it's related (or even necessarily has the right negation parity) or not.
- MiroF 6y agoYou're absolutely right with your impression of GPT/these models - they will generally generate plausible text. And it is also worth noting (as you've sort of mentioned) that appearance of fluency can bias translators who are instructed to just evaluate accuracy, or to evaluate accuracy and fluency separately. There is ongoing research into that problem.