5 ms·
Facebook's AI Just Set a New Record in Translation
- pnloyd 8y ago>> For example, “translate” between neural activity in the brain to videos on a screen, That sounds almost to good to be true. Excited to see what gets developed with these techniques!
- delhanty 8y agoThe referenced link to the original work on Facebook code looks more suitable for HN: https://code.fb.com/ai-research/unsupervised-machine-translation-a-novel-approach-to-provide-fast-accurate-translations-for-more-languages/ https://code.fb.com/ai-research/unsupervised-machine-transla...
- gwern 8y agoAnd was already submitted: https://news.ycombinator.com/item?id=17886827 https://news.ycombinator.com/item?id=17886827
- gok 8y agoIn unsupervised translation, to be specific.
- personjerry 8y agoThis seems like it would only work for similar languages (i.e. romance languages), in that it depends on the embeddings within languages to be similar.
- maneesh 8y agoIs Urdu a romance language? Or related closely to English ?
- lgessler 8y agoIt's a distant relative: Urdu is Indo-Aryan, which is the language family Sanskrit belonged to, and Indo-Aryan is in the Indo-European language family, to which English and Romance languages also belong. GP's point is still a good one though: while Urdu and English have diverged quite a bit despite being of the same stock, they probably still share a lot more typologically than, say, English and Mandarin or Warlpiri.
- rainraingoaway 8y agoUrdu is not even remotely close to being a romance language. Urdu is basically a dialect of hindi, with it's own alphabet derived from Arabic, and loan words from Persian and Arabic.
- yesenadam 8y agoWell, they're both Indo-European - evolved from the same language.
- eganist 8y agoHeads up for those of you obliging Forbes' forced resistance against adblocking: There's (edit: what appears to be) an active exploit in their ad network, one that's getting around Chrome's redirect blocking through an apparent 0day. https://imgur.com/a/sRIB7pn https://imgur.com/a/sRIB7pn I'm on Chrome Beta 69.0.3497.53 on Android, so this may not apply outside that. Chrome team: https://bugs.chromium.org/p/chromium/issues/detail?id=879938 https://bugs.chromium.org/p/chromium/issues/detail?id=879938
- infogulch 8y agoHow did we paint ourselves into this corner where the only way for our websites to exist is to run a different person's arbitrary code on our visitor's devices every time they open a page?
- jdangu 8y agoChrome's protection only works in cross-origin iframes [1] and has been in beta for years. I haven't checked in a while but can't find a source that confirms that it went live. Forbes serves a large portion of their ads in same origin iframes and so is not fully covered by this protection. [1] https://blog.chromium.org/2017/11/expanding-user-protections-on-web.html https://blog.chromium.org/2017/11/expanding-user-protections...
- deleted 8y ago[deleted]
- throwaway2246 8y ago> Instead of giving the system whole words, they give the system the word in parts. For example, the word “hello” might be given as 4 word parts “he” “l” “l” “o”. This means we could learn a translation for the word “he” without the system ever having seen the word “he”. Can anyone add context to this? Can't seem to wrap my head around this part. Doesn't "he" as a part of a word translate differently in different words?
- haneefmubarak 8y agoSure, but I think that would mean it would learn multiple meanings and the pattern groups that use each meaning. Perhaps a more appropriate example (just something I thought of, idk if it would work exactly like this) would be like "antiviral" going to "an" "ti" "vi" "r" "a" "l". It could connect "an" and "ti" and associate it with the meaning of "anti" as a prefix along with connecting "vi" and "r" and giving a possible association of "virus". Finally, it could combine "a" and "l" into the suffix "al" and recognize that meaning. Again, just my two cents. Not particularly sure if this is how that works.
- buboard 8y agoThey reference this paper on Byte Pair Encodings http://www.aclweb.org/anthology/P16-1162 http://www.aclweb.org/anthology/P16-1162 they tokenize text and then learn an embedding for those tokens using (in their case) both the source and target lanuage. This presumably captures regularities for languages that are not very different and has the benefit of having a small dictionary.
- slashcom 8y agoAn easier way to understand it is in the context morphology: word prefixes and suffixes mean things, and words have common roots. For example, polymorphism could be decomposed into poly-morph-ism. Antidisestablishmentarianism, which is unlikely to appear much in the corpus, becomes anti-dis-establish-ment-arian-ism. Now the system can learn how to reuse "anti-" or "establish" from other examples more easily than trying to learn the full word's meaning from the one or two examples it might see in the corpus. BPE is a clever way to induce these sort of decompositions automatically without any linguistic annotation, making them useful in multilingual settings. Other languages are much more morphologically rich than English, and there it really benefits.
- DataJunkie 8y agoOk, but do they actually use it for anything?