3 ms·
> Instead of giving the system whole words, they give the system the word in parts. For example, the word “hello” might be given as 4 word parts “he” “l” “l” “o
by throwaway2246 8y ago
> Instead of giving the system whole words, they give the system the word in parts. For example, the word “hello” might be given as 4 word parts “he” “l” “l” “o”. This means we could learn a translation for the word “he” without the system ever having seen the word “he”.
Can anyone add context to this? Can't seem to wrap my head around this part. Doesn't "he" as a part of a word translate differently in different words?
- haneefmubarak 8y agoSure, but I think that would mean it would learn multiple meanings and the pattern groups that use each meaning. Perhaps a more appropriate example (just something I thought of, idk if it would work exactly like this) would be like "antiviral" going to "an" "ti" "vi" "r" "a" "l". It could connect "an" and "ti" and associate it with the meaning of "anti" as a prefix along with connecting "vi" and "r" and giving a possible association of "virus". Finally, it could combine "a" and "l" into the suffix "al" and recognize that meaning. Again, just my two cents. Not particularly sure if this is how that works.
- buboard 8y agoThey reference this paper on Byte Pair Encodings http://www.aclweb.org/anthology/P16-1162 http://www.aclweb.org/anthology/P16-1162 they tokenize text and then learn an embedding for those tokens using (in their case) both the source and target lanuage. This presumably captures regularities for languages that are not very different and has the benefit of having a small dictionary.
- slashcom 8y agoAn easier way to understand it is in the context morphology: word prefixes and suffixes mean things, and words have common roots. For example, polymorphism could be decomposed into poly-morph-ism. Antidisestablishmentarianism, which is unlikely to appear much in the corpus, becomes anti-dis-establish-ment-arian-ism. Now the system can learn how to reuse "anti-" or "establish" from other examples more easily than trying to learn the full word's meaning from the one or two examples it might see in the corpus. BPE is a clever way to induce these sort of decompositions automatically without any linguistic annotation, making them useful in multilingual settings. Other languages are much more morphologically rich than English, and there it really benefits.
- jstandard 8y agoThanks, this makes much more sense than the author's strange example.
- mlazos 8y agoThis is similar to the word hashing used in deep structured semantic models. [1] This works by representing a word as k-hot vector of trigrams that form the word - it is able to generalize better to unseen words than one-hot whole word embeddings probably due to word morphology as the other answer in this thread suggests. [1] https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/MSRTR2014_wordEmbedding.pdf https://www.microsoft.com/en-us/research/wp-content/uploads/...