3 ms·
No it does not, we are thinking about it. Extra segmentation of tokens is a great idea, I only don't like (who cares? :-)) inconsistencies that BPE sometimes p
by sergeio76 7y ago
No it does not, we are thinking about it. Extra segmentation of tokens is a great idea, I only don't like (who cares? :-)) inconsistencies that BPE sometimes produce. Looking at how BERT tokenizers for example: "tranformers" is one token but "transformer" is segmented into "transform" "##er". Word "outperforms" becomes "out", "##per", "##forms". It is truly amazing that BERT works so well with such tokenization.
- make3 7y agoimho the fact that character level language models work mostly just as well as word level or bpe tokenizers makes it clear that the "embedding" view is wrong nowadays, that the meaning is really built (and stored) in the layers, just fine
- misterman0 7y agoIn my probably even more humble opinion, due to the fact I'm merly a novice NLP researcher, I think you are spot on. I've been experimenting with layering two models on top of each other, one using the utmost banal character based token embedding (bags-of-characters) with clustering, the other using embeddings as wide as there are clusters, and found that even though the tokens "apple" and "elppa" count as one and the same, the context of how the individual tokens are used together is enough for very precisely localizing the right cluster of documents. It seems context is more important than the individual words are. There is also this curious fact that humans have no trouble at all decoding tokens even tough teh the characters have been displaced, which seems to support the idea that a model as trivial (and easy to compute on) as bags of characters is indeed a valid one.
- sergeio76 7y agoMaybe you are right from theoretical point of view, however 1.latency at inference time may be an overkill, 2.look at success of fasttext model from Facebook, it computes vectors for words and ngrams! as well. What you will have to learn with layers and filters this model memorizes. It trains faster hence you can feed huge training data, it runs fast as well.
- misterman0 7y agoPerhaps I'm mistaken because I thought blingfire was mostly a tokenizer. But now I realize it might also be a language model. How are tokens/phrases/documents represented, computationaly? Theoretically, do they live in a vector space?
- lsorber 7y agoLooks like an opportunity to do BPE right. Would love to see your take on it with this engine!