5 ms·
I work in an industry where most of the text is not natural — super wide vocabulary, lots of symbols and abbreviations. I built some utf-8 tokenized models that
by deepsquirrelnet 4y ago
I work in an industry where most of the text is not natural — super wide vocabulary, lots of symbols and abbreviations. I built some utf-8 tokenized models that do very well in comparison to sentence/word piece tokenizers.
Unless your corpus fits the tokenizer vocabulary very well, I think there are many cases where “token free” models can be advantageous. You only have to determine how to deal with the expansion of the input dimension and how it might blow up your attention costs.
- optimalsolver 4y agoInteresting. Could you give some examples of "words" in this field?
- jszymborski 4y agoCan't you just train a piece tokenizer like SentencePiece on your irregular corpus? That's what I've done with my amino-acid model. Or are you saying your tokenizer does better than that.
- deepsquirrelnet 4y agoI think in most applications standard tokenizers work fine. There’s a point where too much of your vocabulary gets chopped up into single tokens anyway, and it ends up performing much worse. I’m not able to quantify it exactly for you. I’d have to benchmark it over some synthetic data to get a better idea. One could imagine a situation where the opportunity for misspelling is very high - say, for example OCR of medium to low quality images. By not using a character-level tokenizer, your model also has to learn to associate complete tokens to any number of incomplete tokens. Instead, it’s simpler to just let the model infer words based on complete or mostly complete character sequences, rather than randomly chopped up word piece tokens.