3 ms·
I think in most applications standard tokenizers work fine. There’s a point where too much of your vocabulary gets chopped up into single tokens anyway, and it
by deepsquirrelnet 4y ago
I think in most applications standard tokenizers work fine. There’s a point where too much of your vocabulary gets chopped up into single tokens anyway, and it ends up performing much worse. I’m not able to quantify it exactly for you. I’d have to benchmark it over some synthetic data to get a better idea.
One could imagine a situation where the opportunity for misspelling is very high - say, for example OCR of medium to low quality images. By not using a character-level tokenizer, your model also has to learn to associate complete tokens to any number of incomplete tokens. Instead, it’s simpler to just let the model infer words based on complete or mostly complete character sequences, rather than randomly chopped up word piece tokens.