3 ms·
Source? I do see this note in the link: > Please note that it is advised to avoid using the Hugging Face fast tokenizer for now, as we’ve observed that the aut
by ehsanu1 3y ago
Source? I do see this note in the link:
> Please note that it is advised to avoid using the Hugging Face fast tokenizer for now, as we’ve observed that the auto-converted fast tokenizer sometimes gives incorrect tokenizations.
If you are correct, maybe it'll work for languages not reliant on whitespace, paired with a code formatter, but we're definitely coming to expect more from these models these days.