3 ms·
The compression algorithm is very similar to a greedy subword tokenizer, which is used in BERT and other older language models, but has become less popular in f
by stephantul 8mo ago
The compression algorithm is very similar to a greedy subword tokenizer, which is used in BERT and other older language models, but has become less popular in favor of BPE.