4 ms·
The 1.4T tokens are what the model was trained on, and not the token range of the embedding.
by bitRAKE 4y ago
The 1.4T tokens are what the model was trained on, and not the token range of the embedding.
- dentalperson 4y agoAh, that makes more sense, thank you. Since this was mentioned in the tokenizer section and the number of unique tokens wasn't mentioned I misunderstood.