3 ms·
Maybe this is why? Most of the training data has the single token version, so the three tokens version was undertrained?
by nialv7 6mo ago
Maybe this is why? Most of the training data has the single token version, so the three tokens version was undertrained?