3 ms·
Andrej Karpathy goes into quite some depth about how tokenization works here: https://www.youtube.com/watch?v=zduSFxRajkE https://www.youtube.com/watch?v=zduSFx
by zellyn 2y ago
Andrej Karpathy goes into quite some depth about how tokenization works here: https://www.youtube.com/watch?v=zduSFxRajkE https://www.youtube.com/watch?v=zduSFxRajkE
tl;dr many of the LLMs use byte-pair encoding to create tokens. You take a set of documents, and then form tokens by repeatedly merging the most common pair of tokens. The initial set of tokens is 256 raw bytes. And the text is typically represented in utf-8.
I expect that although the LLMs can understand arbitrarily but cleanly offset unicode code points by (eventually) noticing the final byte of each sequence, they would do markedly worse on actually processing and completing on them, because they will not have been reduced to the normal set of tokens. However, if the text is actually output cleanly converted, either in internal thinking tokens or in the beginning of the response, they should do fine.
Understanding tokenization is surprisingly useful, even if that video seems awfully long to devote to such a tedious subject. Even Karpathy doesn't like it!