3 ms·
Andrej Karpathy's take from twitter. (https://twitter.com/karpathy/status/1760350892317098371 https://twitter.com/karpathy/status/1760350892317098371) Seeing a
by BryanLegend 3y ago
Andrej Karpathy's take from twitter. (https://twitter.com/karpathy/status/1760350892317098371 https://twitter.com/karpathy/status/1760350892317098371)
Seeing as I published my Tokenizer video yesterday, I thought it could be fun to take a deepdive into the Gemma tokenizer.
First, the Gemma technical report [pdf]:
https://storage.googleapis.com/deepmind-media/gemma/gemma-report.pdf https://storage.googleapis.com/deepmind-media/gemma/gemma-re...
says: "We use a subset of the SentencePiece tokenizer (Kudo and Richardson, 2018) of Gemini for com- patibility. It splits digits, does not remove extra whitespace, and relies on byte-level encodings for unknown tokens, following the techniques used for both (Chowdhery et al., 2022) and (Gemini Team, 2023). The vocabulary size is 256k tokens."
The tokenizer.model file is with this code release:
https://github.com/google/gemma_pytorch/blob/main/tokenizer/tokenizer.model https://github.com/google/gemma_pytorch/blob/main/tokenizer/...
I decoded this model protobuf in Python and here is the diff with the Llama 2 tokenizer:
https://diffchecker.com/TRnbKRMH/ https://diffchecker.com/TRnbKRMH/
Notes:
- vocab size is quite large: 32K -> 256K
- add_dummy_prefix is False. Different from Llama but consistent with GPT. This is a bit more consistent w.r.t. "leave the data alone", as there is no preprocessing step that adds a space to the encoding text.
- the model_prefix is the path of the training dataset, which is amusing to look at: "/cns/mf-d/home/gemini-data-access/tokenizers/final_v1_51GB_run1/bpe_coverage_0_999995_v5/255969". Seems to indicate the tokenizer training corpus was ~51GB (?).
- a lot of user_defined symbols (i.e. special tokens) are present, e.g. "hardcoding" a sequence of up to 31 newlines as tokens, and a large number of other unclear tokens. I tried decoding the octal representations but it's not clear what's happening here. Also a lot of more special tokens for what look like html elements, e.g. <table>, <tr>, <td>, <i>, <b>, etc. Not 100% sure what the unused tokens are for, maybe this is pre-allocated space to make easier future finetunes that try to add more special tokens, as there is no need to resize vocabularies and perform model surgeries (?).
TLDR this is basically the Llama 2 tokenizer, except bigger (32K -> 256K), with a lot more special tokens, and the only functional departure is that add_dummy_prefix is turned off to False. So e.g. tokenizing:
"hello world" becomes:
[17534, 2134]
['hello', 'world']
which otherwise would have been preprocessed to " hello world" (note leading space) and tokenized as:
[25612, 2134]
['hello', 'world']
cool