10 ms·
Anything to do with counting words or letters, manipulating letters, etc is just leading to false inferences about the power of the model, because it doesn't wo
by carrolldunham 3y ago
Anything to do with counting words or letters, manipulating letters, etc is just leading to false inferences about the power of the model, because it doesn't work with words and letters but tokens. It can't see words.
- deleted 3y ago[deleted]
- aezart 3y agoOne thing I'm not clear about is how the LLM decides on the correct tokenization to use. Looking at GPT2's token vocabulary, it has the word "the" as token number 1169, but it also has "t", "h", and "e" separately as tokens 83, 71, and 68. When the tokenizer is reading the prompt, does it greedily create the longest tokens possible from the start of the input, or is there a more sophisticated process going on? I guess in theory it could try multiple tokenizations and see which one results in the overall best response.