4 ms·
"The choice of tokenization method can directly affect the accuracy of character counting. If the tokenization method obscures the relationship between individu
by lbotos 2y ago
"The choice of tokenization method can directly affect the accuracy of character counting. If the tokenization method obscures the relationship between individual characters, it can be difficult for the LLM to count them accurately. For example, if "strawberry" is tokenized as "straw" and "berry," the LLM may not recognize that the two "r"s are part of the same word.
To improve character counting accuracy, LLMs may need to use more sophisticated tokenization methods, such as subword tokenization or character-level tokenization, that can preserve more information about the structure of words."
- marcosdumay 2y agoWhat, again, assumes the LLM understood the question and is making an answer from first principles.
- lbotos 2y agoNo, it does not. You said above that "The LLM gives you the answer it finds on the training set" You and I both agree on that. No first principles there. The training set -- how's it built? With tokens. We have not trained LLMs with a token structure that deals well with compound words. If we trained LLMs with a different token structure, it is more probable that a one-shot answer for these compound word letter counting problems would be accurate. The LLM does not need to understand what "counting is" or even "what a letter is". The LLM will regurgitate the token relationship we train it on.