4 ms·
I thought it was something to do with the way tokens are generated for the word strawberry? https://arbisoft.com/blogs/why-ll-ms-can-t-count-the-r-s-in-strawbe
by lbotos 2y ago
I thought it was something to do with the way tokens are generated for the word strawberry?
https://arbisoft.com/blogs/why-ll-ms-can-t-count-the-r-s-in-strawberry-and-what-it-teaches-us https://arbisoft.com/blogs/why-ll-ms-can-t-count-the-r-s-in-...
- marcosdumay 2y agoThat explanation would require the LLM to actually understand the question and deriving an answer from first principles. It doesn't.
- lbotos 2y ago?? If the input is parsed in to tokens, and the tokens split compound words, nothing about that requires "first principles" thinking to explain why LLMs struggle with getting all of the letters -- the LLM is only going down the vector path of one of the compound words... (I don't think LLMs are sentiment or intelligent btw, I think they are giant probability machines, and the probability that the LLM will get 3 r's on a token of "berry" are very low.)
- marcosdumay 2y agoThe LLM gives you the answer it finds on the training set. All the things on that article are irrelevant for the answer.
- lbotos 2y ago"The choice of tokenization method can directly affect the accuracy of character counting. If the tokenization method obscures the relationship between individual characters, it can be difficult for the LLM to count them accurately. For example, if "strawberry" is tokenized as "straw" and "berry," the LLM may not recognize that the two "r"s are part of the same word. To improve character counting accuracy, LLMs may need to use more sophisticated tokenization methods, such as subword tokenization or character-level tokenization, that can preserve more information about the structure of words."
- marcosdumay 2y agoWhat, again, assumes the LLM understood the question and is making an answer from first principles.
- lbotos 2y agoNo, it does not. You said above that "The LLM gives you the answer it finds on the training set" You and I both agree on that. No first principles there. The training set -- how's it built? With tokens. We have not trained LLMs with a token structure that deals well with compound words. If we trained LLMs with a different token structure, it is more probable that a one-shot answer for these compound word letter counting problems would be accurate. The LLM does not need to understand what "counting is" or even "what a letter is". The LLM will regurgitate the token relationship we train it on.