9 ms·
Testing models on their tokenization has always struck me as kinda odd. Like, that has nothing to do with their intelligence.
by generalizations 2y ago
Testing models on their tokenization has always struck me as kinda odd. Like, that has nothing to do with their intelligence.
- b3nny 2y ago[dead]
- swatcoder 2y agoSurfacing and underscoring obvious failure cases for general "helpful chatbot" use is always going to be valuable because it highlights how the "helpful chatbot" product is not really intuitively robust. Meanwhile, it helps make sure engineers and product designers who want to build a more targeted product around LLM technology know that it's not suited to tasks that may trigger those kinds of failures. This may be obvious to you as an engaged enthusiast or cutting edge engineer or whatever you are, but it's always going to be new information to somebody as the field grows.
- wruza 2y agoIt doesn’t test “on tokenization” though. What happens when an answer is generated is few abstraction levels deeper than tokens. A “thinking” “slice” of an llm is completely unaware of tokens as an immediate part of its reasoning. The question just shows lack of systemic knowledge about strawberry as a word (which isn’t surprising, tbh).
- qeternity 2y agoIt is. Strawberry is one token in many tokenziers. The model doesn't have a concept that there are letters there.
- guywhocodes 2y agoThis is pretty much equivalent to the statement "multicharacter tokens are a dead end for understanding text". Which I agree with.
- sebzim4500 2y agoThat doesn't follow from what he said at all. Knowing how to spell words and understanding them are basically unrelated tasks.
- abdullahkhalids 2y agoIf I ask an LLM to generate new words for some concept or category, it can do that. How do the new words form, if not from joining letters?
- mirekrusin 2y agoNot letters, but tokens. Think that it's translating everything to/from Chinese.
- abdullahkhalids 2y agoHow does that explain why the tokens for strawberry, melon and "Stellaberry" [1] are close to each other? [1] Suggestion from chatgpt3.5 for new fruit name.
- roywiggins 2y agoIlliterate humans can come up with new words like that too without being able to spell, LLMs are modeling language without precisely modeling spelling.
- coder543 2y agoThe tokenizer system supports virtually any input text that you want, so it follows that it also allows virtually any output text. It isn’t limited to a dictionary of the 1000 most common words or something. There are tokens for individual letters, but the model is not trained on text written with individual tokens per letter, it is trained on text that has been converted into as few tokens as possible. Just like you would get very confused if someone started spelling out entire sentences as they spoke to you, expecting you to reconstruct the words from the individual spoken letters, these LLMs also would perform terribly if you tried to send them individual tokens per letter of input (instead of the current tokenizer scheme that they were trained on). Even though you might write a message to an LLM, it is better to think of that as speaking to the LLM. The LLM is effectively hearing words, not reading letters.
- alew1 2y agoIf I show you a strawberry and ask how many r’s are in the name of this fruit, you can tell me, because one of the things you know about strawberries is how to spell their name. Very large language models also “know” how to spell the word associated with the strawberry token, which you can test by asking them to spell the word one letter at a time. If you ask the model to spell the word and count the R’s while it goes, it can do the task. So the failure to do it when asked directly (how many r’s are in strawberry) is pointing to a real weakness in reasoning, where one forward pass of the transformer is not sufficient to retrieve the spelling and also count the R’s.
- viraptor 2y agoThat's not always true. They often fail the spelling part too.
- qeternity 2y agoSure, that's a different issue. If you prompt in a way to invoke chain of thought (e.g. what humans would do internally before answering) all of the models I just tested got it right.
- sebastiennight 2y ago> If I show you a strawberry and ask how many r’s are in the name of this fruit, you can tell me, because one of the things you know about strawberries is how to spell their name. LOL. I would fail your test, because "fraise" only has one R, and you're expecting me to reply "3".
- wruza 2y agoThe thinking part of a model doesn’t know about tokens either. Like a regular human few thousand years ago didn’t think of neural impulses or air pressure distribution when talking. It might “know” about tokens and letters like you know about neurons and sound, but not access them on the technical level, which is completely isolated from it. The fact that it’s a chat of tokens of letters, which are a form of information passing between humans, is accidental.
- probably_wrong 2y agoI would counterargue with "that's the model's problem, not mine". Here's a thought experiment: if I gave you 5 boxes and told you "how many balls are there in all of this boxes?" and you answered "I don't know because they are inside boxes", that's a fail. A truly intelligent individual would open them and look inside. A truly intelligent model would (say) retokenize the word into its individual letters (which I'm optimistic they can) and then would count those. The fact that models cannot do this is proof that they lack some basic building blocks for intelligence. Model designers don't get to argue "we are human-like except in the tasks where we are not".
- pegasus 2y agoOf course they lack building blocks for full intelligence. They are good at certain tasks, and counting letters is emphatically not one of them. They should be tested and compared on the kind of tasks they're fit for, and so the kind of tasks they will be used in solving, not tasks for which they would be misemployed to begin with.
- probably_wrong 2y agoI agree with you, but that's not what the post claims. From the article: "A significant effort was also devoted to enhancing the model’s reasoning capabilities. (...) the new Mistral Large 2 is trained to acknowledge when it cannot find solutions or does not have sufficient information to provide a confident answer." Words like "reasoning capabilities" and "acknowledge when it does not have enough information" have meanings. If Mistral doesn't add footnotes to those assertions then, IMO, they don't get to backtrack when simple examples show the opposite.
- pegasus 2y agoYou're right, I missed that claim.
- mrkstu 2y agoIts not like an LLM is released with a hit list of "these are the tasks I really suck at." Right now users have to figure it out on the fly or have a deep understanding of how tokenizers work. That doesn't even take into account what OpenAI has typically done to intercept queries and cover the shortcomings of LLMs. It would be useful if each model did indeed come out with a chart covering what it cannot do and what it has been tailored to do above and beyond the average LLM.
- SirMaster 2y agoHow is a layman supposed to even know that it's testing on that? All they know is it's a large language model. It's not unreasonable they should expect it to be good at things having to do with language, like how many letters are in a word. Seems to me like a legit question for a young child to answer or even ask.
- stavros 2y ago> How is a layman supposed to even know that it's testing on that? They're not, but laymen shouldn't think that the LLM tests they come up with have much value.
- SirMaster 2y agoI'm saying a layman or say a child wouldn't even think this is a "test". They are just asking a language model a seemingly simple language related question from their point of view.
- groby_b 2y agolayman or children shouldn't use LLMs. They're pointless unless you have the expertise to check the output. Just because you can type text in a box doesn't mean it's a tool for everybody.
- SirMaster 2y agoWell they certainly aren't being marketed or used that way... I'm seeing everyone and their parents using chatgpt.
- meroes 2y agoI hear this a lot but there are vast sums of money thrown at where a model fails the strawberry cases. Think about math and logic. If a single symbol is off, it’s no good. Like a prompt where we can generate a single tokenization error at my work, by my very rough estimates, generates 2 man hours of work. (We search for incorrect model responses, get them to correct themselves, and if they can’t after trying, we tell them the right answer, and edit it for perfection). Yes even for counting occurrences of characters. Think about how applicable that is. Finding the next term in a sequence, analyzing strings, etc.
- antonvs 2y ago> Think about math and logic. If a single symbol is off, it’s no good. In that case the tokenization is done at the appropriate level. This is a complete non-issue for the use cases these models are designed for.
- meroes 2y agoBut we don’t restrict it to math or logical syntax. Any prompt across essentially all domains. The same model is expected to handle any kind of logical reasoning that can be brought into text. We don’t mark it incorrect if it spells an unimportant word wrong, however keep in mind the spelling of a word can be important for many questions, for example—off the top of my head: please concatenate “d”, “e”, “a”, “r” into a common English word without rearranging the order. The types of examples are endless. And any type of example it gets wrong, we want to correct it. I’m not saying most models will fail this specific example, but it’s to show the breadth of expectations.
- baq 2y agoCall me when models understand when to convert the token into actual letters and count them. Can’t claim they’re more than word calculators before that.
- jahsome 2y agoIs anyone in the know, aside from mainstream media (god forgive me for using this term unironically) and civillians on social media claiming LLMs are anything but word calculators? I think that's a perfect description by the way, I'm going to steal it.
- dTal 2y agoI think it's a very poor intuition pump. These 'word calculators' have lots of capabilities not suggested by that term, such as a theory of mind and an understanding of social norms. If they are a "merely" a "word calculator", then a "word calculator" is a very odd and counterintuitively powerful algorithm that captures big chunks of genuine cognition.
- robbiep 2y agoThey’re trained on the available corpus of human knowledge and writings. I would think that the word calculators have failed if they were unable to predict the next word or sentiment given the trillions of pieces of data they’ve been fed. Their training environment is literally people talking to each other and social norms. Doesn’t make them anything more than p-zombies though. As an aside, I wish we would call all of this stuff pseudo intelligence rather than artificial intelligence
- dTal 2y agoI side with Dennett (and Turing for that matter) that a "p-zombie" is a logically incoherent thing. Demonstrating understanding is the same as having understanding because there is no test that can distinguish the two. Are LLMs human? No. Can they do everything humans do? No. But they can do a large enough subset of things that until now nothing but a human could do that we have no choice but to call it "thinking". As Hofstadter says - if a system is isomorphic to another one, then its symbols have "meaning", and this is indeed the definition of "meaning".
- psb217 2y agoHow can I know whether any particular question will test a model on its tokenization? If a model makes a boneheaded error, how can I know whether it was due to lack of intelligence or due to tokenization? I think finding places where models are surprisingly dumb is often more informative than finding particular instances where they seem clever. It's also funny, since this strawberry question is one where a model that's seriously good at predicting the next character/token/whatever quanta of information would get it right. It requires no reasoning, and is unlikely to have any contradicting text in the training corpus.
- viraptor 2y ago> How can I know whether any particular question will test a model on its tokenization? Does something deal with separate symbols rather than just meaning of words? Then yes. This affects spelling, math (value calculation), logic puzzles based on symbols. (You'll have more success with a puzzle about "A B A" rather than "ABA") > It requires no reasoning, and is unlikely to have any contradicting text in the training corpus. This thread contains contradictions. Every other announcement of an llm contains a comment with a contradicting text when people post the wrong responses.
- VincentEvans 2y agoI don’t know anything about LLMs beyond using ChatGPT and Copilot… but unless because of this lack of knowledge I am misinterpreting your reply - it sounds as if you are excusing the model giving a completely wrong answer to a question that anyone intelligent enough to learn alphabet can answer correctly.
- microtonal 2y agoThe problem is that the model never gets to see individual letters. The tokenizers used by these models break up the input in pieces. Even though the smallest pieces/units are bytes in most encodings (e.g. BBPE), the tokenizer will cut up most of the input in much larger units, because the vocabulary will contain fragments of words or even whole words. For example, if we tokenize Welcome to Hacker News, I hope you like strawberries. The Llama 405B tokenizer will tokenize this as: Welcome Ġto ĠHacker ĠNews , ĠI Ġhope Ġyou Ġlike Ġstrawberries . (Ġ means that the token was preceded by a space.) Each of these pieces is looked up and encoded as a tensor with their indices. Adding a special token for the beginning and end of the text, giving: [128000, 14262, 311, 89165, 5513, 11, 358, 3987, 499, 1093, 76203, 13] So, all the model sees for 'Ġstrawberries' is the number 76204 (which is then used in the piece embedding lookup). The model does not even have access to the individual letters of the word. Of course, one could argue that the model should be fed with bytes or codepoints instead, but that would make them vastly less efficient with quadratic attention. Though machine learning models have done this in the past and may do this again in the future. Just wanted to finish of this comment with saying that the tokens might be provided in the model splitted if the token itself is not in the vocabulary. For instance, the same sentence translated to my native language is tokenized as: Wel kom Ġop ĠHacker ĠNews , Ġik Ġhoop Ġdat Ġje Ġvan Ġa ard be ien Ġh oud t . And the word voor strawberries (aardbeien) is split, though still not in letters.
- TiredOfLife 2y agoThe thing is, how the tokenizing work is about as relevant to the person asking the question as name of the cat of the delivery guy who delivered the GPU that the llm runs on.
- ca_tech 2y agoIts like showing someone a color and asking how many letters it has. 4... 3? blau, blue, azul, blu The color holds the meaning and the words all map back. In the model the individual letters hold little meaning. Words are composed of letters but simply because we need some sort of organized structure for communication that helps represents meaning and intent. Just like our color blue/blau/azul/blu. Not faulting them for asking the question but I agree that the results do not undermine the capability of the technology. In fact it just helps highlight the constraints and need for education.
- onlyrealcuzzo 2y ago> Like, that has nothing to do with their intelligence. Because they don't have intelligence. If they did, they could count the letters in strawberry.
- TwentyPosts 2y agoPeople have been over this. If you believe this, you don't understand how LLMs work. They fundamentally perceive the world in terms of tokens, not "letters".
- antonvs 2y ago> If you believe this, you don't understand how LLMs work. Nor do they understand how intelligence works. Humans don't read text a letter at a time. We're capable of deconstructing words into individual letters, but based on the evidence that's essentially a separate "algorithm". Multi-model systems could certainly be designed to do that, but just like the human brain, it's unlikely to ever make sense for a text comprehension and generation model to work at the level of individual letters.
- fmbb 2y ago> that has nothing to do with their intelligence. Of course. Because these models have no intelligence. Everyone who believes they do seem to believe intelligence derives from being able to use language, however, and not being able to tell how many times the letter r is in the word strawberry is a very low bar to not pass.
- roywiggins 2y agoAn LLM trained on single letter tokens would be able to, it just would be much more laborious to train.
- wruza 2y agoWhy would it be able to?
- roywiggins 2y agoIf you give LLMs the letters one a time they often count them just fine, though Claude at least seems to need to keep a running count to get it right: "How many R letters are in the following? Keep a running count. s t r a w b e r r y" They are terrible at counting letters in words because they rarely see them spelled out. An LLM trained one byte at a time would always see every character of every word and would have a much easier time of it. An LLM is essentially learning a new language without a dictionary, of course it's pretty bad at spelling. The tokenization obfuscates the spelling not entirely unlike how verbal language doesn't always illuminate spelling.
- wruza 2y agoMay the effect you see, when you spell it out, be not a result of “seeing” tokens, but a result of the fact that a model learned – at a higher level – how lists in text can be summarized, summed up, filtered and counted? Iow, what makes you think that it’s exactly letter-tokens that help it and not the high-level concept of spelling things out itself?
- furyofantares 2y agoIt's not very interesting when they fail at it, but it will be interesting if they get good at it. Also there are some cases where regular people will stumble into it being awful at this without any understanding why (like asking it to help them with their wordle game.)
- Auracle 2y agoI suppose what models should have are some instructions of things they aren’t good at and will need to break out into python code or what have you. Humans have an intuition for this - I have basic knowledge of when I need to write something down or use a calculator. LLMs don’t have intuition (yet - though I suppose one could use a smaller model for that), so explicit instructions would work for now.