6 ms·
Maybe I'm really stupid... but perhaps if we want really intelligent models we need to stop tokenizing at all? We're literally limiting what a model can see and
by azeirah 2y ago
Maybe I'm really stupid... but perhaps if we want really intelligent models we need to stop tokenizing at all? We're literally limiting what a model can see and how it percieves the world by limiting the structure of the information streams that come into the model from the very beginning.
I know working with raw bits or bytes is slower, but it should be relatively cheap and easy to at least falsify this hypothesis that many huge issues might be due to tokenization problems but... yeah.
Surprised I don't see more research into radicaly different tokenization.
- amelius 2y agoPerhaps we can even do away with transformers and use a fully connected network. We can always prune the model later ...
- cschep 2y agoHow would we train it? Don't we need it to understand the heaps and heaps of data we already have "tokenized" e.g. the internet? Written words for humans? Genuinely curious how we could approach it differently?
- viraptor 2y agoThat's not what tokenized means here. Parent is asking to provide the model with separate characters rather than tokens, i.e. groups of characters.
- skylerwiernik 2y agoCouldn't we just make every human readable character a token? OpenAI's tokenizer makes "chess" "ch" and "ess". We could just make it into "c" "h" "e" "s" "s"
- taeric 2y agoThis is just more tokens? And probably requires the model to learn about common groups. Consider, "ess" makes sense to see as a group. "Wss" does not. That is, the groups are encoding something the model doesn't have to learn. This is not much astray from "sight words" we teach kids.
- TZubiri 2y agoThis is just more tokens? Yup. Just let the actual ML git gud
- taeric 2y agoSo, put differently, this is just more expensive?
- TZubiri 2y agoExpensive in terms of computationally expensive, time expensive, and yes cost expensive. Worth noting that the relationship between characters to token ratio is probably quadratic or cubic or some other polynomial. So the difference in terms of computational difficulty is probably huge when compared to a character per token.
- Hendrikto 2y agoNo, actually much fewer tokens. 256 tokens cover all bytes. See the ByT5 paper: https://arxiv.org/abs/2105.13626 https://arxiv.org/abs/2105.13626
- tchalla 2y agoaka Character Language Models which have existed for a while now.
- cco 2y agoWe can, tokenization is literally just to maximize resources and provide as much "space" as possible in the context window. There is no advantage to tokenization, it just helps solve limitations in context windows and training.
- TZubiri 2y agoI like this explanation
- aithrowawaycomm 2y agoFWIW I think most of the "tokenization problems" are in fact reasoning problems being falsely blamed on a minor technical thing when the issue is much more profound. E.g. I still see people claiming that LLMs are bad at basic counting because of tokenization, but the same LLM counts perfectly well if you use chain-of-thought prompting. So it can't be explained by tokenization! The problem is reasoning: the LLM needs a human to tell it that a counting problem can be accurately solved if they go step-by-step. Without this assistance the LLM is likely to simply guess.
- ipsum2 2y agoThe more obvious alternative is that CoT is making up for the deficiencies in tokenization, which I believe is the case.
- aithrowawaycomm 2y agoI think the more obvious explanation has to do with computational complexity: counting is an O(n) problem, but transformer LLMs can’t solve O(n) problems unless you use CoT prompting: https://arxiv.org/abs/2310.07923 https://arxiv.org/abs/2310.07923
- ipsum2 2y agoWhat you're saying is an explanation what I said, but I agree with you ;)
- aithrowawaycomm 2y agoNo, it's a rebuttal of what you said: CoT is not making up for a deficiency in tokenization, it's making up for a deficiency in transformers themselves. These complexity results have nothing to do with tokenization, or even LLMs, it is about the complexity class of problems that can be solved by transformers.
- ipsum2 2y ago
- jncfhnb 2y agoThere’s a reason human brains have dedicated language handling. Tokenization is likely a solid strategy. The real thing here is that language is not a good way to encode all forms of knowledge
- joquarky 2y agoIt's not even possible to encode all forms of knowledge.
- shaky-carrousel 2y agoI know a joke where half of the joke is whistling and half gesturing, and the punchline is whistling. The wording is basically just to say who the players are.
- layer8 2y agoGoing from tokens to bytes explodes the model size. I can’t find the reference at the moment, but reducing the average token size induces a corresponding quadratic increase in the width (size of each layer) of the model. This doesn’t just affect inference speed, but also training speed.
- famouswaffles 2y agoTokenization is not strictly speaking necessary (you can train on bytes). What it is is really really efficient. Scaling is a challenge as is, bytes would just blow that up.
- ATMLOTTOBEER 2y agoI tend to agree with you. Your post reminded me of https://gwern.net/aunn https://gwern.net/aunn
- gwern 2y agoOne neat thing about the AUNN idea is that when you operate at the function level, you get sort of a neural net version of lazy evaluation; in this case, because you train at arbitrary indices in arbitrary datasets you define, you can do whatever you want with tokenization (as long as you keep it consistent and don't retrain the same index with different values). You can format your data in any way you want, as many times as you want, because you don't have to train on 'the whole thing', any more than you have to evaluate a whole data structure in Haskell; you can just pull the first _n_ elements of an infinite list, and that's fine. So there is a natural way to not just use a minimal bit or byte level tokenization, but every tokenization simultaneously: simply define your dataset to be a bunch of datapoints which are 'start-of-data token, then the byte encoding of a datapoint followed by the BPE encoding of that followed by the WordPiece encoding followed by ... until the end-of-data token'. You need not actually store any of this on disk, you can compute it on the fly. So you can start by training only on the byte encoded parts, and then gradually switch to training only on the BPE indices, and then gradually switch to the WordPiece, and so on over the course of training. At no point do you need to change the tokenization or tokenizer (as far as the AUNN knows) and you can always switch back and forth or introduce new vocabularies on the fly, or whatever you want. (This means you can do many crazy things if you want. You could turn all documents into screenshots or PDFs, and feed in image tokens once in a while. Or why not video narrations? All it does is take up virtual indices, you don't have to ever train on them...)
- numpad0 2y agohot take: LLM tokens is kanji for AI, and just like kanji it works okay sometimes but fails miserably for the task of accurately representating English
- School-Cotton 2y agoWhy couldn’t Chinese characters accurately represent English? Japanese and Korean aren’t related to Chinese and still were written with Chinese characters (still are in the case of Japanese). If England had been in the Chinese sphere of influence rather than the Roman one, English would presumably be written with Chinese characters too. The fact that it used an alphabet instead is a historical accident, not due to any grammatical property of the language.
- stickfigure 2y agoIf I read you correctly, you're saying "the fact that the residents of England speak English instead of Chinese is a historical accident" and maybe you're right. But the residents of England do in fact speak English, and English is a phonetic language, so there's an inherent impedance mismatch between Chinese characters and English language. I can make up words in English and write them down which don't necessarily have Chinese written equivalents (and probably, vice-versa?).
- School-Cotton 2y ago> If I read you correctly, you're saying "the fact that the residents of England speak English instead of Chinese is a historical accident" and maybe you're right. That’s not what I mean at all. I mean even if spoken English were exactly the same as it is now, it could have been written with Chinese characters, and indeed would have been if England had been in the Chinese sphere of cultural influence when literacy developed there. > English is a phonetic language What does it mean to be a “phonetic language”? In what sense is English “more phonetic” than the Chinese languages? > I can make up words in English and write them down which don’t necessarily have Chinese written equivalents Of course. But if English were written with Chinese characters people would eventually agree on characters to write those words with, just like they did with all the native Japanese words that didn’t have Chinese equivalents but are nevertheless written with kanji. Here is a famous article about how a Chinese-like writing system would work for English: https://www.zompist.com/yingzi/yingzi.htm https://www.zompist.com/yingzi/yingzi.htm
- matusp 2y agoI have seen a bunch of tokenization papers with various ideas but their results are mostly meh. I personally don't see anything principally wrong with current approaches. Having discrete symbols is how natural language works, and this might be an okayish approximation.
- malthaus 2y agohttps://youtu.be/zduSFxRajkE https://youtu.be/zduSFxRajkE karpathy agrees with you, here he is hating on tokenizers while re-building them for 2h
- blixt 2y agoI think it's infeasible to train on bytes unfortunately, but yeah it also seems very wrong to use a handwritten and ultimately human version of tokens (if you take a look at the tokenizers out there you'll find fun things like regular expressions to change what is tokenized based on anecdotal evidence). I keep thinking that if we can turn images into tokens, and we can turn audio into tokens, then surely we can create a set of tokens where the tokens are the model's own chosen representation for semantic (multimodal) meaning, and then decode those tokens back to text[1]. Obviously a big downside would be that the model can no longer 1:1 quote all text it's seen since the encoded tokens would need to be decoded back to text (which would be lossy). [1] From what I could gather, this is exactly what OpenAI did with images in their gpt-4o report, check out "Explorations of capabilities": https://openai.com/index/hello-gpt-4o/ https://openai.com/index/hello-gpt-4o/
- PittleyDunkin 2y agoA byte is itself sort of a token. So is a bit. It makes more sense to use more tokenizers in parallel than it does to try and invent an entirely new way of seeing the world. Anyway humans have to tokenize, too. We don't perceive the world as a continuous blob either.
- samatman 2y agoI would say that "humans have to tokenize" is almost precisely the opposite of how human intelligence works. We build layered, non-nested gestalts out of real time analog inputs. As a small example, the meaning of a sentence said with the same precise rhythm and intonation can be meaningfully changed by a gesture made while saying it. That can't be tokenized, and that isn't what's happening.
- PittleyDunkin 2y agoWhat is a gestalt if not a token (or a token representing collections of other tokens)? It seems more reasonable (to me) to conclude that we have multiple contradictory tokenizers that we select from rather than to reject the concept entirely. > That can't be tokenized Oh ye of little imagination.
- Anotheroneagain 2y agoI think on the contrary, the more you can restrict it to reasonable inputs/outputs, the less powerful LLM you are going to need.
- ajkjk 2y agoThis is probably unnecessary, but: I wish you wouldn't use the word "stupid" there. Even if you didn't mean anything by it personally, it might reinforce in an insecure reader the idea that, if one can't speak intelligently about some complex and abstruse subject that other people know about, there's something wrong with them, like they're "stupid" in some essential way. When in fact they would just be "ignorant" (of this particular subject). To be able to formulate those questions at all is clearly indicative of great intelligence.