6 ms·
What's the issue with character-level tokenization(I assume this would be much better at count-the-letter tasks)? The article mentions it as an option but doesn
by IncreasePosts 2y ago
What's the issue with character-level tokenization(I assume this would be much better at count-the-letter tasks)? The article mentions it as an option but doesn't talk about why subword tokenization is preferred by most of the big LLMs out there.
- SEGyges 2y agotokens are on average four characters and the number of residual streams (and therefore RAM) the LLM allocates to a given sequence is proportionate to the number of units of input. the flops is proportionate to their square in the attention calculation. you can hypothetically try to ameliorate this by other means, but if you just naively drop from tokenization to character or byte level models this is what goes wrong
- p1esk 2y ago4x seq length expansion doesn’t sound that bad.
- lechatonnoir 2y agoI mean, it's not completely fatal, but it means an approximately 16x increase in runtime cost, if I'm not mistaken. That's probably not worth trying to solve letter counting in most applications.
- SEGyges 2y agoit is not necessarily 16x if you, e.g., decrease model width by a factor of 4 or so also, but yeah naively the RAM and FLOPs scale up by n^2
- Centigonal 2y agoI think it has to do with both performance (smaller tokens means more tokens per sentence read and more runs per sentence generated) and with how embeddings work. You need a token for "dog" and a token for "puppy" to represent the relationship between the two as a dimension in latent space.
- stephantul 2y agoUsing subwords makes your sequences shorter, which makes them cost less. Besides that, for alphabetic languages, there exists almost no relation between form and meaning. I.e.: “ring” and “wing” differ by one letter but have no real common meaning. By picking the character or byte as your choice of representation, the model basically has to learn to distinguish ring and wing in context. This is a lot of work! So, while working on the character or byte level saves you some embeddings and thus makes your model smaller, it puts all of the work of distinguishing similar sequences with divergent meanings on the model itself, which means you need a larger model. By having subwords, a part of this distinguishing work already has been done by the vocabulary itself. As the article points out, this sometimes fails.
- bunderbunder 2y agoI suspect that the holy grail here is figuring out how to break the input into a sequence of morphemes and non-morpheme lexical units.
- thaumasiotes 2y agoWhat do you mean by non-morpheme lexical units? Syntactic particles, units too small to be morphemes? Lexical items that contain multiple morphemes? In either case, isn't this something we already do well?
- bunderbunder 2y agoPunctuation, for example. And no, at least for the languages with which I'm familiar SOTA tokenizers tend to only capture the easy cases. For example, the GPT-4 tokenizer breaks the first sentence of your post like so: What/ do/ you/ mean/ by/ non/-m/orp/heme/ lexical/ units/? Notice how "morpheme" gets broken into three tokens, and none of them matches "morpheme"'s two morphemes. "Lexical" and "units" are each a single token, when they have three and two morphemes respectively. Or in French, the word "cafetière" gets chopped willy-nilly into "c/afet/ière". The canonical breakdown is "cafe/t/ière".
- 2y ago
- cma 2y agoContext length performance and memory scales N^2. Smaller tokens mean worse scaling, up to a point.