3 ms·
I was a bit surprised by this: > For example, suppose you prompt with a sentence "What is the capital of California: ", it would take ten forward pass iteratio
by sakex 3y ago
I was a bit surprised by this:
> For example, suppose you prompt with a sentence "What is the capital of California: ", it would take ten forward pass iterations to get back the full response of ["S", "a", "c", "r", “a”, "m", "e", "n", "t", "o"]
From the content I have been reading trying to understand LLMs, I thought that the output was a token and not a string of chars. What am I missing here?
- fnordpiglet 3y agoI suspect they’re simplifying things
- toxik 3y agoThe tokens are chosen by a method called BPE, and there are single letter tokens. IOW you can encode the same text many ways. That said, this is probably just shown like this for illustrative purposes.
- gattilorenz 3y agoIt's mentioned in the article: "This example simplifies things a little bit because in actuality tokens do not map 1:1 to ASCII characters (a popular token encoding technique is Byte-Pair Encoding which is beyond the scope of this blog post), but the iterative nature of generation is the same regardless of how you tokenize your sequences." It's an extremization that still is true for character-based models.
- reqo 3y agoYou need to use a decoder to convert the response to a sentence. Example: https://huggingface.co/docs/tokenizers/api/decoders https://huggingface.co/docs/tokenizers/api/decoders
- deleted 3y ago[deleted]
- llm_nerd 3y agoA number of submissions lately have "simplified" and presented every character as a token. This has only confused many readers. LLMs use a vocabulary of statistically chosen tokens. GPT 3 vocab, for instance, splits Sacramento into three tokens- Sac - 38318 rament - 15141 o - 78 There is a rule of thumb that about every four letters in English text becomes a token but that's just the average. California is a single token (25284). As is Canada (17940). And so on.
- jacquesm 3y agoIs that because four letters happen to pack into a word on most architectures or is there some other underlying reason? Also Sac and rament together already form a word so it's kind of logical have the final 'o' as a separate token.
- daveguy 3y agoThe underlying reason is closer to your second observation. Statistically chosen tokens are most likely to be reusable and composable. Much less to do with hardware architecture. Token vocabulary happens before any quantization where processor word size may play a role.
- llm_nerd 3y agoIt's just a function of finding the most efficient solution for encoding a corpus of text into a given vocabulary size. Using something like SentencePiece you can define how big you want your vocabulary to be (the number of discrete tokens), and it will find the best solution of subwords/characters in a sample set. In the case of GPT-3 the vocabulary has a size of 50,257 tokens. GPT-4 increases that past 100k (see https://gist.github.com/s-macke/ae83f6afb89794350f8d9a1ad8a09193 https://gist.github.com/s-macke/ae83f6afb89794350f8d9a1ad8a0...). It's very similar to compression algos, really. Find recurring sets of characters.