3 ms·
Pardon the n00b question, but... How does this relate to vectors? It was my understanding that the tokens were vectors and this seems to show them as an intege
by throwaway2016a 3y ago
Pardon the n00b question, but...
How does this relate to vectors? It was my understanding that the tokens were vectors and this seems to show them as an integer.
It's probably a really obvious question to anyone who knows AI but I figured if I have it someone else does too.
- binarymax 3y agoVery basic overview: A token is assigned a number, that number gets passed into the encoder model with other token numbers, and the encoder model transforms those number sequences into embeddings (vectors)
- nighthawk454 3y agoThe tokens are an integer. The first layer of the model is an 'embedding', which is essentially a giant lookup table. So if a string gets tokenized to Token #3, that means get the vector in row 3 of the embedding table. (Those vectors are learned during model training.) More completely, you can think of the integers as being implicitly a one-hot vector encoding. So say you have a vocab size of 20,000 and you want Token #3. The one-hot vector would be a 20,000 length vector of zeros with a one in position 3. This vector is then multiplied against the embedding table/matrix. Although in practice this is equivalent to just selecting one row directly, so it's implemented as such and there's no reason to explicitly make the large one-hot vectors.
- CuriousSkeptic 3y agoKind of refreshing to see this perspective on lookup vs matrix multiplication, specially with the bias towards the latter as more natural. Is there some reference table somewhere mapping more code idioms like this to equivalent nn representations?
- nighthawk454 3y agoHmm not that I know of, but that would be neat! A lot of the frameworks and model code treat this sort of thing as ‘implementation details’. Which is disappointing because I think it adds perspective and intuition. One other example would be how multi-head attention is implemented with a single matrix. You don’t actually create matrices for each of the N ‘heads’ separately. It’s a logical distinction
- quickthrower2 3y agoAndrej covers this in https://github.com/karpathy/nn-zero-to-hero https://github.com/karpathy/nn-zero-to-hero. He explains things in multiple ways, both the matrix multiplications as well as the "programmer's" way of thinking of it - i.e. the lookups. The downside is it takes a while to get through those lectures. I would say for each 1 hour you need another 10 to looks stuff up and practice, unless you are fresh out of calculus and linear algebra classes. Other idioms I can think of, in my words: Softmax = take the maximum (but in a differentiable way) tanh/sigmoid/relu = a switch. "activation" cross entropy loss = average(-log(probability you gave to the right answer)). Averaged over the current batch you are training on for this step. (Sorry that is still quite mathy).
- nighthawk454 3y agoThanks! I'm gonna check that out
- z3c0 3y agoThe integers represent a position within a vector of "all known tokens". Typically, following a simple bag-of-words approach, each position in the vector would be toggled to 1 or 0 based on the presence of a token in a given document. Since most vectors would be almost completely zeroed, the simpler way to represent these vectors is through a list of positions in the now abstracted vector, aka a sparse vector, ie a list of integers. In the case of more advanced language models like LLMs, a given token can be paired with many other features of the token (such as dependencies or parts-of-speech) to make an integer represent one of many permutations on the same word based on its usage.
- RC_ITR 3y agoAfter training, tokens are vectors, but the number of unique vectors is limited by your vocabulary size (i.e. Should 'The' get a vector or should 'Th' and 'e' each get their own vector?). This step is deciding which clusters of letters (or whatever) get a vector and then giving them a scalar unique ID for conveniences' sake. The training then determines what that vector actually is.
- Ambix 3y agoTokens are just integer numbers, showing their position in the big vocabulary - it's that simple :) And vocabulary is just an array / vector / list - it depends which programming language you use, each has each own terminology for that data structure. For example LLaMA vocabulary has 32,000 tokens.