4 ms·
I'm curious when you say LLM convert words in to vectors, is it similar to vectors in physics, with one part being a magnitude and the other part being a direct
by iamerroragent 3y ago
I'm curious when you say LLM convert words in to vectors, is it similar to vectors in physics, with one part being a magnitude and the other part being a direction?
- jesse_cureton 3y agoIt's similar - basically here a vector/tensor is an array of magnitudes across N dimensions. Whereas in (undergrad-level) physics you might have a 3- or 4-D vector for spatial dimensions and time, here the LLMs are embedding sequences of tokens into N-dimensional space where N is much much larger. There's a pattern called "embedding search" - you precalculate a set of embeddings for a corpus of text. Then to do a search, you calculate embeddings for your search string. Then you can find the closest vector in that N-dimensional space, which finds you the semantically closest neighbor from the original corpus. For an embedding search - the OpenAI Embeddings API gives you a ~1500 dimension output vector. When a LLM is working with input text as a vector, I am not sure what the tokenizer is actually feeding into the model. Hopefully someone else can chime in!
- iamerroragent 3y agoThis is really elucidating. Thank you and your time for writing that out.
- TeMPOraL 3y agoThe magic is really in how absurdly high-dimensional those vectors are. The dimensionality of the latent space in current breed of LLMs is, IIRC, on the order of hundred thousand dimensions. Now, even if all the model does is 1) turn your prompt into high-dimensional vectors, 2) run an adjacency search to find the vectors in the latent space nearest to your vector, and 3) translate those vectors back to tokens, it's more than enough for it to work with concepts. Again, we're talking 1000 to 100 000 dimensional vectors here. Any kind of semantic similarity between words you can think of (tree - green - grass, tree - tall - skyscraper, tree - data structure, tree - files, etc.) can fit in there - the relevant words (tokens) will be close together along some dimensions. So, if you pick a vector somewhere in the latent space, and look around (in hundred thousand dimensions) for its nearest neighbors, the group of points you'd be looking at is, IMHO, a concept in its raw form.
- GuB-42 3y agoVectors in this case are more like an array of numbers, often in the thousands, where each number represent a concept associated with the word. So for example, "elephant" will have have values in the cells representing "animal", "grey", "big", "noun", etc... but low values in "verb", "abstract", "flight", etc... The meaning of these cells is usually not explicit, it is something that emerges from the model learning process, but since these models are usually trained on human language, you often find human concepts in there. And indeed, mathematically, they are the same as vector in physics, just with a lot more dimensions. And indeed there is a concept of magnitude (norm) and direction. Some processes only use the "direction", throwing away the magnitude through a process called normalization, some don't care, and some actually make use of the magnitude.