3 ms·
The most basic of embedding (that I can think of) is one where the number of dimensions corresponds to the number of unique characters in your lexicon. If there
by kreeben 3y ago
The most basic of embedding (that I can think of) is one where the number of dimensions corresponds to the number of unique characters in your lexicon. If there is a number in one of the components of your embedding that is greater than "0", then you know what characters they are. This embedding does not encode the order of the characters, though. They are just a "bag of characters". If you were to then also encode the order of the characters in, say, yet another embedding, you could use those two embeddings to recreate the original word.
Combine the two embeddings into a new vector space and BAM, you've invented "embedding2word".
- jorlow 3y agoGpt (and many others) just add these embeddings together in the model, so you could do that and have one vector that encodes both things together
- TeMPOraL 3y agoThat seems suboptimal, for the same reason LLMs are trained on tokens, and not characters. Tokens seem like a much better "unit of meaning" than characters. Curiously, this also applies to humans: we learn words first, spelling later; we almost always think in terms of words, subwords and phrases, and we're really fast at it - while anything to do with spelling seems to demand much more focus, and some degree of conscious hand-holding. As for encoding sequence, I'm curious how that happens and need to find relevant papers, but I imagine there might be some "n-gram dimensions" in the vector, as in one value for "I'm a first token in a sequence", one value for "I'm a second token in a sequence", etc., which would encode the "occurs before"/"occurs after" relationships using few dimensions, leaving the rest for more interesting relationships.