3 ms·
Embeddings are there to make continuous-space identities so you can run them through a differentiable model. Without this the tokens (no matter your granularity
by thatjoeoverthr 1mo ago
Embeddings are there to make continuous-space identities so you can run them through a differentiable model. Without this the tokens (no matter your granularity) are pure surrogate identities, and you can’t run a gradient through them. You also hit the curse of dimensionality hard because the model can’t perceive similarity. “Cat” and “kitten” for example are simply different atoms of text, but with embeddings, you can leverage what you learned about “cat” when you encounter “kitten”. Look at “A Neural Probabalistic Language Model” (Bengio, 2003).
You can actually rig up an embedding variant of a Markov chain with just a few tokens of context, and no position coding, transformers, attention, none of it, and only minutes of training time. As long as you have the embedding lookup table trainable it will do some neat stuff.