3 ms·
My understanding is limited, but I'm guessing it's to do with the algebraic properties of the embedding vectors. I'm familiar with embeddings that you can add a
by timmb 3y ago
My understanding is limited, but I'm guessing it's to do with the algebraic properties of the embedding vectors. I'm familiar with embeddings that you can add and subtract, which may reveal concepts existing as linear directions within the embedding space.
Here, they're talking about multiplying, dividing and permuting vectors. Multiplying combines concepts, adding creates a superposition.
They also mention randomly selecting embeddings for the concepts in mind. So my guess is that instead of one-hot encoding the classifier, they instead use random encodings on the output, and are working to give those encodings desirable properties.
I would also hazard a guess that the random vectors they choose are close to zero in most components.
- randcraw 3y agoOr a high-dimensional word-to-vec that can represent compound / complex objects?
- chombier 3y agoThis is my understanding as well. There a somewhat accessible introductory video I found useful [0]. The algebra makes it possible to encode sets, key/value associations and sequences to build a knowledge base, and the dot product provides a similarity measure for querying the base. IIUC the key is that for large space dimensions, any two random vectors (say with uniform distribution over {-1, +1}^d) are almost guaranteed to be near- orthogonal. This makes it easy to add new items to the base (by sampling a new random vector and updating the base using algebraic operations with other items), yet the amount of noise introduced by near-orthogonality remains controlled and can be filtered out to keep the algebraic structure working as the base grows. Honestly it seems a bit too good to be true, I'd be very interested to see what are the tradeoffs in practice. [0] https://www.youtube.com/watch?v=oB_mHCurNCI https://www.youtube.com/watch?v=oB_mHCurNCI