4 ms·
Does it mean that discrete representation is enough for capturing high-level semantic info?
by against_entropy 2y ago
Does it mean that discrete representation is enough for capturing high-level semantic info?
- nico 2y agoNot sure about the LLM context But conceptually, everything we put in a computer is represented in discrete binary sequences to be processed and stored. So from that perspective, it wouldn’t be too far fetched
- against_entropy 2y agoThis is definitely true about the 0-1 representation of diodes. I think what marqo demonstrated may indicate that data granularity is not the most essential issue.
- jn2clark 2y agoWith binary representations you still get 2^D possible configurations so its entirely possible from a representation perspective. The main issue (I think at least) is around determining the similarity. Hamming distance gives an output space of D possible scores. As mentioned in the article, going to 0/1 with cosine gives better granularity as it now penalizes embeddings if they have differing amounts of positive elements in the embedding (i.e. living on different hyper-spheres). It is probably well suited to retrieval where there is a 1:1 correspondence for query-document but if the degeneracy of queries is large then there could be issues discriminating between similar documents. Regimes of binary and (small) dense embeddings could be quite good. I expect a lot more innovation in this space.
- amitport 2y agoFloat32 is also a discrete representation.
- against_entropy 2y agoTrue. I was thinking about accuracy and granularity issues