4 ms·
Is it because with transformers and LLMs the embeddings are more powerful than they have been before (they can seem to understand more)?
by quickthrower2 3y ago
Is it because with transformers and LLMs the embeddings are more powerful than they have been before (they can seem to understand more)?
- jaggederest 3y agoBroadly, as I understand it, the way the embeddings are generated "needs" LLMs to produce the superior results you'd see with e.g. OpenAI's text-embedding-ada-002 - we've been working with vectors at least as long as I've been programming, but until recently they didn't mean anything particular. LLMs let the vectors have semantic meaning and relate similar conceptual text, instead of (as in the article) using gzip to generate vectors where the similarity is entirely textual, i.e. similar words and passages and characters, even if they have different meanings. So ultimately they're not at all new (embedding data into a hamming space to compare it to other data), but the underlying meaning is much more useful in a world where you can peek into the internal state of a LLM to generate them.