3 ms·
You need 3 things : a query, a corpus of text to be searched against and a language model that can map text to vectors ("compute the embedding"). Morally, when
by cpa 4y ago
You need 3 things : a query, a corpus of text to be searched against and a language model that can map text to vectors ("compute the embedding").
Morally, when you get a query, you compute its embedding (through an API call) and return parts of your corpus whose embedding are cosine-close to the query.
In order to do that efficiently, you indeed have to pre-compute all the embeddings of your corpus beforehand and store them in your database.
- xrd 4y agoThank you. So, this does imply that you need to know the target of your embeddings in advance, right? If I want to know if my text has something like "dancing" and it does have "tango" inside it, why wouldn't I just generate a list of synonyms and then for-each run those queries on my text and then aggregate them myself? I can download a few synonyms databases here: https://stackoverflow.com/questions/5618304/looking-for-thesaurus-data https://stackoverflow.com/questions/5618304/looking-for-thes... Is the value here that OpenAI can get me a better list of synonyms than I could do on my own? If OpenAI were better at generating this list of synonyms, especially with more current data (I need to search for a concept like "fuzzy-text" and want text with "embeddings" to be a positive match!) that would be valuable. It feels like OpenAI will probably be faster to update their model with current data from the Internet than those synonym lists linked above. Having said that, one of the criticisms of ChatGPT is that it does not have great knowledge of more recent events, right? Don't ask it about the Ukraine war unless you want a completely fabricated result.
- foooobaba 4y agoYou don’t need to know the target queries, if you compute embeddings of your entries and your query you just find which embeddings are closest to your query embedding. The advantage over using synonyms is that the embedding is meant to encode the meaning of the content such that similar embeddings represent similar meaning and you won’t need to deal with the combinatorial explosion of all the different ways you can say the same thing with different words (it can also work for other content, like images, or multi language if ur network is trained for it).
- cpa 4y agoThat's the value prop of large language models (and here, of OpenAI's LLM): because it's been trained on some "somehow sufficiently large" corpus of data, it has internalized a lot of real world concepts, in a "superficial yet oddly good enough" way. Good enough for it to have embeddings for "dancing" and "tango" that will be fairly close. And if you really need to, you can also fine-tune your LLM or do few shot learning to further customise your embeddings to your dataset.
- foooobaba 4y agoBut yes, if you ask OpenAI to predict next set of tokens (which is how chat works), it won’t be up to date with latest information. But if you’re using it for embeddings this is less of a problem since language itself doesn’t evolve as quickly, and using embeddings is all about encoding the meaning of text, which is likely not going to change so much - but not to say it can’t for example the definition of “transformer” pre 2017 is probably not referring to the “transformer architecture”.
- thanatropism 4y agoLook for the Python package "sentence-transformers" and give it a spin.