3 ms·
You might be thinking of finetuning, you should checkout alpaca-lora documentation. Embeddings are just floating point array representation of the underlying t
by JimmyRuska 3y ago
You might be thinking of finetuning, you should checkout alpaca-lora documentation.
Embeddings are just floating point array representation of the underlying text, where 'tokens' that are often used together are numbers close in proximity. You can generate vector representation of any text very easily using the openai apis, or frameworks like langchain https://github.com/openai/openai-cookbook/blob/main/examples/Get_embeddings.ipynb https://github.com/openai/openai-cookbook/blob/main/examples...
Similarity means your text gets turned into an array of numbers and the vector database finds the closest matches, potentially across a huge database of text documents. Vector databases are often used in conjunction to LLMs, for example to pull out all the snippets of relevant text, then feed it within the prompt with your question.