3 ms·
The article is referring to the problem of having a limited context length in LLMs. That is you can only pass X tokens in the prompt. For example, let’s say yo
by dcastm 3y ago
The article is referring to the problem of having a limited context length in LLMs. That is you can only pass X tokens in the prompt.
For example, let’s say you have a prompt that lets you answer questions about a book. If the book is long enough, you won’t be able to include it as is in the prompt, so you have to figure out what are the most relevant passages you must include to answer a given question. What you usually do is find the passages that are the most semantically similar to your question.
Chunk vectors are the vectorized passages of the book (i.e., a numerical vector that represents a passage), and the prompt vector is usually the vectorized question.
To find the most similar vectors you need a distance measure, cosine similarity being the most popular.
The output of finding the most similar vectors is the vectors + it’s metadata (chunk, page, chapter, etc)
- Roark66 3y agoThank you for explaining it. I was aware of the problem of too short context length. I'd love to be able to pass an entire programming project with my prompt for example (or a book in your example). I think I understand how it works now for many kinds of prompts where the information to be extracted is contained in one(or more) of the chunks of much bigger whole. I'm not sure about prompts where it is required to "understand" entire input to answer properly. For example summarising a book. Although even with this vector search could perhaps help by looking for things not near "please provide a summary", but certain hand crafted values such as "important to the plot" etc. I guess I need to do some experimenting with It. I found some open source alternatives to (not at all)OpenAI in form of "Sentence Transformers" to create embeddings. However, what would be really neat is to have a large open source dataset of embeddings already created on some general purpose collection of texts, to try searches etc.
- summarity 3y agoI built something like this at findsight.ai and gave a talk about it (including a discussion on promoting and chunking) here: https://youtu.be/elNrRU12xRc https://youtu.be/elNrRU12xRc There are also open datasets of this, eg https://huggingface.co/datasets/kannada_news https://huggingface.co/datasets/kannada_news for news, or https://sites.google.com/eng.ucsd.edu/ucsdbookgraph/reviews?authuser=0 https://sites.google.com/eng.ucsd.edu/ucsdbookgraph/reviews?...
- hobs 3y agoIts pretty trivial to calculate the embeddings if you have a reasonable GPU with CUDA support, even for decent volumes of data, the problem is that each different model would produce different embeddings and as soon as a new one is released all your previous datasets aren't that useful. Also worth mentioning that vector databases are really useful in Resource Augmented Generation - aka find an answer from an existing corpus and utilize it to supplement the LLM.
- dcastm 3y agoCheck Cohere's embeddings of Wikipedia: https://txt.cohere.com/embedding-archives-wikipedia/ https://txt.cohere.com/embedding-archives-wikipedia/