3 ms·
I still haven't figured what a vector DB is, beyond something something AI.
by NohatCoder 3y ago
I still haven't figured what a vector DB is, beyond something something AI.
- commandlinefan 3y agoI was hoping the article would say and then when it didn't I was hoping the comments would, but so far no luck...
- kelseyfrog 3y agoVectorDBs let you retrieve documents that have textual similarity. They allow you sort results by the cosine similarity[1] between vectors. The idea is that you can attach a vector to documents in the database and then when you pass a vector in the query and get back documents that most match the query vector. The function that creates these vectors(string -> vector) is called an embedding and is constructed in such a way that "semantically similar" strings have vectors that are close together. It's not a very complicated idea, but complicated and powerful are orthogonal concepts. They are useful in AI(LLMs) when you would like to include documents in your prompt that are relevant to your instruction. The best way to describe this is by example. Imagine your query is "What is the capital of France?" Rather than requiring your LLM to encounter this fact during its training, you can embed the question("What is the capital of France?" and retrieve documents (say you've indexed all of Wikipedia in your vector db) and return some snippets from articles that include this information(context). You then pass the the prompt+context to an LLM and given that it now has relevant information, it can answer the question. You can also imagine that it's much easier to update a vector db with new information than it is to retrain a model to ingest new facts. 1. https://en.wikipedia.org/wiki/Cosine_similarity https://en.wikipedia.org/wiki/Cosine_similarity
- danielmarkbruce 3y agoImagine: you need a data structure which allows you to store vectors in memory and then say "here is a vector A, give me the 10 vectors which are closest to the same direction of this vector A, in n dimensional space". One can imagine there is some optimal way to lay out the data in memory such that it would be relatively quick to do that. One can imagine a naive way to do it which probably wouldn't be fast. A vector DB does the first thing - lays out the data in a way which then enables it to do that fast. And then it does all the other stuff a DB does - persisting to disk (which means data needs to be laid out in a sensible way on disk too), handling multiple queries, updates, and the 50 other complicated things databases tend to do. Users generally want to do other operations on vector as well so a vector DB does those too. For a small number of vectors you can build a vector DB yourself. Write a list of vectors to a file, load them into memory in no particular order, then for your "n closest" function, just iterate through the list calculating the difference in direction one by one and keeping the top n. Your simple system will work just fine for a toy demo.
- Salgat 3y agoI think the better question is, what the hell is a vector in the context of normal business logic? Sure, a word embedding vector makes sense because it's all an abstraction anyways, but if I have an "Employee" table with name, address, position, etc columns, how does that translate into a vector?
- danielmarkbruce 3y agoOh, it doesn't. Vectors in this context are used to semantically represent unstructured text, not structured data that you'll find in a table of a sql database (except a big fat text field). Here: https://chat.openai.com/share/9e557a90-e127-4654-9271-7c51fdb67d05 https://chat.openai.com/share/9e557a90-e127-4654-9271-7c51fd...
- nazka 3y agoWow you are so right. I just typed “best vector database” on Google and I have 4 ads with other results all talking about “vector database something AI”.