25 ms·
Glad to see more work on pgvector but why test on such small datasets on a large memory machine? The big ann datasets have 1B points and are much more interesti
by moab 3y ago
Glad to see more work on pgvector but why test on such small datasets on a large memory machine? The big ann datasets have 1B points and are much more interesting/representative of current embedding use cases (eg from dual encoder models).
I’m also curious if there is a way to not store everything in memory for pgvector. Is that possible?
Lastly, what is the parallelism story? Is it just using a thread pool under the hood? OpenMP?
Understanding if pgvector plans to support point insertions and deletions is also important in practice.
- pashkinelfe 3y agoPgvector doesn't need to store everything in memory. It behaves similar to almost any Postgres AM's and store index data on disk. Performance-wise it's better to have enough memory for index data remain in shared memory buffers, but it is not a requirement for pgvector.
- moab 3y agoThanks for the response. I wonder whether HNSW will still perform well if it needs to page neighbor-lists to/from disk. Do you plan to benchmark the setting where the dataset is too large to fit in-memory?
- pashkinelfe 3y agoIf you have 10M or bigger dataset of real-world OpenAI- dimensional vectors, please share, I'll use it in the next benchmarks. Random datasets are too misleading for vector search benchmarks because all ANN engines make use of internal distributions in datasets to struggle with the curse of dimensionality. So I never use random datasets for ann indexed benchmarking. Using simplified less dimensional (eg 128 instead of 1536) vectors also changes performance trends.
- moab 3y agoSee https://big-ann-benchmarks.com/neurips21.html https://big-ann-benchmarks.com/neurips21.html They're not OpenAI embeddings, but they are realistic, and much larger (number of vectors). I think many production embeddings at non-OpenAI companies will use lower-dimensional vectors than 1536, so it makes sense to focus on non-OpenAI embeddings as well in your benchmarking.