5 ms·
Postgres has a extension called cube (https://www.postgresql.org/docs/9.6/cube.html https://www.postgresql.org/docs/9.6/cube.html) that can be used for up to 15
by changelink 7y ago
Postgres has a extension called cube (https://www.postgresql.org/docs/9.6/cube.html https://www.postgresql.org/docs/9.6/cube.html) that can be used for up to 150 dimensions (which is a compile-time limit, you can have more if you compile Postgres yourself).
It's a pretty cool extension that does distance between points, intersections between n-dimensional cubes (hence the name), different distance metrics etc.
It'd be perfect for storing and searching through large amounts of n-dimensional embeddings, I'm guessing it's used for that already.
- orf 7y agoFun fact, there’s a bug in the implementation where you can create cubes with much higher dimensions because one constructor doesn’t do the check. Still wouldn’t recommend it though.
- kaivi 7y agoIn my experience, the cube extension is unusable for >10M x 128D vectors without PCA. I'm using Faiss now with ~500M vectors, and it works great!
- phelmig 7y agoWith how many dimensions are you using Faiss with 100m+ vectors? I’m currently looking a solution to handle 1024 dimensions for ~100m items.
- kaivi 7y agoOn one index I'm using OPQ16_64,IVF262144_HNSW32,PQ16 with 128 dimensions initially. 1024 dimensions is a lot! Could you elaborate on what application requires that many? If it's a DNN layer output, your data must be sparse, so dimensionality reduction won't affect your recall if tuned properly.
- phelmig 7y agoIt's actually a DNN layer output. I haven't considered dimensionality reduction, yet. Thanks for pointing my there, I'll look into it. Probably thats the better way to go. Thanks a lot for your reply!