4 ms·
Ask HN: Disk-backed NumPy for data analysis?
I have a large similarity matrix. It is 300K by 300K entries, and is ~ 100GB in size. To clarify, this is the result of multiplying a 300K by 512 matrix with its transpose.
I want to make queries on this matrix like (what are the rows most similar to row #2), but I don't have a machine with 100GB of RAM.
What would be the best strategy for running queries on this matrix? My current plan is to use duckdb, but I am wondering if there are more elegant alternatives.
- tacosbane 4y agoyou didn't provide any hints why it wouldn't work for you, so i recommend you look at scikit-learn's neighbors module (i.e., construct NearestNeighbors object and query it). https://scikit-learn.org/stable/modules/neighbors.html#unsupervised-nearest-neighbors https://scikit-learn.org/stable/modules/neighbors.html#unsup...
- prirun 4y agoYou can use swap / virtual memory. The tricky part with that or any db is to make sure that your array accesses are as sequential as possible, because in the worst case, each memory reference could turn into a disk access. Make sure your swap is on a fast NVME/SSD device to minimize access time. See mkswap.
- jstx1 4y ago> I want to make queries on this matrix like (what are the rows most similar to row #2) There are libraries for efficient vector similarity search like https://github.com/facebookresearch/faiss https://github.com/facebookresearch/faiss
- kacperlukawski 4y agoFor that scale FAISS might be not enough, as it keeps the data in memory and is simply hard to scale. It's worth considering to use a proper vector database in a distributed mode to spread the load, or just allow storing some data on disk: https://qdrant.tech/articles/memory-consumption/ https://qdrant.tech/articles/memory-consumption/
- PaulHoule 4y agoTry https://www.dask.org/ https://www.dask.org/ Also consider https://www.pinecone.io/ https://www.pinecone.io/ The obvious way to do a similarity query is to do a full-scan which can be done in a straightforward way against data on disk if the data is appropriately packaged. It may be slow but it works. There are n-dimensional indexes that can accelerate similarity queries, pinecone uses them, they are not as effective as 1-d and 2-d (geospatial) indexes but they do help. Sparse similarity search is its own problem, solved by full-text search engines like Lucene. Dimensional reduction, say going from 300K features to 100 features, might improve the quality of your results as well as compactifying your data. It could be something you do once on a monster machine and then do many lookups on a smaller machine. That dimensional reduction might take a lot of resources, see https://scikit-learn.org/stable/modules/decomposition.html https://scikit-learn.org/stable/modules/decomposition.html but in your case you might get embarassingly good results with random projections https://en.wikipedia.org/wiki/Random_projection https://en.wikipedia.org/wiki/Random_projection
- vsroy 4y agoThanks for the response. One thing to clarify: It's a 300K x 300K similarity matrix, which means I have 300K embeddings. Each embedding itself only has dimension 512. In other words, the similarity matrix is the similarity between each embedding & every other embedding in the 300K set of embeddings. Regardless, I think Dask will be useful here.
- yorwba 4y agoUse mmap to load only the region of the file that corresponds to row #2, interpret it as an array using np.frombuffer, proceed as with normal in-memory data. This works as long as you control the file format (so you can ensure that rows are stored contiguously) and as long as you don't need too many rows in memory at the same time.