4 ms·
how does this do with large-scale searches >10M+ rows? any benchmarks on performance?
by phenkdo 5y ago
how does this do with large-scale searches >10M+ rows? any benchmarks on performance?
- minxomat 5y agoOh it's terrible. More educational. You'd never ever want to do a full-scan cosine matching in production. Use locality sensitive hashing (with optimizations like amplification, SuperBit or DenseFly) for real world workloads.
- gk1 5y agoSince the answer, in minxomat's words, is "terrible," maybe look at Pinecone (https://www.pinecone.io https://www.pinecone.io) which makes light work of searching through 10M+ rows. It sits alongside, not inside, your Postgres (or whatever) warehouse. Disclosure: I work there.
- fzysingularity 5y agoOut of curiosity, what kind of performance do you get with 100M rows with pinecone? Looking at your pricing tiers, ~100M rows would need ~200GB memory, and @ $0.1 / GB / hr that's $20 / hr if I'm not mistaken? Also, can you join with existing SQL tables to do hybrid searches in pinecone?
- gk1 5y agoOur performance is independent of the collection size, thanks to dynamic sharding. You can expect 50-100ms pretty much regardless of size. We don't support external SQL joins yet. Depending on what you're trying to do, we have an upcoming feature that might do the job. At 100M vectors you're well into "volume discount" territory. Even more so with 3B vectors, as you mentioned in another comment. The free trial is capped at 10GB but shoot me an email (greg@pinecone.io) to get around it for a proper test and pricing estimate.
- fzysingularity 5y agoOthers have pointed this out already, but have a look at Milvus (https://github.com/milvus-io/milvus https://github.com/milvus-io/milvus). I was able to get a simple version of it running with searches over ~3B vectors in under a second on a single machine (with out of the box configuration, and practically no optimization done just yet).