3 ms·
Congrats to Qdrant's team, $28M for a Series is really nice. There are a lot of OSS vector search databases out there, we could probably list the main ones: -
by francoismassot 3y ago
Congrats to Qdrant's team, $28M for a Series is really nice.
There are a lot of OSS vector search databases out there, we could probably list the main ones:
- Qdrant: https://github.com/qdrant/qdrant https://github.com/qdrant/qdrant
- Weaviate: https://github.com/weaviate/weaviate https://github.com/weaviate/weaviate
- Milvus: https://github.com/milvus-io/milvus https://github.com/milvus-io/milvus
What else?
- andre-z 3y agoThese are the major ones, correct.
- mritchie712 3y agoWe use pgvector which if you're already using postgres should be in the running for your use case. I also like https://github.com/lancedb/lancedb https://github.com/lancedb/lancedb
- tajd 3y agoThis website for comparing vector database solutions might be handy https://vdbs.superlinked.com/ https://vdbs.superlinked.com/
- lmeyerov 3y agoIt's funny taking a numbers view. The most popularly used might not even be these, but vector indexes in existing popular OSS DBs and storage systems people are already using. Afaict earliest would be faiss on disk and vectors in opensearch & elasticsearch, and I'd be curious how say databricks, pgvector and other big ones are getting picked up now that they are out. Most of these supported fast & large-scale indexes even early on (ivfpq, ...) by wrapping faiss and friends. ~All OSS DBs we use now, esp managed, have or are getting vector indexes. Another one most similar to qdrant we track internally is lancedb. They are clever by supporting an embedded architecture, so an architectural reason to prefer over most existing OSS DBs. In our survey 2 years ago, we predicted specialized vector DBs having regular OSS DBs be the elephant in the room, and missed embedded as a fundamentally different category: https://gradientflow.com/the-vector-database-index/ https://gradientflow.com/the-vector-database-index/ . (Good luck to qdrant! I'm happy they waited before raising, hopefully this means they can operate more healthily than otherwise and easier to maintain the discipline to do that!)
- jillesvangurp 3y agoThere are a few more. Pinecone comes to mind. And then there are traditional databases and search products that are integrating vector search capabilities as well: Postgres, Elasticsearch, Opensearch, Solr. They each have their limitations of course but the 28M round suggests a moat that I'm not seeing that clearly in terms of tech. What's so special about qdrant relative to their competition? At least they are Apache licensed for now. So, that's nice. But that also means e.g. Apache Lucene could borrow some code from them to beef up their vector search capabilities. Which would benefit Elasticsearch, Opensearch, and Solr which all depend on Lucene. Which raises the question what the point is of QDrant long term and why investors are betting on this as opposed to other things. It seems to me that the main challenge with vector search is inference cost (at index and query time), not storing the vectors. A secondary concern is the vector comparisons at query time. A good way to cut down on that is to reduce the overall result set using traditional search or query mechanisms. In other words, you need
- lsaferite 3y agoIs Pinecone OSS? I ask because this was the statement from the PP There are a lot of OSS vector search databases out there, we could probably list the main ones ... What else?
- utopcell 3y agoPinecone is not OSS.
- manishsharan 3y agoI think there will be enough of market to justify a few more dedicated VectorDB vendors. From the enterprise perspective, which of these vendors proved the best combination of security, availability, performance and pricing will matter. when we run benchmarks on our (self hosted) LLMs, we do not a clear idea of where we have bottlenecks and we end up assuming its the GPU/memory. And our pilot implementation will never go into production as the security model is nearly non existent in our implementations; the execs AND qa are getting the same RAG outputs. It is all very new to us and our teams. If a vendor can outperform its competition in our tests and show credible security model with segmentation of knowledge, that would be the choice.
- epistasis 3y agoIt's fascinating to see the diversity of vector databases! I've chosen to prototype with two, ChromaDB, and LanceDB, based on the ease of using embeddings with them, and had not even heard of these others here. I'm also very excited to go through VectorHub's table of databases: https://vdbs.superlinked.com https://vdbs.superlinked.com (discovered from sibling comment here: https://news.ycombinator.com/item?id=39103322 https://news.ycombinator.com/item?id=39103322)
- daveed 3y agoI think Activeloop(YC) is too: https://github.com/activeloopai/deeplake/ https://github.com/activeloopai/deeplake/
- rgbrgb 3y agoI've been looking at this one to embed in desktop apps https://github.com/unum-cloud/usearch https://github.com/unum-cloud/usearch
- alfalfasprout 3y agoSure, but frankly it's historically very hard to build a business around a specialized database, especially if you have competitors that are even 80% as good but free. The cases where I've seen this work are when the DB offers something way ahead of what their competitors offer. For example, KDB+ was historically unrivaled when it came to ultra high performance time series storage and Aerospike is very hard to beat for extremely high performance multi-node K/V. Otherwise there's little to stop a larger company from offering the OSS competitor to your DB as a service for a lower cost and invest eng resources to close the gap.
- shenli3514 3y agoChroma looks good. https://github.com/chroma-core/chroma https://github.com/chroma-core/chroma 10k+ stars, very easy to use, and can be used as an embedding database
- morgango 3y agoElasticsearch: https://www.elastic.co/platform https://www.elastic.co/platform
- bumberg 3y agoMarqo.ai (https://github.com/marqo-ai/marqo https://github.com/marqo-ai/marqo) is doing some interesting stuff and is oss. We handle embedding generation as well as retrieval (full disclosure, I work for Marqo.ai)