17 ms·
Qdrant, the Vector Search Database, raised $28M in a Series A round
- redwood 3y agoAnyone using Qdrant in prod?
- andre-z 3y agoMany: https://testimonial.to/qdrant/all https://testimonial.to/qdrant/all https://techcrunch.com/2024/01/23/qdrant-open-source-vector-database/ https://techcrunch.com/2024/01/23/qdrant-open-source-vector-...
- swalsh 3y agoI built a little proof of concept that uses it in the RAG pipeline, it's been proving quite useful, so we're just starting the move to production. It's probably going to stay, but I'm also evaluating databricks new vector store as we're using databricks for all the analytics parts of the app already, and having them all on the same infrastructureis appealing.
- crucio 3y agoWe are for a few projects. We've been using them for over a year and have been impressed. We have 10s millions of items in there with lots of daily inserts/deletions etc. There's been a couple of gotchas but generally it is quite predictable and scalable. We use 768 dimensional vectors for our items with several other payload filters (e.g. language). Performance has been good and I think the qdrant team focus on the right features without creeping into other areas.
- inertiatic 3y agoI do, and it's very very rough around the edges to be honest. Lots of things broken, things are even breaking between releases suddenly in unexpected places. Or at least, I'm used to working with more robust data stores. If my work was more high stakes, I'd have already advocated for moving our vector search to something more robust. Thankfully it's not and I can just maintain what we're making with not too much stress, and enjoy seeing this OS project grow from a user perspective (haven't seen a data store go through this very initial phase in my career yet). Support from the team is great however, and congrats to them for this round!
- esafak 3y agoPlease elaborate. What would you have moved to, for example? This is valuable information.
- clbrmbr 3y agoWhat’s the mote here? Seems to be a risky investment when it’s such a crowded space and likely to be decent open source alternatives for those with small budgets and homegrown solutions for companies with bigger budgets and requirements.
- jjackson5324 3y agoIt's a series A. Thinking about the moat in a hot market like vector DBs is a great way to miss out on unicorns.
- tw1984 3y agothe OP argued that it is not a hot market - companies like openai is going to eventually use its own while small players are going to just use openai's assistant APIs, they don't have to operate their own "vector database". it is also worth to mention that even if there is going to be a market called "vector databases", which is highly unlikely, you can't just written off all existing regular databases and pretend that they are not going to just walk in and take over. all in all, there is no reason to believe it is a hot market. it is much better to ask is there going to be a market at all.
- mritchie712 3y agoOpenAI uses Qdrant https://news.ycombinator.com/context?id=38611608 https://news.ycombinator.com/context?id=38611608
- tw1984 3y agofor now. let me repeat what I have already explained - when compared to today's leading AI tech, a "vector database" is just ancient tech. major players are going to build their inhouse solution or they'd conclude it to be some kind of labor intensive & low profit margin baggage and outsource it. you can build a business around it, just like all major tech companies have cleaning guys work for them one way or another, people have to realize that it doesn't make carpet cleaning a high tech or strategically important business.
- weinzierl 3y agoA while ago I read in a thread here that they are used in OpenAI's products and at another popular company. I am not sure but vaguely remember X/Grok. They are also a Rust shop. Who says Germany has no cool startups. EDIT: Yes, it was Grok.
- sroecker 3y agoYes, used by Grok: https://twitter.com/qdrant_engine/status/1721097971830260030 https://twitter.com/qdrant_engine/status/1721097971830260030 Oh wow, completely missed that they're German. Should have noticed their "Impressum"..
- treprinum 3y agoIs Open AI using it in their assistants API for retrieval? Answer performance of those is really bad and retrieval is slow compared to Pinecone.
- simonw 3y agoYes, it's used for the RAG implementation - though we only know that due to information leaked in an error message I believe: https://twitter.com/altryne/status/1721989500291989585 https://twitter.com/altryne/status/1721989500291989585
- saliagato 3y agoHard to believe OpenAI uses Quadrant when they are backed by Microsoft, thus having Azure Cognitive Search (now "AI" Search)
- wodenokoto 3y agoDo you have any experience in AI search to compare it to other products? I’m genuinely curious to know if it’s any good.
- chintler 3y agoCognitive Search is nowhere as good as a 'pure' vector DB. Behind the scenes, it's a managed elasticsearch/opensearch with some vector search capabilities. The 'AI' implementations I've done with Cognitive Search always boil down to hybrid(vector+fts) text search.
- mindvirus 3y agoCongrats to them! What have your experiences with vector databases been? I've been using https://weaviate.io/ https://weaviate.io/ which works great, but just for little tech demos, so I'm not really sure how to compare one versus another or even what to look for really.
- danielbln 3y agoWe're using Postgres with the pg_vector extension for basically all of our projects. We know and love Postgres, it has a big track record, the extension is supported on all major managed cloud offerings, no new tooling needed, pg_vector supports HNSW indices for performance as well. Once in a whole supabase slips into a project, but that's basically just Postgres with some bells and whistles on top. I got nothing bad to say about Pinecone, We aviate, Chroma etc. but when it comes to dbs, I like to go with the devil I know.
- andre-z 3y agoYou should use whatever works best for you unless you face some limitations. The issue is that vector databases are not databases but search engines. It is ACID vs BASE. A few thoughts on this https://qdrant.tech/articles/dedicated-service/ https://qdrant.tech/articles/dedicated-service/
- Fendyfd 3y agoThere are multiple vectordb available in the market, open source ones include Milvus, Qdrant, Weaviate etc. Cloud services include Zilliz Cloud (managed milvus), Qdrant Cloud, Weaviate Cloud etc. Try using a benchmark tool to evaluate them. Here is an open-source option for your reference: VectorDBBench (https://github.com/zilliztech/VectorDBBench https://github.com/zilliztech/VectorDBBench)
- tw1984 3y agowhy would anyone use a "benchmark" tool from a vendor (zilliz here) to test the performance of its competitors?
- shanghaikid 3y agoCongratulations. What do you think you milvus? https://milvus.io/ https://milvus.io/. The difference seems significant from the architecture perspective.
- _mh56 3y agoI applied to Qdrant a while back and got this response: "We are getting many applications for this position. Usually, a test task would help preselect suitable candidates. However, since we develop open-source software, we rely on contribution. You can build an open-source Qdrant connector to another framework or library. The simplest one would be, for example, a Streamlit data connector. But other ideas are more than welcome! No limitations and no deadline. As long as this job position is online, we accept submissions. After you are done, send us an email to career@qdrant.com with the link to the repo. We will review it and get back to you asap." No interviews, conversation before this email. Hope they see and fix this. Edit : No Pay.
- Oras 3y agoDo you expect a tech interview by just applying? From your perspective, you're filling out an application, maybe writing a cover letter, but on the other side, there are 100+ applications like yours. Not all of them are qualified, CVs are not a trustable source anyway. That's why companies add tests to filter first, then interview later.
- _mh56 3y agoI don't expect interviews. But I also don't want to spend 20 hours working before getting a "Unfortunately we've decided not to move forward" message. As a thumb rule, I'm happy to put 4x more effort than the company. If they interview me for 1 hour, I spend 4 hours doing the take-home. Anything more feels like exploitation.
- epolanski 3y agoWell, they said they receive a lot of applications so they are in the position of setting the rules. You are absolutely right into setting your own rules as well, those haven't overlapped.
- pclmulqdq 3y agoAs a general rule of thumb, random series A startups are in much lower demand for top-tier talent than top-tier talent is in demand for these companies. That would mean that the good engineers should set the rules of engagement, and that any startup that thinks they set the rules is attracting worse talent.
- anonzzzies 3y agoOfftopic: Is there a good OSS mixed (vectors + traditional) that can be embedded in our own solution and allows storing indexes in a pluggable kv storage? Besides rolling one, I cannot really find anything. Rust or Go would be best.
- anonzzzies 3y agoThe sourcecode is very readable of this product. And good license, no agpl or worse stuff.
- treprinum 3y agoWhy is AGPL bad? It would prevent Amazon from taking it from the founders and making money off it without giving anything back, like they did to dozens other products.
- JoachimS 3y agoIt also makes it less interesting for other potential customers to use it. Reducing the potential market is probably not what a VC funding a startup wants.
- menaerus 3y agoIMHO it looks arcane to understand or to debug, and with most likely a lot of negative performance implications, due to its shared-ptr-in-disguise all-over-the-place design. $ git grep "Arc<" | wc -l 451 It could be probably related to the fact that the main author of the codebase is coming from the Java/Scala world. Or perhaps it's the Rust safety guarantees.
- nemothekid 3y agoQdrant is a async Rust project, so there will be lots of Arc. Rust safety guarantees doesn't really let share references across threads haphazardly.
- menaerus 3y agoSomething being async doesn't imply shared-ptr design. But perhaps this is what Rust makes you to to achieve its safety guarantees?
- nemothekid 3y ago
- rvz 3y agoWell deserved funding round for a company that underpins most of the AI hype happening all over the place and probably always overlooked by many analysts. Let’s see what they can do in a year or more with that new capital.
- infecto 3y agoI am excited to see how the vector search space plays out. Most of my work is not constrained by a low latency chat type user experience and I have not touched most of the vector search apis. I wonder what the difference is between competitors. The way I picture it is everyone is starting up their own Elasticsearch hosted solution and while there are some differences in functionality, the real bet is cost and scale.
- ankit219 3y agoI think alpha lies in how good the embedding space is rather than which db you use to store and retrieve. A typical tradeoff between accuracy and performance, and here accuracy will be more important in many cases esp for businesses and enterprises. With that, and existing database providers introducing their own support for vectors, this space might be commoditized in near term. Re embeddings, you would likely get better results if you train your own embeddings model. A popular approach is ColBERT, which anecdotally outperforms vector search in border cases[1]. Second is training an embedding model using initial layers of an LLM. [2]. In Colbert's case once it's trained, you dont need a db to store the vectors. [1]: https://twitter.com/arjunkmrm/status/1744741903646773674 https://twitter.com/arjunkmrm/status/1744741903646773674 [2]: https://huggingface.co/intfloat/e5-mistral-7b-instruct https://huggingface.co/intfloat/e5-mistral-7b-instruct
- infecto 3y agoI agree with you. I was ignoring the accuracy/performance tradeoff. Even in that space while there is certainly a lot of innovation left, there is already so much that is available commercially open source. If that holds true, you are really left with competing on price and scale in the long run.
- lettergram 3y agoNot to knock Qdrsnt, but generally the whole “vector search database” rush is insane. I’ve been working with vectors for over a decade; particularly with embeddings used in AI. We’re talking projects from 100k to 100B+ records, used for AI applications Postgres, particularly with pgvector and derivatives, can handle to millions of records very rapidly no problem. It’s very cheap, scales great, and is accurate. I’m sure some of these open source solutions are improvements. That said, weigh vendor lock in, cost, risk and in the end it usually makes very little sense.
- therealdrag0 3y agoWhat’s you use for 100B records?
- lettergram 3y agoIdk if I can talk about my exact project. But I worked alongside these folks: https://www.capitalone.com/tech/machine-learning/learning-embeddings-of-financial-graphs/ https://www.capitalone.com/tech/machine-learning/learning-em... Running an R&D group in the department.
- francoismassot 3y agoCongrats to Qdrant's team, $28M for a Series is really nice. There are a lot of OSS vector search databases out there, we could probably list the main ones: - Qdrant: https://github.com/qdrant/qdrant https://github.com/qdrant/qdrant - Weaviate: https://github.com/weaviate/weaviate https://github.com/weaviate/weaviate - Milvus: https://github.com/milvus-io/milvus https://github.com/milvus-io/milvus What else?
- andre-z 3y agoThese are the major ones, correct.
- mritchie712 3y agoWe use pgvector which if you're already using postgres should be in the running for your use case. I also like https://github.com/lancedb/lancedb https://github.com/lancedb/lancedb
- tajd 3y agoThis website for comparing vector database solutions might be handy https://vdbs.superlinked.com/ https://vdbs.superlinked.com/
- lmeyerov 3y agoIt's funny taking a numbers view. The most popularly used might not even be these, but vector indexes in existing popular OSS DBs and storage systems people are already using. Afaict earliest would be faiss on disk and vectors in opensearch & elasticsearch, and I'd be curious how say databricks, pgvector and other big ones are getting picked up now that they are out. Most of these supported fast & large-scale indexes even early on (ivfpq, ...) by wrapping faiss and friends. ~All OSS DBs we use now, esp managed, have or are getting vector indexes. Another one most similar to qdrant we track internally is lancedb. They are clever by supporting an embedded architecture, so an architectural reason to prefer over most existing OSS DBs. In our survey 2 years ago, we predicted specialized vector DBs having regular OSS DBs be the elephant in the room, and missed embedded as a fundamentally different category: https://gradientflow.com/the-vector-database-index/ https://gradientflow.com/the-vector-database-index/ . (Good luck to qdrant! I'm happy they waited before raising, hopefully this means they can operate more healthily than otherwise and easier to maintain the discipline to do that!)
- tw1984 3y agoI don't think such business model is going to last. There is no reason for AI giants like OpenAI to stick with such external "vector databases". There is not much technical stuff there. Unless you want to argue that "vector searching" is just some labor work when compared to AI, in that case, sure.
- manishsharan 3y agoThere are huge segments e.g. banking, insurance,legal, which are wary of using OpenAI and they would much rather host their own LLMs. I think these vector databases will find a ready market in this segment
- tw1984 3y agoTell me what makes you believe that those big techs are not going suit those "banking, insurance,legal" orgs by providing them their own LLMs? For example, ever heard about github enterprise? you pay a stupid amount, github setup everything almost identical to the public github, just on your servers for your employees. Why big techs won't do the same here? Those high profit margin part of the LLM business is for big players only, they don't burn hundreds of billions to offer you opportunties to cut their profit by capitalizing on their core business. Communism doesn't exist in high tech. People don't work their xxx off to pave ways for your free lunch for life.
- manishsharan 3y agoI wanted to respond to you but your hostile tone implies you are not looking for a conversation.
- tw1984 3y ago[flagged]
- mritchie712 3y agoOpenAI uses Qdrant for chatgpt and a few other products. https://news.ycombinator.com/context?id=38611608 https://news.ycombinator.com/context?id=38611608
- tw1984 3y agoWe have to be honest - "vector" database is a low tech stuff when compared to today's AI. You shouldn't be expecting to walk into the battle of AI, which is arguable the most important one in our life time, to dig a chunk of significant profit from major AI players' pocket by just having some low tech stuff. They use external "vector databases" for now because they don't want to invest R&D resources on such non-key issues for now. for now is the keyword here. When the company grow to 10k or 30k people, there will be teams competing for visibility, someone is going to build their inhouse "vector database" to get his/her slice of the pie. Do you still believe that any AI major player is going to reply on some external vector databases?
- coffeebeqn 3y agoAre in-house databases that common? I thought generally we’ve found as an industry that to be a great thing to purchase. I do wonder how many will need anything other than the vector support in their already existing Postgres instances though
- tw1984 3y ago> I do wonder how many will need anything other than the vector support in their already existing Postgres instances though exactly! if there is a real & strong demand, we'd be seeing open source ones get upgraded & ready in months. it is more like just one of those "I want to build something easy in the core but fancy in its name to get some quick VC $"
- Prosammer 3y agoMy understanding is that these vector search databases are generally used by people who want to use an existing AI and extend it with RAG etc. Was anyone ever expecting major AI players to use a tool like this as you are suggesting?
- tw1984 3y ago> use an existing AI and extend it with RAG etc companies like openai will fill the gap by offering such features out of the box. there is no logical reason why and how a big AI tech giant is going to take all hard work and letting someone else to take the profit by ignore the last mile issue. in fact, openai has already released such APIs in their last devday event.
- klebe 3y ago[flagged]
- pclmulqdq 3y agoThe "unpaid labor" companies never seem to attract good people. Also, using interviews as a way to get free labor is illegal.
- yujian 3y agoGood on them, I know the crustaceans are out here happy about this raise for a Rust based Vector DB! (now I'm gonna plug what I work on) If you're interested in a more scalable vector database written in Go, check out Milvus (https://github.com/milvus-io/milvus https://github.com/milvus-io/milvus)
- andre-z 3y agoThe open-source benchmarks show different results. Feel free to make a PR to improve. ;) https://qdrant.tech/benchmarks/ https://qdrant.tech/benchmarks/
- softwaredoug 3y agoSomeone has to ask the question: How many vector DBs do we really need? How do the vector DB companies differentiate themselves? And why do we need a company at all when there are increasingly awesome open source options? I genuinely ask - there are a lot of other problems in the RAG, fine tuning, AI/LLM, retireval space, to solve. And more and more vector retrieval is, while not 100% solved, at least is something the community has a grasp on the tradeoffs. Solved to the point that squeezing a bit more recall out of vector retrieval isn't the problem anymore.
- inertiatic 3y ago>Solved to the point that squeezing a bit more recall out of vector retrieval isn't the problem anymore. I think this is a bit of a strawman. I don't think recall is the main point these systems are trying to sell us on, it's more about robustness and ease of use compared to building something inhouse or using a lower level library to build a system on top of it just for this small part of your overall project/product (be it RAG, search, whatever). I guess Lucene-based solutions, while very mature overall in terms of engineering, lagged behind this functionality (out of caution, trying to build what's going to be long term useful) and are also perceived a bit too cumbersome. So these stores do make sense, I think. The core functionality is nothing too complex (at least HNSW), but hiding it behind a stable black box with just a few inputs and levers, has value for people that are likely to use these stores.
- sanp 3y agoAgree but then the same argument applies to RDBMSs and multiple vendors seem to be doing OK in that space. I think it ultimately comes down to "stuff" (sales journey, price, support etc.) other than the technology itself. I am sure any RDBMS can meet most of the requirements of any given customers (in most cases) but we still see customers buying across vendors.
- esafak 3y agoqdrant is open source. Being open source is not in opposition to running a company; it is part of their strategy. There is still work to be done in vector databases. None of the products have perfected hybrid search yet, for example, and performance varies a lot between products; they are not fungible.
- hartator 3y ago> For example, it can automatically map ‘frontend engineer’ to ‘web developer’ Small revolution indeed. Ref: https://qdrant.tech/use-cases/ https://qdrant.tech/use-cases/
- avereveard 3y agoI really don't understand that sample the similarity capability is provided by the external embeddings model not by quadrant per se, unless they have some proprietary embeddings.
- minimaxir 3y agoCorrect, but that one-pager is more aimed toward project managers than engineers. Marketing copy is weird like that.
- ancorevard 3y agoHonest question, how long before EU makes it unattainable for Qdrant to remain in Germany/EU?
- wahnfrieden 3y agoWhat’s the best vector db for text similarity that can run in browser front ends too?
- spullara 3y agoHonestly there is no reason, except huge scale, to have a separate vector db. Every normal database and search engine now support vector search.
- braza 3y agoOutside AI and LLMs, there are some solid use cases for those Vector Search Databases? Maybe I am not seeing something, but it’s hard to see it gaining traction outside tech companies.
- esafak 3y agoVector databases enable semantic and similarity search. What company does not need that?
- beernet 3y agoCompanies that don't want/can build it by themselves, so the majority of enterprises. It's a nice series A by the numbers, at the same time, generating relevant revenue will very likely not happen (given the valuation at this round was probably around 200M€). It's hype all over but can't blame them, would do the same I guess.
- yding 3y agoCongrats! Amazing milestone.