3 ms·
> we never needed any "vector DB" ... the integration of real data is complex, embeddings get created through platforms like Airflow, metadata in a document sto
by shepardrtc 3y ago
> we never needed any "vector DB" ... the integration of real data is complex, embeddings get created through platforms like Airflow, metadata in a document store, and made available for fast retrieval on low-gravity disk (e.g. built on top of something like rocks db... then something like ANN is easy on tools like FAISS)
A good vector DB would take care of all of that.
- james-revisoai 3y agoI think the point is they already had those components or familiarity with them in their system hooked into existing data. You could add semantic search with text vectors 3 years ago within 1 hour and so the need for a service isn't there. They didn't need a simultaneous second DB for records just matching by a foreign ID key based on a new service. I have seen it too (as somebody who studied Information Retrieval and these systems and implemented aspects in production at 2 startups, including more recently these semantic search approaches) - you can store vectors (floats) of length 384, 512, 768 or 2048 like most of the current ones easily alongside existing data without typically making a dent in overall data storage, and load into memory for semantic search according to your needs (for example, for big data you can use more bucketing approaches with don't guarantee the nearest neighbour, but one of the roughly closest ones, which is often a trade off worth making). It's not only not a hassle to do that, but overall superior to using the vector DB online tools for developer experience too. I suspect at 3 million+ rows you begin to have struggles, but barely any services will get there. And if you do, you may try other techniques (dimensionality reduction to 3x smaller often only loses a tiny portion of the data accuracy, but saves 2/3rds on costs), or employ cutting edge embedding techniques like Custom Matrix Optimisation, multi modal alignment, and others that these services don't necessarily offer. Sure it's good for the beginner/a tutorial website, but it doesn't solve a problem for developers in information retrieval imo.