5 ms·
I think the move towards vector databases might be more hype than necessity. Traditional databases, when properly optimized, can handle vector data for many use
by loondri 3y ago
I think the move towards vector databases might be more hype than necessity. Traditional databases, when properly optimized, can handle vector data for many use cases. The push for specialized vector databases could be re-evaluated in terms of efficiency and cost-effectiveness compared to optimizing existing scalar databases.
- avereveard 3y agoWell you could store numbers all fine, but indexing vectors for similarity queries seems fairly recent and not all that widespread in the transactional world. As the traditional db move forward in the space the need for dedicated vector databases will likely shrink, except for some very specific implementation that offer unique enough features (I.e. deeplake does vector search over object storage, which is very convenient for certain specific scenarios)
- sgu999 3y agosqlite has r-trees for instance [0]. Could it be good enough for most use cases? If it's to query a knowledge base for instance, a couple dimensions should be sufficient. With the added benefit of being able to query your data in other ways. [0] https://www.sqlite.org/rtree.html https://www.sqlite.org/rtree.html
- mattashii 3y agor*-trees work doesn't work well when the number of dimensions stored in the index is much higher than the logarithm of the number of indexed entries, and this is a prevailing property of divide-and-conquer spatial index types when the keyspace is divided based on a single dimension at a time. As vectors regularly have 100+ dimensions, normal spatial indexing methods applied to vectors wouldn't be very efficient for anything with much less than 2^100 index entries; which is quite suboptimal for most datasets that you would want to have indexed.
- KirillPanov 3y agoAlso the distance metric for r*-trees is just plain wacky for anything other than low-dimensional Euclidean space. Even if you could make it perform well, it would not do what you want.
- modulovalue 3y agoAre you saying this because r-trees expect a proper metric space, and people have the need to index datasets over non-metric spaces?
- pornel 3y agoThe curse of dimensionality creates a seemingly paradoxical situation where you have a vast vast search space, but everything is incredibly close to each other. Space subdivision algorithms become ineffective.
- jpcapdevila 3y agoHere is a SQLite extension that uses Faiss under the hood. https://github.com/asg017/sqlite-vss https://github.com/asg017/sqlite-vss Not associated with the project, just love SQLite and find it very useful.
- thesz 3y agoWhat is "vector search over object storage?" Does deeplake performs some computations on objects and search on their embeddings?
- avereveard 3y agoIt stores everything on cheap storage, with no compute associated (i.e. S3) and uses the client to compute the query embeddings and to retrieve the embedding index and to run an indexed search to identify the data to be retrieved, and likewise the client does the work of updating the index structure on writing. The benefit is that you don't have to pay for the compute part of a database, and the storage layer is as cheap as it could be on the cloud.
- deepGem 3y agoAt the expense of latency ? when in fact latency is the most important aspect for any search. Any idea on how fast the searches from the client are ? retrieve the embedding index and to run an indexed search to identify the data to be retrieved. Please bear with the layman like questioning - So if the data is {obj: "obj1, "data": {"name": "atlas", "embedding": "1123124234" } What is an embedding index ? Is it something like {"1123124234": "obj1"} ? From what I understand the query will be "geography" whose embedding will be "12311111" and now you have to run a KNN for a match which will return {"name": "atlas", "embedding": "1123124234"} Not sure where the embedding index comes into play here.
- avereveard 3y agoEh, sure latency is suboptimal. But if you have a LLM in the mix, that latency will dominate the overall response time. At that point you might not care about how performant your index is, and since performance/cost is non linear, it can translate to very significant savings
- deepGem 3y agoHow is indexing a vector different from indexing a varchar or an integer ? If you convert a vector into a byteaarray it should be no different from a bytearray of varchar but for the bytearray contents. Now if you want to do similarity search you have to measure the distance between 2 or more vectors and that's independent of the indexing. No ? So any database with sufficient memory should be able to accomplish this as evidenced by the vector similarity search feature of Redis. ( I don't know how Redis folks have implemented vector similarity but they do support KNN search )
- sgarland 3y agoMostly the number of dimensions. Assuming your vectors are float16, so 2 bytes each, you’d run into Postgres’ B+tree index limit (2704 bytes) very quickly. You could index a 512-dimension vector fine, but I believe most models are well beyond that. There are alternative index types, of course, or you could index the hash of the vector. These both come with tradeoffs.
- smilliken 3y agoBtree isn't a very useful index type for a vector, though. GIN, GIST, and the handful of new extensions optimizing for vector search are what you'd want (and don't have this limitation). Aside, you can increase the size of tuples you can index in a postgresql btree by increasing the postgresql page size (requires a recompile and creating a new database instance).
- sgarland 3y agoAgreed to the first, but you have to first know that those exist (and what they’re good for). This leads into my second point: IME, the Venn diagram for “people making AI stuff” and “people capable of compiling and running their own DB in a reliable manner” has no overlap.
- JustLurking2022 3y agoThe distance computation could be separate from the indexing, but it will be inefficient relative to having an index organized to support the task.
- mnky9800n 3y agoTo be fair, Vector databases does sound more official as "new and important technology" compared to the last db hype of NOSQL.
- Guvante 3y agoI mean NOSQL was hype with no substance but "you can scale more if you deal with not having ACID" is just generally true. Of course ACID scales to well into the Fortune 500 scale so...
- smitty1e 3y ago"No substance" seems a bit harsh. They mostly seem a tarted up associative array, sure, but a key-value store is a thing.
- mnky9800n 3y agoBut don't you prefer your key value stores to be wearing red lipstick and a pushup bra?
- threeseed 3y ago> but a key-value store is a thing DynamoDB underpins much of AWS which in turn underpins a ridiculous number of web services. So definitely more than just a thing.
- rakoo 3y agoA naturally distributed key/value store. NoSQL wasn't a product, it was a radical rethinking of the balance between what we wanted and what we needed. Turns out some absolute necessities of using a RDBMS or even using SQL are not that important, and relaxing those actually-not-requirements allows massive scalability, something we desperately needed, or some evolved data structures and computation, like what redis provides.
- deleted 3y ago[deleted]
- 3y ago
- jjtheblunt 3y agoWhat would you use to compute proximity of vectors, for example?