9 ms·
Will Amazon S3 Vectors kill vector databases or save them?
- Fendy 1y agowhat do you think?
- sharemywin 1y agoit's annoying to me that there's not a doc store with vectors. seems like the vector dbs just store the vectors I think.
- jeffchuber 1y agochroma stores both
- nkozyra 1y agoAs does Azure's AI search.
- intalentive 1y agoI just use sqlite
- storus 1y agoPinecone allows 40k of metadata with each vector which is often enough.
- whakim 1y agoElasticsearch and Vespa both fit the bill for this, if your scale grows beyond the purpose-built vector stores.
- simonw 1y agoElasticsearch and MongoDB Atlas and PostgreSQL and SQLite all have vector indexes these days.
- KaoruAoiShiho 1y ago> MongoDB Atlas It took a while but eventually opensource dies.
- CuriouslyC 1y agoMy search service Lens returns exact spans from search, while having the best performance both in terms of latency and precision/recall within a budget. I'm just working on release cleanup and final benchmark validation so hopefully I can get it in your hands soon.
- resters 1y agoBy hosting the vectors themselves, AWS can meta-optimize its cloud based on content characteristics. It may seem like not a major optimization, but at AWS scale it is billions of dollars per year. It also makes it easier for AWS to comply with censorship requirements.
- barbazoo 1y ago> It also makes it easier for AWS to comply with censorship requirements. Does it, how? Why would it be the vector store that would make it easier for them to censor the content? Why not censor the documents in S3 directly, or the entries in the relational database. What is different about censoring those vs a vector store?
- resters 1y agoOnce a vector has been generated (and someone has paid for it) it can be searched for and relevant content can be identified without AWS incurring any additional cost to create its own separate censorship-oriented index, etc. AWS can also add additional bits to the vector that benefit its internal goals (scalability, censorship, etc.) Not to mention there is lock-in once you've gone to the trouble of using a specific embedding model on a bunch of content. Ideally we'd converge on backwards-compatible, open source approaches, but cloud vendors want to offer "value" by offering "better" embedding models that are not open source.
- simonw 1y agoThis is a good article and seems well balanced despite being written by someone with a product that directly competes with Amazon S3. I particularly appreciated their attempt to reverse-engineer how S3 Vectors work, including this detail: > Filtering looks to be applied after coarse retrieval. That keeps the index unified and simple, but it struggles with complex conditions. In our tests, when we deleted 50% of data, TopK queries requesting 20 results returned only 15—classic signs of a post-filter pipeline. Things like this are why I'd much prefer if Amazon provided detailed documentation of how their stuff works, rather than leaving it to the development community to poke around and derive those details independently.
- speedysurfer 1y agoAnd what if they change their internal implementation and your code depends on the old architecture? It's good practice to clearly think about what to expose to users of your service.
- altcognito 1y agoKnowing how the service will handle certain workloads is an important aspect of choosing an architecture.
- libraryofbabel 1y agoIf you can truly abstract away an internal detail, then great. But often there are design decisions that you cannot abstract away because they affect e.g. performance in a major way. For example, I don't care whether some AWS service is written in Java or Go or C++. I do care a bit about how its indexing and retrieval works, because I need to know that to plan my query workloads. I actually think AWS did a reasonably good job of this with DynamoDB. Most of the performance tradeoffs, indexing etc., is pretty clear if you ready enough docs without exposing a ton of unnecessary internals.
- alanwli 1y agoThe alternative is to find solutions that can reasonably support different requirements because business needs change all the time especially in the current state of our industry. From what I’ve seen, OSS Postgres/pgvector can adequately support a wide variety of requirements for millions to low tens of millions of vectors - low latencies, hybrid search, filtered search, ability to serve out of memory and disk, strong-consistency/transactional semantics with operational data. For further scaling/performance (1B+ vectors and even lower latencies), consider SOTA Postgres system like AlloyDB with AlloyDB ScaNN. Full disclosure: I founded ScaNN in GCP databases and am the lead for AlloyDB Semantic Search. And all these opinions are my own.
- storus 1y agoDoes this support hybrid search (dense + sparse embeddings)? Pure dense embeddings aren't that great for specific search, they only hit meaning reliably. Amazon's own embeddings also aren't SOTA.
- infecto 1y agoThat’s where my mind was rolling and also if not, can this be used in OpenSearch hybrid search?
- danielcampos93 1y agoI think you would be very surprised by the number of customers who don't care if the embeddings are SOTA. For every Joe who wants to talk GraphRAG + MTEB + CMTEB and adaptive rag there are 50 who just want whatever IT/prodsec has approved
- qaq 1y ago"I recently spoke with the CTO of a popular AI note-taking app who told me something surprising: they spend twice as much on vector search as they do on OpenAI API calls. Think about that for a second. Running the retrieval layer costs them more than paying for the LLM itself. That flips the usual assumption on its head." Hmm well start sending full documents as part of context see it flip back :).
- heywoods 1y agoEgress costs? I’m really surprised by this. Thanks for sharing.
- qaq 1y agoSry maybe should've being more clear it was a sarcastic remark. The whole point of doing vector db search is to feed LLM with very targeted context so you can save $ on API calls to LLM.
- infecto 1y agoThat’s not the whole point it’s in the intersection of reducing tokens sent but also getting search both specific and generic enough to capture the correct context data.
- j45 1y agoIt's possible to create linking documents between the documents to help smooth out things in some cases.
- heywoods 1y agoNo worries. I should probably make sure I have at least a token understanding of the topic cloud based architecture before commenting next time haha.
- andreasgl 1y agoThey’re likely using an HNSW index, which typically requires a lot of memory for large data sets.
- scosman 1y agoAnyone interested in this space should look at https://turbopuffer.com https://turbopuffer.com - I think they were first to market with S3 backed vector storage, and a good memory cache in front of it.
- nosequel 1y agoTurbopuffer was mentioned in the article.
- k9294 1y agoTurbopuffer is awesome, really recommend it. Also they have extra features like automatic recall tuning based on you data, option to choose read after write guarantees (trading latency for consistency or vice versa), BM25 search, filtering on the filed and many more. Really recommend to check them out if you need a vector DB. I tried qdrant and zilli cloud solutions and in terms of operational simplicity turbopuffer just killing it. https://turbopuffer.com/docs/query https://turbopuffer.com/docs/query
- redskyluan 1y agoAuthor of this article. Yes, I’m the founder and maintainer of the Milvus project, and also a big fan of many AWS projects, including S3, Lambda, and Aurora. Personally, I don’t consider S3Vector to be among the best products in the S3 ecosystem, though I was impressed by its excellent latency control. It’s not particularly fast, nor is it feature-rich, but it seems to embody S3’s design philosophy: being “good enough” for certain scenarios. In contrast, the products I’ve built usually push for extreme scalability and high performance. Beyond Milvus, I’ve also been deeply involved in the development of HBase and Oracle products. I hope more people will dive into the underlying implementation of S3Vector—this kind of discussion could greatly benefit both the search and storage communities and accelerate their growth.
- redskyluan 1y agoBy the way, if you’re not fully satisfied with S3Vector’s write, query, or recall performance, I’d encourage you to take a look at what we’ve built with Zilliz Cloud. It may not always be the lowest-cost option, but it will definitely meet your expectations when it comes to latency and recall.
- pradn 1y agoThanks for writing a balanced article - much easier to take your arguments seriously! And a sign of expertise.
- Shakahs 1y agoWhile your technical analysis is excellent, making judgements about workload suitability based on a Preview release is premature. Preview services have historically had significantly lower performance quotas than GA releases. Lambda for example was limited to 50 concurrent executions during Preview, raised to 100 at GA, and now the default limit is 1,000.
- cpursley 1y agoPostgres has pgvector. Postgres is where all of my data already lives. It’s all open source and runs anywhere. What am I missing with the specialty vector stores?
- CuriouslyC 1y agolatency, actual retrieval performance, integrated pipelines that do more than just vector search to produce better results, the list goes on. Postgres for vector search is fine for toy products or stuff that's outside the hot loop of your business but for high performance applications it's just inadequate.
- cpursley 1y agoFor the vast majority of applications, the trade off is worth keeping everything in Postgres vs operational overhead of some VC hype data store that won’t be around in 5 years. Most people learned this lesson with Mongo (postgrest jsonb is now good enough for 90% of scenarios).
- cpursley 1y agoAlso, no way retrieval performance is going to match pgvector because you still have to join the external vector with your domain data in the main database at the application level, which is always going to be less performant.
- CuriouslyC 1y agoFor a large class of applications, the database join is the last step of a very involved pipeline that demands a lot more performance than PGVector can deliver. There are also a large class of applications that don't even interface with the database directly, except to emit logging/traceability artifacts.
- jitl 1y agoi'll take a 100ms turbopuffer vector search plus a 50ms postgres-select-where-id-in over a 500ms all-in-one pgvector + join query. When you only need to hydrate like 30 search result item IDs from Postgres or memcached i don't see the join being "too expensive" to do in memory.
- rubenvanwyk 1y agoI don’t think it’s either-or, this will probably become the default / go-to - if you aren’t storing your vectors in your db like Neon or Turso. As far as I understand, Milvus is appropriate for very large scale, so will probably continue targeting enterprise.
- giveita 1y agoBetteridge can answer No to two questions at once!
- janalsncm 1y agoS3 vectors has a topK limit of 30, and if you add filters it may be less than that. So if you need something with higher topK you’ll need to 1) look elsewhere or 2) shard your dataset into N shards to get NxK results, which you query in parallel and merge afterwards. I also didn’t see any latency info on their docs page https://docs.aws.amazon.com/AmazonS3/latest/API/API_S3VectorBuckets_QueryVectors.html https://docs.aws.amazon.com/AmazonS3/latest/API/API_S3Vector...
- mediaman 1y agoAnd a topk of 30 also means reranking of any sort is out, except for maybe limited reranking of 30->10, but that seems kind of pointless with today’s LLMs that can handle a bit more context.
- janalsncm 1y agoYeah exactly, so you could do something like shard by the first 4 bits of md5 of the text (gives you 16 buckets) but now you’re adding extra complexity to work around their limitations.
- catlifeonmars 1y ago3) ask TAM for a service quota increase
- conradev 1y agoAt a glance, it looks like a lightweight vector database running on top of low-cost object storage—at a price point that is clearly attractive compared to many dedicated vector database solutions. They also didn’t mention LanceDB, which fits this description but with an open source component: https://lancedb.github.io/lancedb/ https://lancedb.github.io/lancedb/
- kjfarm 1y agoThis may be because LanceDB is the most attractive with a price point of standard S3 storage ($0.023/GB vs $0.06/GB). I also like that Lancedb works with S3 compatible stores, such as Backblaze B2 which is even cheaper (~70% cheaper).
- nickpadge 1y agoI love lancedb. It’s the only way I’ve found to performantly and cheaply serve 50m+ records of 768 dimensions. Runs on s3 a bit too slow, but on EFS can still be a few hundred millis.
- factsaresacred 1y agoFor low cost, there's also Cloudflare Vectorize ($0.05 per 100 million stored vectors), which nobody seems to know exists: https://www.cloudflare.com/developer-platform/products/vectorize/ https://www.cloudflare.com/developer-platform/products/vecto...
- hbcondo714 1y agoIt would be great to have the vector database run on the edge / on-device for offline-first and be privacy-focused. https://objectbox.io/ https://objectbox.io/ does this but i would like to see AWS and others offer this as well.
- greenavocado 1y agoI am already using Qdrant very heavily for code dev (RAG) and I don't see that changing any time soon because its the primary choice for the tools I use and it works well
- j45 1y agoThe cloud is someone else's computer. If it's this sensitive, there's a lot of companies staying on the sidelines until they can compute in person, or limiting what and how they use it.
- teaearlgraycold 1y ago> Not too long ago, AWS dropped something new: S3 Vectors. It’s their first attempt at a vector storage solution Nitpick: AWS previously funded pgvector (the slow down in development indicates to me they have stopped). Their hosted database solutions supported the extension. That means RDS and Aurora were their first vector storage solutions.
- softwaredoug 1y agoI’m not sure S3 vectors is a true vector database/search engine in the way something like Elasticsearch, Turbopuffer or Milvus is. It’s more a convenient building block for simple high scale retrieval. I think of a search system doing quite a lot from sparse/lexical/hybrid search, metadata filtering, numerical ranking (recency/popularity/etc), geo, fuzzy, and whatever other indices at its core. These are building blocks for getting initial candidates. Then you need to be able to combine all these into one result set for your users - usually with a query DSL where you can express a ranking function. Then there’s usually ancillary features that come up (highlighting, aggregations, etc). So while S3 vectors is a fascinating primitive, I’m not sure I’d reach for it outside specific circumstances.
- anonu 1y agoIf you like to die in a slow and expensive way - sure.
- jhhh 1y ago"That gap isn’t just theoretical—it shows up in real bills." "That’s not linear growth—it’s a quantum leap" "The performance and recall were fantastic—but the costs were brutal" "it’s not a one-size-fits-all solution—it’s the right tool for the right job." "S3 Vectors is excellent for cold, cheap, low-QPS scenarios—but it’s not the engine you want to power a recommendation system" "S3 Vectors doesn’t spell the end of vector databases—it confirms something many of us have been seeing for a while" "that’s proof positive that vector storage is a real necessity—not just “indexes wrapped in a database." "the vector database market isn’t being disrupted—it’s maturing into a tiered ecosystem where different solutions serve different performance and cost needs" "The golden age of vector databases isn’t over—it’s just beginning." "The bigger point is that Milvus is evolving into a system that’s not only efficient and scalable, but AI-native at its core—purpose-built for how modern applications actually work."
- turing_complete 1y agoSince when was everything no longer "announced" or "released", but "dropped"? Is this an LLMism?
- Urahandystar 1y agoNo you're just old. Come sit with us in a nice comfy chair.
- fragmede 1y agoStarted in the 1988, with music, then expanded from there. https://english.stackexchange.com/questions/632983/has-drop-recently-acquired-the-meaning-of-releasing-digital-content-beyond-mu https://english.stackexchange.com/questions/632983/has-drop-...
- iknownothow 1y agoS3 has much bigger fish in its sight than the measely vector db space. If you see the subtle improvements in features of S3 in recent years, it is clear as day, at least to me, that they're going after the whale that is Databricks. And they're doing it the best way possible - slowly and silently eating away at their moat. AWS Athena hasn't received as much love for some reason. In the next two years I expect major updates and/or improvements. They should kill off Redshift.
- antonvs 1y ago> … going after the whale that is Databricks. Databricks is tiny compared to AWS, maybe 1/50th the revenue. But they’re both chasing a big and fast-growing market. I don’t think it’s so much that AWS is going after Databricks as that Databricks happens to be in a market that AWS is interested in.
- iknownothow 1y agoI agree, Databricks is one of many in the space. If S3 makes Databricks redundant, then they also make others like Databricks redundant too.
- physicsguy 1y agoThe biggest killer of vector dbs is that normal DBs can easily store embeddings, and the vector DBs just don’t then offer enough of a differentiator to be a separate product. We found our application was very sensitive to context aware chunking too. You don’t really get control of that in many tools.
- curtisszmania 1y ago[dead]
- vincirufus 1y agoThis could be game changing