9 ms·
It's interesting that folks go down this hacky route when they can use something like Vespa, which is orders of magnitude better from a performance, relevance,
by binarymax 3y ago
It's interesting that folks go down this hacky route when they can use something like Vespa, which is orders of magnitude better from a performance, relevance, scalability, and developer ergonomics perspective.
- deleted 3y ago[deleted]
- moomoo11 3y agoIs that a different system? Sorry I’m on spotty mobile that can’t open anything besides HN lol (God bless this website). Sometimes it is just easier to use the existing systems and squeeze them as much as possible. Especially when it’s a small team or solo without much $$
- binarymax 3y agoWhen it comes to search I cannot disagree more. https://vespa.ai https://vespa.ai is a purpose built search engine. If you start bolting search onto your database, your relevance will be terrible, you'll be rewriting a lot of table stakes tools/features from scratch, and your technical debt will skyrocket.
- hot_gril 3y agoOr it'll be good enough for whatever minimal search use case you have, and you upgrade to vespa (or whatever new thing) later when it's actually needed. If we jumped right to the most capable long-term solution for every feature we had, our systems would be nuts.
- LunaSea 3y agoThe advantage of pg_vector is that you don't need a second, specialised database and you also don't need to synchronise data. It makes much more operational sense to use pg_vector if your use case can be implemented tha way.
- binarymax 3y agoIt makes terrible operational sense. What are the HA/DR, sharding, replica, and backup strategies and tools for pg_vector? What are the embedding integration and relevance tools? What are the reindexing strategies? What are the scaling, caching, and CPU thrashing resolution paths? You're going to spend a bunch of time writing integrations that already exist for actual search engines, and you're going to be stuck and need to back out when search becomes a necessity rather than an afterthought.
- philipbjorge 3y agoWhat makes most operational sense is going to depend on your context. From my vantage point, you’re both right in the appropriate context.
- deleted 3y ago[deleted]
- whakim 3y agoWhat if you don't need those things yet and you just have some embeddings you want to query for cosine similarity? A dedicated vector database is way, way overkill for many people.
- djbusby 3y agoThe HA/DR, Sharing, Replica and Backup would all be the same as before. Its all in PG so you use the existing method. If you have two systems, then you have two (unique) answers for HA,DR,Shard,Replica,Backup - the PG set and the Vespa. That's more complicated, from an operational perspective. PG FTS is quite good, and there are in-pg methods that can improve it. And, from experience, when it's item to upscale to Solr/ES/etc it's not a very heavy lift.