10 ms·
MongoDB acquires Voyage AI
- schnebbau 2y agoHow does MongoDB still have that much available to spend? Everyone I know moved off it years ago.
- geodel 2y agoEveryone you know put a dollar in donation basket while moving off. Mongo collected all and brought Voyage AI
- porridgeraisin 2y agoThere are a lot of people still on it, including the place I worked at last. It was starting to get expensive though, so we were experimenting with other document stores (dynamodb was being trialled, since we were already AWS for most things, just around the time I left)
- bithavoc 2y agothat’s what I thought, but every single candidate I interviewed mentioned MongoDB as their recent reference document database, I asked the last candidate if they were self-hosting, the answer is no, they used MongoDB cloud.
- winrid 2y agoI self host a handful of mongodb deployments for personal projects and manage self hosted mongo deployments of almost a hundred nodes for some companies. Atlas can get very expensive if you need good IO.
- slt2021 2y agoif you a developer you wanna use MongoDB as database, not be MongoDB SRE and DBA thats the reason for using Atlas
- skatanski 2y agoPrecisely, and if you are enterprise, you want to have an option to request priority support and have a lot of features out of the box. Also some of the search features are only available in Atlas unfortunately.
- rpep 2y agoYou cant use the embeddings/vector search stuff this refers to in self hosted anyway, it’s only implemented in their Atlas Cloud product. It makes it a real PITA to test locally. The Atlas Dev local container didn’t work the same when I tried it earlier in the year.
- tiltowait 2y agoAtlas has a generous free tier that is great for hobby projects.
- isoprophlex 2y agoPretty sure they achieved fiscal nirvana by exploiting enterprise brain rot. You hook em, they accumulate tech debt for years, all their devs leave, now they can't move away & you can start increasing prices. Eventually the empty husk will topple over but that's still years away.
- dimgl 2y agoIs it possible that they simply have a good product?
- isoprophlex 2y agoImpossible! It's not based on sqlite, postgres or written in rust, so it must be terrible!
- flessner 2y agoI never understood this argument, there are many great products running on Java, PHP, Ruby, JavaScript... All of these languages have a "crowd" that hates them for historic and/or esoteric reasons. Great products are in my opinion a function of skill and care. The only benefit a "popular" tool or language gets you is a greater developer pool for hiring.
- vosper 2y agoThey do have a good product, but "they accumulate tech debt for years, all their devs leave, now they can't move away" is the story of the place I worked at a few years ago. The database was such a disorganized, inconsistent mess that no-one had the stomach (or budget) to try and get off it.
- codr7 2y agoBasically every MongoDB I've come across; same story, different faces.
- dkjaudyeqooe 2y ago
- DarmokJalad1701 2y agoBecause they are web-scale obviously.
- deleted 2y ago[deleted]
- Cshelton 2y agoWe use it a lot for a specific use-case and it works great. Mongo has come a long long way since the release over a decade ago, and if you keep it in Majority Read and Write, it's very reliable. Also, on some things, it allows us to pivot much faster. And now with the help of LLMs, writing "Aggregation Pipelines" are very fast.
- burningion 2y agoI've been using Mongo while developing some analysis / retrieval systems around video, and this is the correct answer. Aggregation pipelines allow me to do really powerful search around amorphous / changing data. Adding a way to automatically update / recalculate embeddings to your database makes even more sense.
- magarnicle 2y agoDo you have any tricks for writing and debugging pipelines? I feel like there are so many little hiccups that I spend ages figuring out if that one field name needs a $ or not.
- codr7 2y agoPretending a pile of json is a database is great for pivoting, not so great for anything else. Maintaining apps built on MongoDB is soul killing.
- rudolph9 2y agoThis
- SilasX 2y agoWell, it's referred to as a cash-and-stock deal but I can't find any more detail about how much is stock: https://seekingalpha.com/news/4412466-mongodb-acquires-voyage-ai-for-220m-to-boost-ai-search https://seekingalpha.com/news/4412466-mongodb-acquires-voyag...
- mgfist 2y ago$2.3B in cash as of last quarter
- yfontana 2y agoThis may be a shock to many HN readers, but MongoDB's revenue has been growing quite fast in the last few years (from 400M in 2020 to 1.7B in 2024). They've been pushing Atlas pretty hard in the Enterprise world. Have no experience with it myself, but I've heard some decently positive things about it (ease of set up and maintenance, reliability).
- yla92 2y agoMongo Atlas (their cloud offering) is really solid (and expensive)!
- paxys 2y agoMongoDB is a public company. Its quarterly financial reports will give you a much more accurate picture of the company's health than "everyone you know".
- aitchnyu 2y agoAre they profitable, and at which point in time? How good of an investment was it? Sorry, my eyes were swimming in their financial report hosted in their domain.
- thecleaner 2y agoCurious - do people migrate due to the price tag, ease of use, sth else ?
- ChrisArchitect 2y agoVoyage AI post: https://blog.voyageai.com/2025/02/24/joining-mongodb/ https://blog.voyageai.com/2025/02/24/joining-mongodb/
- BlairCurrey 2y agoand the mongo blog post for how it will be used: https://www.mongodb.com/blog/post/redefining-database-ai-why-mongodb-acquired-voyage-ai https://www.mongodb.com/blog/post/redefining-database-ai-why...
- infecto 2y agoOnly skimmed through the release..I hope they continue supporting the API but it comes with a little higher confidence that the company behind it is not collecting all your data. Voyage has some interesting embedding models that I have been hesitant to fully utilize due to the lack of confidence in the startup behind it.
- kaycebasques 2y agoThis blog post outlines the new roadmap: https://www.mongodb.com/blog/post/redefining-database-ai-why-mongodb-acquired-voyage-ai https://www.mongodb.com/blog/post/redefining-database-ai-why...
- __jl__ 2y agoThey commit to supporting the API in step 1 but it's not entirely clear to me whether that commitment continues with step 2-3...
- Beefin 2y agowhat's the calculus here? if i'm a developer choosing a low-level primitive such as a database, i'm likely quite opinionated on which models i use.
- crowcroft 2y agoIf I had to guess they might see embedding models become small and optimised enough to the point that they can pull them into the DB layer as a feature instead of being something devs need to actively think about and build into their app. Or it could just be an expansion to their cloud offering. In a lot of cases embedding models just need to be 'good enough' and cheap and/or convenient is a winning GTM approach.
- cpursley 2y agoHow is MongoDB still a thing when there's already several ways to handle json in Postgres including Microsofts new documentdb extension: https://gist.github.com/cpursley/c8fb81fe8a7e5df038158bdfe0f06dbb#nosql https://gist.github.com/cpursley/c8fb81fe8a7e5df038158bdfe0f... What am I missing? Are Mongo users simply front end folks who didn't have time to learn basic SQL or back end architecture?
- frankfrank13 2y agoEnterprise sales
- amazingamazing 2y agoMongoDB is not the same as Postgres and jsonb. also, I'd challenge your thinking - ultimately the goal is to solve problems. you don't necessarily need SQL, or relations for that matter. that being said, naively modeling your stuff in mongodb (or other things like dynamodb) will cause you severe pain... what's also true, which people forget, is naively modeling your stuff with a relational database will also cause you pain. as they sometimes say, normalize until it hurts, and then denormalize to scale and make it work the amount of places I've seen that skip the second part and have extremely normalized databases makes me cringe. it's like people think joins are free...
- pphysch 2y agoThen your implementation can be as simple as CREATE TABLE documents (content JSONB);. But I suspect a PK and some metadata columns like timestamps will come in handy.
- amazingamazing 2y agosigh - mongoDB is not the same as creating a table with jsonb. for one, you don't have to deal with handling connections. that being said, postgres is great, but it's not the same.
- 2y ago
- hartator 2y agoI rather them focus on performance. Last MongoDB is still slower than MongoDB 3.4. An almost 10-year old release. For both reads and writes.
- amazingamazing 2y agomongodb had consistency issues before v5 if I recall, so take that for what it's worth.
- memco 2y agoCan you share more details about the conditions under which it is slow in recent versions? We moved from 3.x to 7 for our main database and after adding a few indexes we were missing we have seen at least an order of magnitude speed up.
- hartator 2y agoMost regular inserts and regular selects: https://medium.com/serpapi/mongodb-benchmark-3-4-vs-4-4-vs-5-0-vs-6-0-cb65146ae5ee https://medium.com/serpapi/mongodb-benchmark-3-4-vs-4-4-vs-5... We have internally a benchmark with MongoDB 8.x, but same pattern of disappointing results.
- winrid 2y agoAs someone that has ran every version from 3.2 to 8 on small nodes and large clusters (~100+ nodes)... 8 is waaay faster in the real world. It's not really comparable. Your micro benchmark is comparing the few nanoseconds of the heavier query planner, but in the real world that query planner gives real benefits. Not to mention aggregations, memory management improvements, and improvements when your working set size is very large/larger than memory.
- hartator 2y agoCan you share some data about this? Here's another dataset about performance regression doing `$inc`s as fast as possible on the same object. Mongo 3.4.24: 332,037 stats update in 100s. (3,321 stats updates per s) Mongo 8.0.4: 287,553 stats update in 100s. (2,876 stats updates per s) (higher is better)
- kaycebasques 2y agoBloomberg says it was a $220M cash & stock deal: https://www.bloomberg.com/news/articles/2025-02-24/mongodb-buys-voyage-ai-for-220-million-to-bolster-ai-search https://www.bloomberg.com/news/articles/2025-02-24/mongodb-b...
- markus_zhang 2y agoLooks like everyone is jumping into the AI game. Is there a bubble?
- codr7 2y agoWhatever respect I had left for MongoDB just went out the window, the last thing I want in my database is AI.
- eudoxus 2y agoAside from the MongoDB of it all, wheres the issue with adding "AI" here? As I understand it this is just vector types, similarity searches, embedding indexes, and RAG capabilities. All of which are just data storage/retrieval mechanics and custom types. This isn't adding some omnipotent AI agent to run/manage/optimize your DB or otherwise turn it into some blackbox gizmo.
- codr7 2y agoOh I'm pretty sure that will be the next step, given the direction we're moving in and the lack of common sense and responsibility on display. I see GenAI as a stop gap solution at best, not really optimal for any problems; and AGI is a major distraction from finding good solutions to important problems. The wild goose chase to apply GenAI to everything has serious consequences. People are so excited about the fact that a computer can sort of drive a car that they don't even stop to consider that a human driver that randomly fails the same way would never get a license, and rightly so. So excited about the fact that a computer can sort of write functional code that they don't stop to consider that any human developer that fails randomly the same way would never get a job, and rightly so. We're already applying it to weapons/warfare, which is obviously a very bad idea. I'm sure the technology will improve, but never to the point where it's reliable. It will fail randomly less often, but the magnitude of its failures isn't going anywhere.
- whalesalad 2y agolol dude where have you been
- htrp 2y agoVoyage AI basically builds embedding models for vector search
- crowcroft 2y agoYou don't hear the big AI providers talk about embeddings much, but I have to believe in the long run that companies building SOTA foundational LLMs are going to ultimately have the best embedding models. Unless you can get to a point where you can make these models small enough that basically sit in the DB layer of an application...
- htrp 2y agoThat and because the embedding models are much easier to improve with at scale usage (hence why everyone has a deep search/research/RAG tool built into their AI web app now).
- rich_stable 2y agoThis is essentially my prediction; either that or something functionally equivalent.
- connectsnk 2y agoI understand the criticisms, but in my experience, MongoDB has come a long way. Many of the earlier issues people mention have been addressed. Features like sharding, built-in replication, and flexible schemas have made scaling large datasets much smoother for me. It’s not perfect, but it’s a solid choice.
- beoberha 2y agoI think the amount of people working on large enterprise systems here is a lot smaller than one would think. Whenever a fly.io post about sqlite ends up in here, there are a scary amount of comments about using sqlite in way more scenarios than it should be.
- connectsnk 2y agoTrue. I have that feeling many times that the enterprise crowd doesnt visits hacker news.
- dkjaudyeqooe 2y agoWhy would they be here? They use Oracle. But mainly because management hasn't worked out how to cancel their licenses without breaking their budgets.
- koakuma-chan 2y agoWhy would I use anything other than sqlite?
- margalabargala 2y agoEasy. Sometimes it's more than you need, and there's no reason to use sqlite when you can just write things to a flat text file that you can `grep` against.
- koakuma-chan 2y ago
- moralestapia 2y ago10x exit in a couple years, quite nice on the VC side! On the tech side ... no idea what Mongo's plan is ... their embedding model is not SOTA, does not even outperform the open ones out there, and reranking is a dead end in 2025. I think the value is on Voyage's team, their user base and having a vision that aligned with Mongo's. Congrats!
- serjester 2y agoWe've benchmarked a ton of the open models and voyage dramatically outperforms them. I think MTEB is a bad benchmark.
- touche_bag 2y agoInteresting take. Have you benchmarked models on your own data? Cause at this point everything is contaminated so I find it impossible to tell what proper sota is. Also - most folks still just use openai. Last time I checked, reranking always performs better than pure vector search. And to my knowledge it's still the superior fusion method for keyword and vector results.
- moralestapia 2y agoIn my experience, storing RAG chunks with a little bit of context helps a lot when doing the retrieval, then you can skip the whole "rerank" bit and halve your cost and latency. With embedding/generative models becoming better with time, the need for a rerank step will be optimized away.
- touche_bag 2y agoHuh? Rerank is always a boost on top of retrieval. So regardless of the chunking method or model you use, reranking with a good model will always result in higher MRR. And improvements in embedding models also will never solve the problem of merging lexical and vector search results. Rank/score fusion are flawed since both are hardly comparable and boosting only works sometimes. Whereas rerankers generally do a pretty good job at this. Performance is indeed the biggest issue here. Rerankers are slow as hell and simply not feasible for some use cases.
- lpapez 2y agoIs Voyage AI web-scale yet?
- jamesrr39 2y agoGenuine question: I appreciate the comments about MongoDB being much better than it was 10 years ago; but Postgres is also much better today than then as well. What situations is Mongo better than Postgres? Why choose Mongo in 2025?
- 999900000999 2y agoSimple. Postgres is hard, you have to learn SQL. SQL is hard and mean. Mongo means we can just dump everyone into a magic box and worry about it later.No tables to create. But their is little time, we need to ship our CRUD APP NOW! No one on the team knows SQL! I'm actually using Postgres via Supabase for my current project, but I would probably never use straight up Postgres.
- chpatrick 2y agoEven as a JSON document store I'd rather use postgres with a jsonb column.
- tiltowait 2y agoWhy is that? I found Postgres's JSONB a pill to work with beyond trivial SELECTs, and even those were less ergonomic than Mongo.
- chpatrick 2y agoBecause you get the convenience of having a document store with a schema defined outside of the DB if you want it, along with the strong guarantees and semantics of SQL.
- chpatrick 2y agoFor example: let's say you had a CRM. You want to use foreign keys, transactions, all the classic SQL stuff to manage who can edit a post, when it was made, and other important metadata. But the hierarchical stuff representing the actual post is stored in JSON and interpreted by the backend.
- rich_stable 2y agoFor context, I am a startup founder, an Atlas user, and not what anyone would call a "major account". I'm also in my early 30's, so I wasn't around for the whole "web scale" meme era of MongoDB. I have personally been incredibly impressed with the way MongoDB has "looked out" for my company over the past year. I'll try to be satisfactorily specific their with privacy in mind, so this may come out a bit fuzzy. Their technical teams, both overseas and US, have produced some of the most thorough, detailed recommendations I have ever seen and their communication/followup was excellent. I've run many designs and ideas by their team, out of habit at this point, and have always been pleased with the response. They remember who I am. All of this is really unusual for a company at my growth stage. My use case requires full technical depth on text searching and vectorization; I use every aspect of Atlas Search that is available. A downside of building "bleeding edge" is that my tooling needs always seem to be just inches beyond what is available, so just about every release seems to have something that is "for me." It's hard to say if my feedback specifically has an impact on their roadmap - but they really do seem to build things I want. I think they reported ~50% better performance on bulkWrite() in the 8 release, but it was closer to 500% for my use case. Speaking of, this acquisition is like providence for me, because I've shared my various solutions with them for synchronously vectorizing "stuff" for use with local LLMs. It's a reasonably hard technical problem without a lot of consensus on standards; I think a lot of people believe there are standards, yet any discussion will quickly devolve into something like the "postgres/mongo" conversations you see here (I won't be visiting that topic). I strongly agree with the "this should be a database level feature" approach they are taking here; that's certainly how my brain wants to think about it and currently I have to do a lot of "glue"-ing to make it work the way I require. I hope they win.
- hodgesrm 2y agoSo basically you see this as helping to get working vector search in MongoDB? It sounds as if the attraction, then, is that it integrates easily with your existing Atlas usage. Or is there more?
- rich_stable 2y agoThere is more. They explain it better than I could in the roadmap links.
- redskyluan 2y agoI'm actually very disappointed about the performance of Mongo vector search after I test on it. Any vector database, is better than mongodb performance wise.