11 ms·
Databases in 2025: A Year in Review
- A1aM0 9mo agoPavlo is right to be skeptical about MCP security. The entire philosophy of MCP seems to be about maximizing context availability for the model, which stands in direct opposition to the principle of Least Privilege. When you expose a database via a protocol designed for 'context', you aren't just exposing data; you're exposing the schema's complexity to an entity that handles ambiguity poorly. It feels like we're just reinventing SQL injection, but this time the injection comes from the system's own hallucinations rather than a malicious user.
- Miyamura80 9mo agoTotally agree, unfettered access to databases are dangerous There are ways to reduce injection risk since LLMs are stateless and thus you can monitor the origination and the trustworthiness of the context that enters the LLM and then decide if MCB actions that affect state will be dangerous or not We've implementeda mechanism like this based on Simon Willison's lethal trifecta framework as an MCP gateway monitoring what enters context. LMK if you have any feedback on this approach to MCP security. This is not as elegant as the approach that Pavlo talks about in the post, but nonetheless, we believe this is a good band-aid solution for the time bein,g as the technology matures https://github.com/Edison-Watch/open-edison https://github.com/Edison-Watch/open-edison
- quotemstr 9mo ago> Totally agree, unfettered access to databases are dangerous Any decent MVCC database should be able to provide an MCP access to a mutable yet isolated snapshot of the DB though, and it doesn't strike me as crazy to let the agent play with that.
- thesz 9mo agoFor this database has to have nested transactions, where COMMITs do propagate up one level and not to the actual database, and not many databases have them. Also, a double COMMIT may propagate changes outside of agent's playbox.
- quotemstr 9mo ago> For this database has to have nested transactions, where COMMITs do propagate up one level and not to the actual database, Correct, but nested transaction support doesn't seem that much of a reach if you're an MVCC-style system anyway (although you might have to factor out things like row watermarks to lookaside tables if you want to let them be branchy instead of XID being a write lock.) You could version the index B-tree nodes too.
- thesz 9mo ago> but nested transaction support doesn't seem that much of a reach if you're an MVCC-style system anyway You are talking about code that have to be written and tested. Also, do not forget about double COMMIT, intentional or not.
- SpaceL10n 9mo agoWas the trade-off so exciting that we abandoned our own principles? Or, are we lemmings? Edit: My apologies for the cynical take. I like to think that this is just the move fast break stuff ethos coming about.
- anthonypasq 9mo agoi dont know anyone with a brain that is using a DB mcp with write permissions in prod. i mean trying to lay that blame on a protocol for doing something as nuts as that seems unfair.
- nijave 9mo agoYes and no. Least privilege has existed in databases for a very long time. You need to implement correct DB privileges using user/roles, views, and other best practices. The MCP server is more like a dumb client in this setup. However, that's easy for people to forget and throw privileged creds at the MCP and hope for the best. The same stands for all LLM tools (MCP servers or otherwise). You always need to implement correct permissions in the tool--the LLM is too easily tricked and confused to enforce a permission boundary
- p2hari 9mo agoThe author mentions about it in the name change for edgeDb to gel. However, it could also have been added in the Acquisitions landscape. Gel joined vercel [1]. 1. https://www.geldata.com/blog/gel-joins-vercel https://www.geldata.com/blog/gel-joins-vercel
- djsjajah 9mo agoYou just ruined my day. The post makes it sound like gel is now dead. The post by Vercel does not give me much hope either [1]. Last commit on the gel repo was two weeks ago. [1] https://vercel.com/blog/investing-in-the-python-ecosystem https://vercel.com/blog/investing-in-the-python-ecosystem
- kaelwd 9mo agoFrom discord: > There has been a ton of interest expressed this week about potential community maintenance of Gel moving forward. To help organize and channel these hopes, I'm putting out a call for volunteers to join a Gel Community Fork Working Group (...GCFWG??). We are looking for 3-5 enthusiastic, trustworthy, and competent engineers to form a working group to create a "blessed" community-maintained fork of Gel. I would be available as an advisor to the WG, on a limited basis, in the beginning. > The goal would be to produce a fork with its own build and distribution infrastructure and a credible commitment to maintainership. If successful, we will link to the project from the old Gel repos before archiving them, and potentially make the final CLI release support upgrading to the community fork. > Applications accepted here: https://forms.gle/GcooC6ZDTjNRen939 https://forms.gle/GcooC6ZDTjNRen939 > I'll be reaching out to people about applications in January.
- wussboy 9mo agoI would think a two-week break over the holiday season wouldn't be a death knell.
- apavlo 9mo agoThanks for catching this. Updated: https://www.cs.cmu.edu/~pavlo/blog/2026/01/2025-databases-retrospective.html#erratum-gel https://www.cs.cmu.edu/~pavlo/blog/2026/01/2025-databases-re... I need to figure out an automatic way to track these.
- danielfalbo 9mo agoMaybe off-topic but, If you're not familiar with the CMU DB Group you might want to check out their eccentric teaching style [1]. I absolutely love their gangsta intros like [2] and pre-lecture dj sets like [3]. I also remember a video where he was lecturing with someone sleeping on the floor in the background for some reason. I can't find that video right now. Not too sure about the context or Andy's biography, I'll research that later, I'm even more curious now. [1] https://youtube.com/results?search_query=cmu+database https://youtube.com/results?search_query=cmu+database [2] https://youtu.be/dSxV5Sob5V8 https://youtu.be/dSxV5Sob5V8 [3] https://youtu.be/7NPIENPr-zk?t=85 https://youtu.be/7NPIENPr-zk?t=85
- sirfz 9mo agoIndeed, I was delighted when I read the part about wutang's time capsule and obviously OP is a wu-tang and general hip hop fan. The intro you shared is dope!
- deleted 9mo ago[deleted]
- znpy 9mo agoI can't understand if their "intro to database systems" is an introductory (undergrad) level course or some advanced course (as in, introduction to database (internals)). Anyone willing to clarify this? I'm quite weak at database stuff, i'd love to find some undergrad-level proper course to learn and catch up.
- Tostino 9mo agoIt's the internals. He is training up people to work on new features for existing databases, or build new ones. Not application developers on how to use a database. Knowing some of the internals can help application developers make better decisions when it comes to using databases though.
- Tostino 9mo agoHere is the playlist: https://www.youtube.com/playlist?list=PLSE8ODhjZXjYMAgsGH-GtY5rJYZ6zjsd5 https://www.youtube.com/playlist?list=PLSE8ODhjZXjYMAgsGH-Gt... You can tell from the topics, it's related to building databases, not using them.
- santiagobasulto 9mo agoI love these yearly review posts. Thanks Andy and team.
- shrx 9mo agoNothing about time series-oriented databases?
- speedgoose 9mo agoNot much happened I guess. Clickhouse has got an experimental time series engine : https://clickhouse.com/docs/engines/table-engines/special/time_series https://clickhouse.com/docs/engines/table-engines/special/ti...
- shrx 9mo agoQuestDB at least is gaining some popularity: https://questdb.com/ https://questdb.com/ I was hoping to learn about some new potentially viable alternatives to InfluxDB, alas it seems I'll continue using it for now.
- speedgoose 9mo agoI'm running an experimental side project where I doing some kind of glue between various time-series APIs and storage engines. For example it has an InfluxDB compatible ingestion API, so Telegraf can push its data to it or InfluxDB can replicate to it. It also has a Prometheus remote read and remote write API, so it's compatible with Prometheus. The storage can be done in various systems, including ClickHouse, SQLite, DuckDB, TimescaleDB… I should try to include QuestDB.
- apavlo 9mo ago> Nothing about time series-oriented databases? https://www.cs.cmu.edu/~pavlo/blog/2026/01/2025-databases-retrospective.html#you-didnt-read-this https://www.cs.cmu.edu/~pavlo/blog/2026/01/2025-databases-re...
- TekMol 9mo agoFrom my perspective on databases, two trends continued in 2025: 1: Moving everything to SQLite 2: Using mostly JSON fields Both started already a few years back and accelerated in 2025. SQLite is just so nice and easy to deal with, with its no-daemon, one-file-per-db and one-type-per value approach. And the JSON arrow functions make it a pleasure to work with flexible JSON data.
- delaminator 9mo agoFrom my perspective, everything's DuckDB. Single file per database, Multiple ingestion formats, full text search, S3 support, Parquet file support, columnar storage. fully typed. WASM version for full SQL in JavaScript.
- sanderjd 9mo agoThis is a funny thread to me because my frustration is at the intersection of your comments: I keep wanting sqlite for writes (and lookups) and duckdb for reads. Are you aware of anything that works like this?
- SchwKatze 9mo agoI think you could build an ETL-ish workflow where you use SQLite for OLTP and DuckDB for OLAP, but I suppose it's very workload dependent, there are several tradeoffs here.
- sanderjd 9mo agoRight. This is what I want, but transparently to the client. It seems fairly straightforward, but I keep looking for an existing implementation of it and haven't found one yet.
- nlittlepoole 9mo agoDuckDB can read/write SQLite files via extension. So you can do that now with DuckDB as is. https://duckdb.org/docs/stable/core_extensions/sqlite https://duckdb.org/docs/stable/core_extensions/sqlite
- maximgeorge 9mo ago[dead]
- pjmlp 9mo agoOver here, it is DB2, SQL Server or Oracle if using a plain RDMS, or whatever DB abstraction layer is provided on top of a SaaS product, where we get to query with some kind of ORM abstraction preventing raw SQL, or GraphQL, without knowing the implementation details.
- sandos 9mo agoThis sounds like a flashback to J2EE. Which I know is still alive and well. Banks, insurance companies and the tax agency do not much care for fancy new stuff, but that it works.
- pjmlp 9mo agoYep, Fortune 500 enterprise consulting, boring technology that pays the bills. Java, .NET, C++, nodejs, Sitecore, Adobe Experience Manager, Optimizely, SAP, Dynamics, headless CMSes,...
- chasd00 9mo agoI describe these techs like garbage trucks. No one likes to see them but they’re there every day doing a decent part of what it takes to hold society together hah.
- pjmlp 9mo agoScott Hanselman has a good term for all these kind of jobs, the dark matter developers. https://www.hanselman.com/blog/dark-matter-developers-the-unseen-99 https://www.hanselman.com/blog/dark-matter-developers-the-un...
- beders 9mo agoWhile the author mentions that he just doesn't have the time to look at all the databases, none of the reviews of the last few years mention immutable and/or bi-temporal databases. Which looks more like a blind spot to me honestly. This category of databases is just fantastic for industries like fintech. Two candidates are sticking out. https://xtdb.com/blog/launching-xtdb-v2 https://xtdb.com/blog/launching-xtdb-v2 (2025) https://blog.datomic.com/2023/04/datomic-is-free.html https://blog.datomic.com/2023/04/datomic-is-free.html (2023)
- anonymousDan 9mo agoWhy fintech specifically?
- postexitus 9mo agoBecause, money.
- falcor84 9mo agoI would assume that it's because in fintech it's more common than in other domains to want to revert a particular thread of transactions without touching others from the same time.
- postexitus 9mo agoNot only transactions - but state of the world.
- defo10 9mo agocompliance requirements mostly (same for health tech)
- groestl 9mo agoDestructive operations are both tempting to some devs and immensely problematic in that industry for regulatory purposes, so picking a tech that is inherently incapable of destructive operations is alluring, I suppose.
- backtogeek 9mo agoI can't believe that article has no mention of SQLite ??
- apavlo 9mo ago> I can't believe that article has no mention of SQLite ?? https://www.cs.cmu.edu/~pavlo/blog/2026/01/2025-databases-retrospective.html#you-didnt-read-this https://www.cs.cmu.edu/~pavlo/blog/2026/01/2025-databases-re...
- bob1029 9mo agoNo MSSQL, DB2 or Oracle either. Anything this proven & stable is probably not worth blogging about in this context. SQLite gets a lot of attention on HN but that's a bit of an exception.
- astrostl 9mo agoSame. CMD-F, 'sqlite', no hits, skip and go straight to comments.
- deleted 9mo ago[deleted]
- lvl155 9mo agoI want to thank Andy and the entire DB Group at CMU. They’ve done a great job of making database accessible to so many people. They are world class.
- techsystems 9mo agoWhat did they do?
- swyx 9mo agolook up the cmu db youtube
- gr4vityWall 9mo agoDidn't know MongoDB was suing the company behind FerretDB. That's disgusting.
- beembeem 9mo agoAndy has a balanced and appropriate take here.
- codeulike 9mo agoBarely any mention of Oracle or MS Sql Server, commonly reckoned to be #1 and #3 most used databases in the world https://db-engines.com/en/ranking https://db-engines.com/en/ranking
- qcnguy 9mo agoOracle is mentioned at the start, where he proclaims the "dominance" of Postgres and then admits its newest features have been in Oracle for nearly a quarter of a century already. The dominance he's talking about is only about how many startups raise how many millions from investors, not anything technical. And then of course at the end he has a whole section about Larry Ellison, like always.
- sanderjd 9mo agoIsn't it because it's about news, as in what's changing, rather than being about what's staying the same? He's a researcher, so his interests are always going to be more oriented toward new systems and new companies more than the big dominant systems.
- bzGoRust 9mo agoI would like to mention that vector databases like Milvus got lots of new features to support RAG, Agent development, features like BM25, hybrid search etc..
- jereze 9mo agoNo mention of DuckDB? Surprising.
- mariocesar 9mo agoSame surprise here. However in practice, the community tends to talk about DuckDB more like a client-side tool than a traditional database
- dujuku 9mo agoAlso somewhat surprised. DuckDB traction is impressive and on par with vector databases in their early phases. I think there's a good chance it will earn an honorable mention next year if adoption holds and becomes more mainstream. But my impression is that it's still early in its adoption curve where only those "in the know" are using it as a niche tool. It also still has some quirks and foot-guns that need moderately knowledgeable systems people to operate (e.g. it will happily OOM your DB)
- jimmar 9mo ago> "The Dominance of PostgreSQL Continues" It seems like the author is more focused on database features than user base. Every metric I can find online says that MySQL/MariaDB is more popular than PostgreSQL. PostgreSQL seems "better" (more features, better standards compliance) but MySQL/MariaDB works fine for many people. Am I living in a bubble?
- spprashant 9mo agoI think author is basing his observations on where the money is flowing. PostgreSQL adjacent startups and businesses are seeing a lot of investment.
- alexpadula 9mo agoWell yeah.
- apavlo 9mo ago> Am I living in a bubble? There are rumblings that the MySQL project is rudderless after Oracle fired the team working on the open-source project in September 2025. Oracle is putting all its energy in its closed-source MySQL Heatwave product. There is a new company that is looking to take over leadership of open-source MySQL but I can't talk about them yet. The MariaDB Corporation financial problems have also spooked companies and so more of them are looking to switch to Postgres.
- Sesse__ 9mo ago> There are rumblings that the MySQL project is rudderless after Oracle fired the team working on the open-source project in September 2025. Not just the open-source project; 80%+ (depending a bit on when you start counting) of the MySQL team as a whole was let go, and the SVP in charge of MySQL was, eh, “moving to another part of the org to spend more time with his family”. There was never really a separate “MySQL Community Edition team” that you could fire, although of course there were teams that worked mostly or entirely on projects that were not open-sourced.
- SchwKatze 9mo agoCan we even say that Anyblox is a file format? By my understanding of the project it's "just" a decoder for other file formats to solve the MxN problem.
- tiemster 9mo agoAlso emmer (which is perhaps too niche to get mentioned in an article like this), which I focuses more on being a quick/flexible 'data scratchpad', rather than just scale. https://hub.docker.com/r/tiemster/emmer https://hub.docker.com/r/tiemster/emmer
- furrball010 9mo agonice to see it get mentioned here :), I like using it also for scripts etc. Quite flexible because you can do everything with the api.
- divan 9mo ago> Acquisitions ... Gel → Vercel is a bit misleading. Gel (formerly EdgeDB) is sunsetting it's development. (extremely talented) Team joins Vercel to work on other stuff. That was a hard hit for me in December. I loved working with EdgeQL so much.
- senderista 9mo agoIt is a beautifully designed language and would make a great starting point for future DB projects.
- dmarwicke 9mo agowe had to restrict ours to views only because it kept trying to run updates. still breaks sometimes when it hallucinates column names but at least it can't do anything destructive
- cloutiertyler 9mo agoHow is SpacetimeDB not mentioned here?
- apavlo 9mo ago> How is SpacetimeDB not mentioned here? https://www.cs.cmu.edu/~pavlo/blog/2026/01/2025-databases-retrospective.html#you-didnt-read-this https://www.cs.cmu.edu/~pavlo/blog/2026/01/2025-databases-re...
- cryptica 9mo agoIt's so weird how everyone nowadays is using Postgres. It's not like end users can see your database. It's disturbing how everyone is gravitating towards the same tools. This started happening since React and kept getting worse. Software development sucks nowadays. All technical decisions about which tools to use are made by people who don't have to use the tools. There is no nuance anymore. There's a blanket solution for every problem and there isn't much to choose from. Meanwhile, software is less reliable than it's ever been. It's like a bad dream. Everything is bad and getting worse.
- da02 9mo agoWhich alternatives to PostgreSQL would you like to see get more attention?
- cryptica 9mo agoAll of them. Nothing wrong with Postgres, I like Postgres. But the more alternatives the better. My favorite database is RethinkDB but officially, it's a dead project. Unofficially it's still pretty great.
- esafak 9mo agoWhat's wrong this postgres?
- throw0101d 9mo agoRegarding distributed(-ish) Postgres, does anyone know if something like My/MariaSQL's multi-master Galera† is around for Pg: > MariaDB Galera Cluster provides a synchronous replication system that uses an approach often called eager replication. In this model, nodes in a cluster synchronize with all other nodes by applying replicated updates as a single transaction. This means that when a transaction COMMITs, all nodes in the cluster have the same value. This process is accomplished using write-set replication through a group communication framework. * https://mariadb.com/docs/galera-cluster/galera-architecture/introduction-to-galera-architecture https://mariadb.com/docs/galera-cluster/galera-architecture/... This isn't necessarily about being "web scale", but having a first-party, fairly-automated replication solution would make HA easier for a number internal-only stuff much simpler. † Yes, I am aware: https://aphyr.com/posts/327-jepsen-mariadb-galera-cluster https://aphyr.com/posts/327-jepsen-mariadb-galera-cluster
- nijave 9mo agoCitus, sort of Cockroach For HA, Patroni, stolon, CNPG Multimaster doesn't necessarily buy you availability. Usually it trades performance and potentially uptime for data integrity.
- senderista 9mo agoAnd DSQL if you use a very loose definition of "Postgres"
- npalli 9mo agoAndy is probably the only person who adores Larry Ellison (Oracle) unironically.
- zjaffee 9mo agoWhat an amazing set of articles, one thing that I think he's missed is the clear multi year trends. Over the past 5 years there's been significant changes and several clear winners. Databricks and Snowflake have really demonstrated ability to stay resilient despite strong competition from cloud providers themselves, often through the privatization of what previously was open source. This is especially relevant given also the articles mentioning of how cloudera and hortonworks failed to make it. I also think the quiet execution of databases like clickhouse have shown to be extremely impressive and have filled a niche that wasn't previously filled by an obvious solution.
- quotemstr 9mo agoWhy does "database" is surveys like this not include DuckDB and SQLite, which are great [1] embedded answers to Clickhouse and PostgreSQL. Both are excellent and useful databases; DuckDB's reasonable syntax, fast vectorized everything, and support for ingesting the hairiest of data as in-DB ETL make me reach for it first these days, at least for the things I want to do. Why is it that in "I'm a serious database person" circles, the popular embedded databases don't count? [1] Yes, I know it's not an exact comparison.
- shekispeaks 9mo agoTiDB has gained some momentum in silicon valley with companies looking to adopt it. Does he have any commentary on TiDB which is an OLTP and OLAP hybrid?
- felipelalli 9mo agoI think it's time for a big move towards immutable databases that weren't even mentioned in this article. I've already worked with Datomic and immudb: Datomic is very good, but extremely complex and exotic, difficult learning curve to achieve perfect tuning. immudb is definitely not ready for production and starts having problems with mere hundreds of thousands of records. There's nothing too serious yet.
- andersmurphy 9mo agoWith a trend towards immutable single writer databases MMAP seems like a massive win.
- mtndew4brkfst 9mo agoAndy is very critical of using mmap in database implementations.
- andersmurphy 9mo agoWhy? Sqlite and LMDB make fantastic use of it. For anyone doing a single writer db it's a no brainer. It does so much for you and it does it very well. All the things you don't have to implement because it does it for you: - Reading the data from disk - Concurrency between different threads reading the same data - Caching and buffer management - Eviction of pages from memory - Playing nice with other processes in the machine Why would you not leverage it? It's such a great fit for scaling reads.
- alexpadula 9mo ago“ It's such a great fit for scaling reads.” And losing them.
- andersmurphy 9mo agoHow so? LMDB, boltdb/bbolt and sqlite (with mmap) are all rock solid. Just because mongodb used mmap badly does not make it any less valuable.
- cmrdporcupine 9mo agoThe strongest argument as far as I can see it is... the problem is you now lose control over all those things. It's a black box with effectively no knobs. Anyways, read for yourself, Pavlo & Leis get into it in detail, and there's benchmarks: https://db.cs.cmu.edu/papers/2022/cidr2022-p13-crotty.pdf https://db.cs.cmu.edu/papers/2022/cidr2022-p13-crotty.pdf https://db.cs.cmu.edu/mmap-cidr2022/ https://db.cs.cmu.edu/mmap-cidr2022/
- ComputerGuru 9mo agoPg18 is an absolutely fantastic release. Everyone flaks about the async IO worker support, but there’s so much more. Builtin Unicode locales, unique indexes/constraints/fks that can be added in unvalidated state, generated virtual (expression) columns, skip scans on btree indexes (absolutely huge), uuidv7 support, and so much more.
- thesurlydev 9mo agoSupabase seems to be killing it. I read somewhere they are used by ~70% of YCombinator startups. I wonder how many of those eventually move to self-hosted.
- alexpadula 9mo agoBeen reading these for a few years. I enjoy them, thank you Andy. I hope you’re doing better.
- qinchencq 9mo agoWas hoping to read about graph database, AI-related changes..., but didn't expect this: "I almost died in the spring semester...surprisingly hard to concentrate on important things like databases when you can't breathe." Hope Prof. Pavlo has been breathing better, stellar review.
- cluckindan 9mo ago”I still haven't met anybody who is actively using Dgraph.” That’s because it is mostly used in national security and military applications in several countries.
- ugamarkj 9mo agoInteresting article and I've enjoyed reading all the comments. The focus on Postgres is fascinating to me. From an analytics database perspective, we've had great success with Exasol, which isn't so well known in the US. It is very low overhead (no index management) and extremely fast and scalable. They have a free version as well as a licensed MPP version -- cloud hosted or on-prem. It is a blank slate, but it can do all the things.
- budapest05 9mo agoFor analytics workloads, Exasol is a great choice: high performance & MPP scale. They offer a free Personal Edition for download and testing (cloud and on‑prem options available): https://downloads.exasol.com/exasol-personal https://downloads.exasol.com/exasol-personal worth a quick benchmark with your own data.