5 ms·
"We need to keep track of our customers’ account balances. We need them to trust that their Wave balance is correct and Wave won’t lose the money." "At Wave, t
by _vvhw 5y ago
"We need to keep track of our customers’ account balances. We need them to trust that their Wave balance is correct and Wave won’t lose the money."
"At Wave, though, we prefer to use boring technology, and a simple relational database like Postgres does an equally good job at this."
How does Wave ensure multi-region durability for Postgres, high availability, and strict serializability during failover, without risk of split-brain in the presence of network partitions?
I'm also interested in Wave's storage fault model with regards to Postgres? How are misdirected reads/writes, lost reads/writes, corruption in the middle of the WAL, and corruption at the end of the WAL handled? To ensure that account balances and transactions are never lost? As far as I understand, Postgres does not provide too many guarantees when it comes to interesting disk failures?
I don't mean to advocate for cryptocurrency in any way.
I'm genuinely curious, because I would have thought that a more modern distributed database that guarantees strict serializability in the face of network partitions would be the more obvious choice for tracking account balances safely?
- lkrubner 5y ago> multi-region durability Why is this important? Why can't they get by with a traditional main server plus some backups with a sync, plus failover when the main server goes down? It seems to me that all of your questions are issues that have been often discussed over the last 20 years, and they all have well known answers, except perhaps "multi-region durability". Leave that out, and you've well known answers for your other questions. And it seems to me, what they are saying is that they found it easier to go with a boring and traditional setup, that is, a boring and traditional setup is easier than any other setup.
- _vvhw 5y ago"Why is this important?" Because without multi-region or multi-AZ durability, Wave could lose balances. For example, in the event of a DC fire. "Why can't they get by with a traditional main server plus some backups with a sync" Because without distributed strict serializability, Wave would also be risking the loss of data back to the last backup. However, I'm assuming here that Wave are running consensus around Postgres. That they're not taking any of these chances. But if they're doing it right, then Postgres just doesn't seem like the most boring way to get distributed failover safely, especially when the data is as valuable as account balances, where so much compliance is at stake. Why not simply use a distributed database to begin with?
- knorker 5y agoNot losing transactions with a database is not exactly terra incognita. > Because without distributed strict serializability Depends what you mean by "distributed". A simple 2-phase commit would do it. Or hell, just a write-ahead log on the application layer. > Postgres just doesn't seem like the best way to get distributed failover safely That may be so. But are you arguing in favour of blockchain? When was the last time a real bank forgot your bank balance? They never used blockchains for this.
- _vvhw 5y ago"But are you arguing in favour of blockchain?" No, please see my original comment, where I made this clear: "I don't mean to advocate for cryptocurrency in any way." By "modern distributed database", I don't mean blockchain. I just mean "modern distributed database", e.g. things like FoundationDB, Spanner, Aurora or CockroachDB. "A simple 2-phase commit would do it. Or hell, just a write-ahead log on the application layer." Of course, and that leads into my second question, also in my original comment: How are storage faults in the middle of the committed WAL handled? Or is the rest of the committed WAL simply discarded, conflated with a torn write from a system crash? These questions are important when it comes to storing balances safely. Banks can use all kinds of techniques for defense-in-depth (not to suggest they do), they're not limited to Postgres. But a distributed database is a good thing these days, no?
- knorker 5y agoGotcha. I'm also excited about the modern distributed databases. They're pretty awesome. But yeah, a HN comment is not the right place to design a resilient DB setup. Like I said it's not terra incognita, but it's also not just "spin up a postgres".
- _vvhw 5y ago"Gotcha. I'm also excited about the modern distributed databases. They're pretty awesome." Awesome! "HN comment is not the right place to design a resilient DB setup." Why not? In my experience, HN is always the place to talk about building stuff, running stuff, and how to do that better. :)
- deleted 5y ago[deleted]
- kleinsch 5y agoYou seem to be vastly overthinking this. People have been building applications (including banking applications) using boring old relational databases for decades without losing data. Single master, replicated to another region. If relational databases were as faulty as you describe, every web app you use would be losing data all the time. Not everything has to be an overthought system design interview question.
- _vvhw 5y agoSingle DB corruption and disk failures are more common than you think. "Single master, replicated to another region." Why do you recommend against using a modern distributed database? Is it that hard to spin up something like CockroachDB? In fact, it wouldn't surprise me that Wave are running CockroachDB as their Postgres foundation under the hood.
- lkrubner 5y ago> Why do you recommend against using a modern distributed database? Call me maybe. This is a newer technology and the growing pains have been dramatic. I assume you've followed Jespen? https://aphyr.com/tags/jepsen https://aphyr.com/tags/jepsen I use MongoDB all the time, for a lot of projects. But I'm also keenly aware of how many years they have struggled to pass the Jespen tests. I can also recall several times engineers from MongoDB have said "We have finally resolved all of the issues raised by Jespen" and then another bug was discovered. Again, I use MongoDB all the time. But I'm aware it is simpler and safer to use Postgres, for many, many use cases.
- _vvhw 5y ago"I assume you've followed Jespen?" Yes! You might also want to check out autonomous deterministic testing, which is like Jepsen, but where you can explore many more state spaces by speeding up time, and where you can replay any distributed bugs instantly, just by replaying a seed. This is a brilliant talk on these techniques: FoundationDB — How I Learned to Stop Worrying and Trust the Database — https://www.youtube.com/watch?v=OJb8A6h9jQQ https://www.youtube.com/watch?v=OJb8A6h9jQQ I also work on a distributed database called TigerBeetle [1] where we use these new techniques. You can run our simulation tests yourself even. It's as simple as cloning the repo and running "scripts/vopr.sh", and let me know if you want a tour! These testing techniques work so well that we've been running a bug bounty challenge with awards up to $8192 for anyone who can find a way to break it or crash it. It comes with the simulation testing all included so it's pretty easy to get started. [2] [1] https://github.com/coilhq/tigerbeetle https://github.com/coilhq/tigerbeetle [2] https://github.com/coilhq/viewstamped-replication-made-famous https://github.com/coilhq/viewstamped-replication-made-famou...