7 ms·
>I don't find it surprising that so many distributed databases fail at ensuring consistency; it's a very hard problem It's not hard — it's slow. Sending reads
by acconsta 11y ago
>I don't find it surprising that so many distributed databases fail at ensuring consistency; it's a very hard problem
It's not hard — it's slow. Sending reads and writes through Paxos or Raft gives you sequential consistency. But not surprisingly, touching a quorum of nodes for every operation is too slow to be practical for many workloads.
And that's usually fine — most data aren't bank account balances.
- simoncion 11y agoThough, it's not even remotely okay if your documentation and marketing material indicates that consistency isn't being sacrificed for performance. :) I suspect that no experienced programmer expects to get data safety for free. We do expect that the documentation doesn't lie to us when describing the level of data safety provided.
- antirez 11y agoI totally agree, just a note: write safety and consistency are not the same thing. An example is a CP store that breaks linearizability in order to serve reads without asking other nodes, but just providing the local value. Writes are as safe as a true CP store, but the consistency model is different and weaker. It is true that many times the replication model employed also has effects on the ability of writes to survive, I just mentioned the above example to show it is not always the case.
- acconsta 11y agoRight, that's the real value here — holding databases accountable for their claims. Zookeeper clearly documented that writes are consistent but reads aren't, so it "passed" Jepsen cleanly. The Etcd guys claimed reads are linearizable, so Aphyr tested for linearizability and it failed. https://aphyr.com/posts/291-call-me-maybe-zookeeper https://aphyr.com/posts/291-call-me-maybe-zookeeper https://aphyr.com/posts/316-call-me-maybe-etcd-and-consul https://aphyr.com/posts/316-call-me-maybe-etcd-and-consul
- deleted 11y ago[deleted]
- nemo44x 11y agoI think you touched on a key point a lot of people don't consider. Not all data is the same nor does it have to be treated the same. Many distributed systems will choke on a piece of data here and there and may not give the appropriate response. But you need to consider what you're working on and if that's OK from time to time given the advantages the system you're considering offers. What might be considered an appropriate data store for one data set may not be for another. And the type of information a lot of distributed systems are handling are far from bank accounts, for example. And if you're handling ultra important, sensitive data there are techniques that have been available for many years (two phase commit, for example) that can help. I love this series but I'm blown away by how many engineers here automatically assume a system is a failure because it doesn't pass a certain type of test from time to time. I do agree with the series that the marketing materials shouldn't claim things, however. Companies should be honest with what their systems weaknesses currently are.
- superuser2 11y agoBanks are eventually consistent. Online transaction processing creates "pending" transactions, and the data is often inconsistent. Your charge may exist in the merchant's database but not post to your online banking for several hours. Or it may be a wildly different amount - i.e. gas stations will place a $100 hold on your debit card and it will stay that way for days until it's settled for the actual (lesser) amount. If you were accidentally double-charged, than rather than processing a separate refund, the merchant may simply not settle the duplicate charge, and it will drop off your pending transactions... eventually. The lag time may be several days or weeks. If it's a debit card, you can't spend the money during that time and you may be temporarily broke because of it. If you make an ACH transfer, money will disappear from your account one night, spend a business day in the aether, and then post to the recipient's account on the third day. The system is in inconsistent state (i.e. money is missing) for at least a day, possibly a whole weekend. The actual transaction is settled and goes into "posted" state with a lag of 2.5 * 10^8 ms - i.e. 3 business days. That's if you're lucky. Banks do need strong consistency, but not in anything approaching realtime. Even ZooKeeper could probably handle the U.S.'s financial transactions faster than current infrastructure.
- baudehlo 11y agoAnd what would shock you most is some of the methods used to send those transactions (eg ftping a text file, and if it gets corrupt, someone opens it in vi to fix it - I have a friend who used to do exactly that).
- kalleboo 11y agoFor anyone who wants to read more, this blog post series goes into quite a lot of detail on ACH (including the FTPing and the fixed-field file format) http://engineering.zenpayroll.com/how-ach-works-a-developer-perspective-part-4/ http://engineering.zenpayroll.com/how-ach-works-a-developer-...
- z92 11y agoLooks like ACID is used by everyone, but banks.
- 11y ago