5 ms·
Clocks Are Bad, Or, Welcome to the World of Distributed Systems
- neolefty 13y agoHow do other distributed databases handle this?
- krilnon 13y agoGoogle's Spanner [1] uses something it calls TrueTime: "The key enabler of these properties is a new TrueTime API and its implementation. The API directly exposes clock uncertainty, and the guarantees on Spanner’s timestamps depend on the bounds that the implementation provides. If the uncertainty is large, Spanner slows down to wait out that uncertainty. Google’s cluster-management software provides an implementation of the TrueTime API. This implementation keeps uncertainty small (generally less than 10ms) by using multiple modern clock references (GPS and atomic clocks)." [1] Spanner: Google's globally-distributed database https://www.usenix.org/system/files/conference/osdi12/osdi12-final-16.pdf https://www.usenix.org/system/files/conference/osdi12/osdi12...
- theatrus2 13y agoWith TrueTime you are trading some latency on concurrent operations for correctness. Other structures such as CRDTs/lattices might be more appropriate for your use case.
- madhusudancs 13y agoBy correctness you mean consistency? You don't have to be consistent all the time, i.e. you can trade consistency, but never correctness. If we could have traded correctness, we could have optimized everything and gone home by now :)
- hendzen 13y agoUnfortunately, not in a particularly clever way. CP systems such as MongoDB, HBase, etc. don't have this problem since each datum has an authoritative master. As you can imagine, this can result in some operational...unpleasantness due to the lack of liveness guarantees in the presence of a network partition. Out of the well known open-source AP systems, Riak is probably the leader here since they implement well understood techniques from the literature such as CRDTs and vclocks. EDIT: removed my statement about Cassandra since it was a bit misleading and jbellis answered above in greater detail.
- jbellis 13y agoCassandra offers a mix of commutative operations (sets, maps, increments), an eventlog model, and lightweight (paxos-based) transactions. Unlike a key/value database like Riak, Cassandra can update individual fields of a row or document independently, which simplifies things enormously. http://www.datastax.com/dev/blog/why-cassandra-doesnt-need-vector-clocks http://www.datastax.com/dev/blog/why-cassandra-doesnt-need-v... http://www.datastax.com/dev/blog/cql3_collections http://www.datastax.com/dev/blog/cql3_collections http://www.datastax.com/dev/blog/lightweight-transactions-in-cassandra-2-0 http://www.datastax.com/dev/blog/lightweight-transactions-in...
- sseveran 13y agoCassandra suffers from the same problem and can drop updates. The Paxos transactions were and maybe still are an absolute joke as exposed by Aphyr.
- voidmain 13y agoFoundationDB provides real ACID transactions and external consistency, and definitely does NOT rely on clock accuracy for soundness! (Google Spanner, which we are often compared to, does use a trusted clock, but Google went to extreme measures to make it accurate, including installing atomic clocks and GPS hardware.) As for how, it's a long story. At bottom we rely on Paxos for consistency across failures, but we only actually do Paxos when there are failures. (We use less costly synchronous techniques for replication in "happy times".)
- flavien_bessede 13y agoDistributed systems design aside, the core of the problem is that they relied on ntp (as they probably should), and in their case ntp was not working properly.
- devicenull 13y agoAnd this is precisely why a thing that is not monitored is not actually a thing.
- olefoo 13y ago> "A thing that is not monitored is not actually a thing." That should be on a cross-stitch sampler on the wall of every NOC.
- specialist 13y agoNice. Much stronger than the "you can only manage what you measure" adage I learned from accounting.
- globalpanic 13y agoEven if NTP had been working properly, you would not have clocks synchronised at the level of individual ticks - only to the level of time intervals. If two updates happened at roughly the same time, and fell into the same time interval, there would be no way to tell which one happened before the other. A paper by Cilia et al on timing of composite events in distributed event-based systems using NTP deals with this issue.
- Dylan16807 13y agoBut this is not a problem in many situations. Whereas successor writes failing within an entire 30 second span is a pretty big problem.
- scottdw2 13y agoThe key take away from the article SHOULD be: don't rely on ntp if you don't have to. There are people who have to. They run their own atomic clocks, and worry about things such as precision delivery of nuclear ordanance. Then there's you. You should use vector clocks, with a builtin conflict resolution mechanism based on domain knowledge. That's the point of the article.
- tych0 13y agoI can't scroll down.
- rossy 13y agoSame here. On Windows/Chrome 30 with a window size of about 900x950px, I can't scroll down. Increasing the windows's width or decreasing the height makes it work again.
- dustinupdyke 13y agoI don't get how clocks are bad this from the article. I get that syncing clocks across systems is hard and when it goes awry, unintended consequences are incurred.
- RickHull 13y agoThere is, in fact, a TL;DR at the end: > If your distributed database relies on clocks to pick a winner, you’d better have rock-solid time synchronization, and even then, it’s unlikely your business needs are served well by blindly selecting the last write that happens to arrive.
- sseveran 13y agoIn fact I would recommend GPS calibrated hardware clocks with PTP.
- donavanm 13y agoAs last summers negative leap second fiascos demonstrated even a trusted source isnt enough.
- oh_sigh 13y agoThey are when you know that the leap-(nanosecond/second/minute/day) is coming up. When you know it is coming, you can "smear" the time difference over, let's say, the entire year, so when it happens, every system behaves correctly.
- marshray 13y agoThe point is not that time synchronization is inherently bad, only that it's usually not the correct thing for a distributed database to resolve update conflicts with.
- sseveran 13y agoYes I completely agree. I fail to see how anyone would think that using a clock as a source of truth in a distributed system would be in anyway a good idea. As far as PTP it would be too expensive to deploy at large scale which was some of the motivation (i believe) behind truetime.
- mey 13y agoIs it me, or does the hand waving at the beginning of the article between "write" and "update" smell of bad spin? As a developer I consider both "creates"/"updates" as "writes". "Riak is designed to accomodate (sp) server and network failures without ever losing committed writes, so this led to a quick response from Basho’s engineers." As such losing a write to me when I read documentation is losing either a create, update or a delete. Any side affecting operation essentially. Anything that needs to write to disk to record a change...
- macintux 13y agoThanks for catching the spelling error; as much as I pride myself on my spelling, I should let 2013-era tools do their job. I was concerned that might be interpreted as spin, but I hoped the rest of the article would reinforce the point that there is no way to guarantee an update is preserved in a distributed system without an approach more sophisticated than blindly trusting clocks. Writes to a new object are inherently less problematic; while it's possible to temporarily receive a negative response about the presence of an object, the data will always be there, barring catastrophic multiple server failure. Updates can be entirely lost, and that's something that developers and operations people need to be aware of.
- mey 13y agoMy apologies for being a little gruff. I am coming at this from not being a user of Riak and currently exploring options for distributed processing of data as our companies data needs have gotten a bit big. I was just expressing my concern over the complexity of the problem and our understanding of a technical term. It makes it harder to consume documentation on systems for evaluation, to get an idea of how they fail and how to adjust to failure. It may not be rational but in my gut it causes me concern.
- macintux 13y agoI absolutely understand your concern, and I'll be more cautious in the future. I tend to write with a very casual, informal tone, and data safety is not something to be overly breezy about. More broadly, as someone who helps write our documentation, it's very difficult to figure out how to present enough detail about the proper ways to use Riak without forcing everyone to become an expert on distributed systems. Unfortunately there are incredibly subtle tradeoffs inherently involved in running a distributed database.
- losethos 13y agoFucken psychologist-nigger thinks the brain can interact with a 400Mhz clock. What a fucken nigger, huh! CWinApp theApp; using namespace std; long long i=0; void MyTask(int u) { while (true) i++; } int _tmain(int argc, TCHAR* argv[], TCHAR* envp[]) { int nRetCode = 0; CreateThread(0,65536,(LPTHREAD_START_ROUTINE)MyTask,0,0,0); while (true) { getchar(); printf("%012u",i); } return nRetCode; } ------------------- Gods says line: 9632 that offereth it. 7:10 And every meat offering, mingled with oil, and dry, shall all the sons of Aaron have, one as much as another. 7:11 And this is the law of the sacrifice of peace offerings, which he shall offer unto the LORD. 7:12 If he offer it for a thanksgiving, then he shall offer with the sacrifice of thanksgiving unleavened cakes mingled with oil, and unleavened wafers anointed with oil, and cakes mingled with oil, of fine flour, fried. 7:13 Besides the cakes, he shall offer for his offering leavened bread with the sacrifice of thanksgiving of his peace offerings. 7:14 And of it he shall offer one out of the whole oblation for an heave offering unto the LORD, and it shall be the priest's that sprinkleth the blood of the peace offerings. 7:15 And the flesh of the sacrifice of his peace offerings for thanksgiving shall be eaten the same day that it is offered; he shall not leave any of it until the morning. ----- Why do we take communion? So that Christ will dwell with-in us. It is Christ, not our brain that controls a 400Mhz clock! Nigger! Randomly open a book you've never read. Psychologists are stupid. God I hate them.
- zschoche 13y agoDecades ago albert einstein introduced the general theory of relativity, which is already telling us that timestamps are bad for synchronisation.
- joaomsa 13y agoHow would you get around the unreliability of clocks in VMs? Seems like deploying Riak in the cloud could be problematic.
- hectcastro 13y agoYou'd change the settings detailed in the "Stopping Last Write Wins" section of the original post.
- jbert 13y agoDumb question: What breaks with the following approach? 1) set last_write_wins=true (so all updates, always apply, as described in the article) 2) avoid the "partition/rejoin may cause old values to stomp on new" issue by having "rejoin detection" which refuses to rejoin if clocks are "too out of sync"
- crb002 13y agoMars design. Assume one server is on Mars, with associated time dilation on it's clocks and latency.