4 ms·
Your post made me read this: https://www.cockroachlabs.com/blog/living-without-atomic-clocks/ https://www.cockroachlabs.com/blog/living-without-atomic-clo... on
by ketzu 4y ago
Your post made me read this: https://www.cockroachlabs.com/blog/living-without-atomic-clocks/ https://www.cockroachlabs.com/blog/living-without-atomic-clo... on how spanner and cockroadDB do this. And the answes are surprisingly simple, but with consequences attached.
* Spanner creates very tight bounds on the clock synchronization (through atomic clocks/GPS) and ... just waits out the length of the bound.
* CockroachDB seems to do lamport-clocks-with-real-timestamps for linearizability (track the highest seen timestamp for causality chains). For preventing consistency violations with reads, they also track the bounds on the clock and potentially attempt to read again.
So they approve of the overall message (of not just comparing conventional timestamps) but work around those using the uncertainty of the clocks they have available.
My professor used to introduce logical clocks with "as we usually can't use atomic clocks on all nodes" because of that.
- hansvm 4y agoThe "as we usually can't use atomic clocks on all nodes" constraint is more common than you might think. Even code as simple as ts = get_atomic_time_with_no_delay_because_this_node_has_an_atomic_clock(); do_stuff_with_time(ts); is really broken if you have to wait for garbage collection, wait for your process/thread/coroutine/... to be scheduled, wait for a nearly invisible pause as a cache line is refreshed, .... The time you're using the atomic timestamp is at some future moment, and if that delay matters for your algorithm you need to track your uncertainty just like Spanner/Cockroach do (you just get significant benefits because the uncertainty is low).
- kurthr 4y agoIf your get_atomic() isn't also capturing or resetting an internal hardware timer atomically it's kinda broken code anyway. It's important to remember that you're getting data from the firmware of an atomic clock (or GPS receiver) and that they will go to great lengths to compensate for latency variations so not doing it in the sampling code would be nuts (or just not using isochronous coms). Then the local clock/timer can be calibrated from the that and you can have near clock-cycle correct timing. Typically they only send out 1 pulse per second so anything finer than that is locally interpolated. We used 1ms internal intervals with an accuracy of ~1ppm. Last time I worked on something like this was 20 years ago, but it was a Pentium box. The clock firmware to converted a 5MHz sinewave output to 1Pps, but could also output a synthetic UTC through LAN in cases where there was no other connectivity. I guess I'm saying this is really specialized code for whatever hardware/OS you're using. (edit) Here's some documentation for a very similar box: https://www.manualslib.com/manual/1281621/Symmetricom-5071a.html?page=27#manual https://www.manualslib.com/manual/1281621/Symmetricom-5071a....
- jrockway 4y agoThe GPS clock is really used to provide long-term stability for the computer's internal oscillator, which provides fine short-term stability. Your calls to gettimeofday() are handled by the internal oscillator, not by asking the GPS unit directly. Additionally, NTP daemons are impressive in their ability to get the "right answer" on current internal oscillator frequency as affected by the entire system. With 86400 very accurate data samples per day, you can do a lot! In a distributed system, the absolute value of the time is important, since that's what your transaction timestamps are based on. There is a lot to do to ensure that your time offset from UTC ends up being correct; antenna cable length compensation, oscillator quantization compensation, etc. These are not strictly necessary for Spanner but can make microsecond-level differences which are quite noticeable. "The API directly exposes clock uncertainty, and the guarantees on Spanner’s timestamps depend on the bounds that the implementation provides. If the uncertainty is large, Spanner slows down to wait out that uncertainty. Google’s cluster-management software provides an implementation of the TrueTime API. This implementation keeps uncertainty small (generally less than 10ms) by using multiple modern clock references (GPS and atomic clocks)." Basically, there is some tradeoff between clock synchronization perfection and transaction processing speed that can be made. You can build dedicated hardware that's clocked by a GPSDO and provide an API to get hardware timestamps, and maybe process transactions a little faster. Your good old C++ program running on Linux with an internal oscillator adjusted by NTP with a GPS PPS input still beats communicating between datacenters to agree on event ordering, though. Light is slow!
- eternalban 4y ago> In a distributed system, the absolute value of the time is important Only for a co-variant data cohort is total temporal order necessary for correctness. It doesn't need to be system wide. Distinct (data independent) process groups can be partially ordered. There really is no such thing as "now" outside of a shared frame (or intersecting frames) of observation.
- hansvm 4y agoAll I was pointing out is that worst-case latency on a throughput-optimized OS can be unbounded, and it's not totally unexpected to see 1ms+ delays between those two lines of code, even if the clock implementation is flawless. Even in a RTOS you have jitter from cache misses and whatnot, just to a lesser degree. Code sensitive to clocks not being correct is often sensitive no matter how small the deviance is, and smaller deviations simply make bugs less frequent. Treating that point-in-time estimate as anything other than an estimate from the recent past can lead to code that looks flawless but occasionally breaks.
- photochemsyn 4y agoThat's an interesting link, it leads to a this 1991 Liskov paper, which might seem out-of-date, but apparently was very foundational to the whole concept. A little searching turned up this modern discussion of it: https://muratbuffalo.blogspot.com/2022/11/practical-uses-of-synchronized-clocks.html https://muratbuffalo.blogspot.com/2022/11/practical-uses-of-... Seems the bottom line is: "Since clock synchronization can fail occasionally, it is most desirable for algorithms to depend on synchronization for performance but not for correctness."
- jasonwatkinspdx 4y agoFYI the general technique Cockroach is using is commonly called Hybrid Logical Clocks, though Cloudera call their very similar idea Virtual Time as I recall. Anyhow, here's a nice summary: http://muratbuffalo.blogspot.com/2014/07/hybrid-logical-clocks.html http://muratbuffalo.blogspot.com/2014/07/hybrid-logical-cloc...
- preseinger 4y agohttps://www.cockroachlabs.com/blog/living-without-atomic-clocks/ https://www.cockroachlabs.com/blog/living-without-atomic-clo... > perfectly synchronized clocks are a holy grail of sorts for distributed systems research. They provide, in essence, a means to absolutely order events, regardless of which node an event originated at two events that occur on two physically discrete nodes at exactly the same time have no well-defined order, even if their clocks are synchronized > before a node is allowed to report that a transaction has committed, it must wait 7ms. Because all clocks in the system are within 7ms of each other, waiting 7ms means that no subsequent transaction may commit at an earlier timestamp, even if the earlier transaction was committed on a node with a clock which was fast by the maximum 7ms. Pretty clever. so sending information between two antipodes on earth takes 66ms in one direction, 132ms round-trip, minimum. the speed of light dictates this lower limit if two transactions are made on those two nodes at exactly the same time, there is no objective order between the two. you can choose an order, but only with knowledge of both transactions, and that information physically cannot traverse space-time in 7ms so it's really not clear to me how this works. as an optimization, sure -- maybe there is some optimistic concurrency control that will fail transactions that conflict in the way i've described?
- layer8 4y ago“Exactly the same time” isn’t even physically well-defined, due to relativity.
- thfuran 4y agoWell, we can make some pretty accurate guesses about the frame of reference if we know the nodes' geographic positions and ban airplanes.
- preseinger 4y agoyes but those guesses have "smudge windows" which are functions of the physical distance between relevant nodes if you want to make assertions about order between events from a set of nodes, and you want to use per-node physical clocks to determine that order, then even if those clocks are perfectly synchronized, order can only be decided when the light-cones of all nodes intersect, which is the maximum distance between any two nodes times the speed of light 7ms works for up to 2098km, that's in the best case
- layer8 4y agoStrictly speaking, atomic clocks by themselves wouldn’t fully solve the problem either, because they tick at slightly different rates in different locations, due to variations in the gravitational field.