8 ms·
Making every (leap) second count with our new public NTP servers
- leephillips 10y agoA lot of interesting geophysics in the unpredictable need for leap seconds. I mention Google's "smearing" approach here: http://arstechnica.com/science/2016/04/the-leap-second-because-our-clocks-are-more-accurate-than-the-earth/ http://arstechnica.com/science/2016/04/the-leap-second-becau...
- deleted 10y ago[deleted]
- brandmeyer 10y ago> Instead of adding a single extra second to the end of the day, we'll run the clocks 0.0014% slower across the ten hours before and ten hours after the leap second, and “smear” the extra second across these twenty hours. Holy leaping second, batman! Unilaterally being off by up to a half second from the rest of the world's clocks is a pretty aggressive step. I think I would have preferred to see a resolution made by an independent body on something this drastic.
- subway 10y agoThey've been doing this since at least 2011.
- brandmeyer 10y agoThat doesn't make it any less unilateral. IMO, it makes it rather worse to have been doing this for five years. The situation would be much better if they had been working to build a broader consensus over all that time. As near as I can tell, they don't even have consensus within Linux, let alone POSIX or the ITU.
- detaro 10y agoHonest question: what's so bad about others' systems running on slightly off time? I get why people care about internal consistency, and why deviations should be quite small, but this?
- brandmeyer 10y agoSlightly contrived example: Lets say that you were running a distributed database, and you had distributed instances across different cloud providers for increased reliability. if your database relies on high-resolution timestamps for distributed conflict resolution, then you're going to have a hard time. Another example: Suppose that a portion of an industrial monitoring system processes remote sensor data in a cloud datacenter with smeared time, while the sensor nodes keep strict UTC time. Your SCADA system had better not have any hard-baked assumptions like "messages cannot come from the future", or you're going to have a hard time, too. Lets say that a company's internal NTC servers include several sources for reliability and redundancy. Much like Google DNS, perhaps one of the sources is Google NTP, while another is derived from the NTP pool. How do you expect the NTP daemon to behave in this situation? It will certainly be able to observe a 500ms difference between its source timeservers.
- detaro 10y agofirst example: ok, yes, if they offer the DB as a service, that would be bad. If you run it on a VM it's IMHO your responsibility to make sure your time sensitive database nodes have shared time. Same for the second example, interesting point for SaaS scenario, although it seems like that could break through normal deviations already. EDIT: ok, the blog post actually mentions "local clocks in sync with VM instances running on Google Compute Engine", my bad. Not sure what to think about that. In comparison, Amazon recommends running NTP on your VMs and their Linux AMIs come with pool.ntp.org configured as default. </edit> Third: It's going to figure out some solution (if Google is only one source it's probably going to drop it as faulty), but you probably should not have added a time source that's officially documented to not strictly follow standards. It's not like Google offered a NTP service for years and now suddenly switched how it works. I guess I underestimate the amount of trust people put into random time sources: practice is probably messier than theory.
- JoshTriplett 10y agoSmearing leap seconds does make sense, but it's an odd step to take unilaterally, rather than coordinating with other NTP servers and with Linux timekeeping (which currently handles leap seconds via a 61-second minute instead).
- ac29 10y ago>Linux timekeeping (which currently handles leap seconds via a 61-second minute instead) Google doesn't think so: "No commonly used operating system is able to handle a minute with 61 seconds"
- jmgao 10y agoBits of pieces of the operating systems might handle leap seconds properly, but it's doubtful that every single component that uses time does the right thing. The last two leap seconds have revealed bugs in the kernel: https://lwn.net/Articles/504744/ https://lwn.net/Articles/504744/ for the one in 2015 and https://lwn.net/Articles/648313/ https://lwn.net/Articles/648313/ for the one in 2016, and I don't think it's unlikely that the one scheduled for December will reveal another.
- JoshTriplett 10y agoSome things did go wrong on a few of the previous leap-second injections, and the Linux timekeeping maintainer had talked about changing the approach to handling them (which has already changed at least once in the past). I don't, however, think it makes sense to unilaterally change this, without (any obvious signs of) coordination with the timekeeping maintainers and the maintainers of major NTP servers.
- fanf2 10y agoThings go wrong on every leap second, sometimes catastrophically. They go wrong on non-leap-seconds because of falsely advertised leap seconds. They go wrong 4 months before a leap second because a leap indicator got set and some software had an incorrect idea of when it was due. Never mind the theory, the practice is a clusterfuck.
- KayEss 10y agoIt's far better than what POSIX clocks do. They'll just drop back a second and you'll get the same time twice.
- paulajohnson 10y agoThat is a historical artifact. The original Unix developers decided to treat time as seconds since the start of 1970, implicitly assuming that every day has 86400 seconds. Back then UTC was in its infancy, most programmers had not even heard of leap seconds, and most computer clocks were set by the sysadmin looking at his watch. If we were starting from scratch we would have a date-time type with a day number field and a seconds-since-midnight field. However that would be a breaking change for every piece of software out there, so we are stuck with a time_t that cannot handle leap seconds.
- leni536 10y agoOr we would use the NTP timestamp directly (or a lower precision version) as time_t, which (AFAIK) doesn't suffer from leap seconds. On can always convert time_t to truct tm, tm_sec is defined to be in the range of [0..60].
- leni536 10y agoLooks like I was wrong: https://www.eecis.udel.edu/~mills/leap.html https://www.eecis.udel.edu/~mills/leap.html
- magicalist 10y agoYou're going to have a bad time if you assume "the rest of the world" isn't doing their own, different adjustment https://developers.google.com/time/smear#othersmears https://developers.google.com/time/smear#othersmears > preferred to see a resolution made by an independent body Independent bodies have spent the last decade debating if leap seconds should even exist. Agreeing on how to treat them if we keep them is way down the priority list.
- orblivion 10y ago> You're going to have a bad time... Hilarious. Was this intentional?
- leni536 10y agoWhy the hell aren't time servers and clients sync to TAI instead? Dealing with leap seconds should be a client side problem.
- russdill 10y agoYup, it seems awesome, but software needs to be written to handle it properly. I think it'd work no problem with software already written against monotonic clocks, but everything else would probably need some fixing.
- paulajohnson 10y agoThe problem is that time_t (seconds since 1970) implicitly assumes 86400 seconds per day. You would have to redefine time_t and rewrite every piece of code that uses it.
- kazinator 10y agoLeap seconds is what allows that assumption to work. Leap seconds exist only in real time, not in historic recorded time. There are in fact 86400 "calendar seconds" in a day, exactly. Essentially, when a day is done, we call it 84600, even though it's actually 86400.epsilon. Only special applications need to know the exact physical number of seconds between two calendar times, rather than the calendar seconds.
- azinman2 10y agoRight... Because leap seconds are high on everyone's priority list. It's the programmers fault! Or... recognize that super rare bugs are inevitable and create a higher level way to avoid them entirely. I vote for option 2.
- kazinator 10y agoThat "new higher level way" is leap seconds. The concept of leap seconds allows us to do date and time calculations in "calendar seconds", and not care about the discrepancy between physical seconds and calendar seconds. Leap seconds basically add a corrective jump to physical time (what is measured by our super accurate clocks that use physical seconds and not calendar seconds) to match calendar time. Leap seconds matter if you're doing some scientific or engineering calculation (astronomy, aerospace or whatever) and you need an exact physical time down the fraction of a second between two events that are far apart in the calendar. They do not enter into everyday calculations, like using time_t seconds to calculate the number of days between two dates.
- newman314 10y agoDoes anyone know if Google has open sourced the time smearing algorithm?
- deleted 10y ago[deleted]
- detaro 10y agoThey discuss various ways of smearing here, but I haven't seen their actual implementation code: https://developers.google.com/time/smear https://developers.google.com/time/smear
- liotier 10y ago“Leap Smearing must not be used for public-facing NTP servers” - https://tools.ietf.org/html/draft-ietf-ntp-bcp-02 https://tools.ietf.org/html/draft-ietf-ntp-bcp-02
- klodolph 10y agoWow, that's a really boneheaded thing to put in a standard. I think we can all agree that it's important to make leap smearing available for those who want to use it, especially considering the bugs in leap second handling for common NTP clients.
- paulajohnson 10y agoI disagree. The point of NTP, and of time services in general, is that everyone agrees about the time. If an organisation wants to use non-standard time it can, but public-facing NTP servers should all agree and all provide the standard time. Google, for whatever reasons, is making its NTP servers deliberately wrong, and there is no mechanism in NTP for a server to say "I'm using time-smearing". So they shouldn't be doing this on public-facing NTP.
- toomuchtodo 10y ago> and there is no mechanism in NTP for a server to say "I'm using time-smearing". That should most definitely be in the standard, along with communicating to the client full details about how smearing is configured.
- creshal 10y agoSure, but until then, public-facing NTP servers should stick to the current standard.
- klodolph 10y agoThen NTP has already failed. Most systems are already incapable of agreeing on whether it is 23:59:59 or 23:59:60 on days with leap seconds. There is simply not an API that will let you distinguish the two. It is better to be deliberately wrong in a controlled fashion than to be accidentally wrong because you never expected your clock to be non-monotonic. You seem to be arguing for the status quo, are you aware of just how deeply broken the status quo is?
- deleted 10y ago[deleted]
- nullc 10y agoI predicted this for leap smear a while back-- we have time sync because having systems with different times is a source of problems... logical fix: get them onto the same time. Smear is a workaround for those who care about phase alignment but don't care about frequency error. ... and who don't need to exchange times with anyone else. This last point reduces the set to no one, since it can't extend to everyone (some parties care a lot more about frequency error than phase error!). This circus is enhanced by NTP's inability to tell you what timebase it's using (or, god forbid, offsets between what its giving you and other timebases...) It's going be especially awesome when NTP daemons with both smear and non-smear peers get both the smear frequency error AND get a leap second. I for one welcome this great opportunity for an enhanced trash fire to help convince the world that we need to stop issuing leap seconds. (It's absurd-- causes tens of millions in disruption easily, -- and it would take 4000 years to even drift an hour off solar time, at which point timezones could be rotated if anyone really cared).
- detaro 10y ago> Smear is a workaround for those who care about phase alignment but don't care about frequency error. ... and who don't need to exchange times with anyone else. This last point reduces the set to no one, since it can't extend to everyone (some parties care a lot more about frequency error than phase error!). I don't quite understand that point. E.g. the typical web server doesn't have much of a need to exchange precise time with others. HTTP, TLS, ... require timestamps, timestamps are shown to users occasionally, but as long as they are roughly right that is enough. As long as all internal systems work off the same standard it is fine. Which seems to be the reasoning under which Google choose to use it, even though one might argue that with their cloud offerings they are not as insular.
- antoncohen 10y agoFor people talking about Google unilaterally doing this, it has been common to smear the leap second for the last couple years. Usually companies do it internally by having their NTP servers skew time, either with Chrony or `ntpd -x`. Standards bodies have not been able to react quickly enough to the need to smear the leap second in a consistent way. I'm thankful that Google has decided to run public NTP servers with consistently smeared leap seconds. Here are two Red Hat articles on how to deal with the leap second, from 2016 and 2015: https://access.redhat.com/articles/15145 https://access.redhat.com/articles/15145 http://developers.redhat.com/blog/2015/06/01/five-different-ways-handle-leap-seconds-ntp/ http://developers.redhat.com/blog/2015/06/01/five-different-...
- JdeBP 10y agoI hope that there are people on standards bodies who remember or learned what it was like before UTC when civil time seconds were not one SI second long, and in effect "smearing" happened all the time.