21 ms·
Leap second causing Linux server crashes?
- richurd 14y agoReally? They were running busybox? If not, it's not Linux. It's GNU/Linux.
- raverbashing 14y agoOuch! My Debian GNU/Linux 6.0 is still standing Oh well, reading the issue, the machine date is Sat Jun 30 16:11:31 EDT 2012 Stopped ntpd just in case
- deleted 14y ago[deleted]
- rbanffy 14y agoSame here. Set ntp to restart in 12 hours.
- raverbashing 14y agoWith ntp stopped, no problem whatsoever
- drivebyacct2 14y agoMy Ubuntu servers seem unaffected thus far.
- brongondwana 14y agoIt seems to depend on high load, so you could be lucky!
- rarrrrrr 14y agoUnfortunately I can confirm that Ubuntu 10.04 is vulnerable. We're proceeding with the fixtime.pl workaround.
- boyter 14y agoI can confirm it too, but didn't catch it in time. A reboot however and everything is back to normal.
- henrikschroder 14y agoWe didn't catch it in time either. It was oh so much fun to wake up to our service not working at all, all java and mysqld processes spinning like crazy, and having to reboot all servers. :-/
- alexkus 14y agoI found mysqld (5.5.24) running at 159% CPU on a Ubuntu 11.04 (64-bit) box this morning. ntpd had drifted in the leap second between 1am and 2am (GMT) this morning (NTP drift info is one thing I graph with MRTG). [EDIT] Ah, covered elsewhere. Fixed by manually setting the date on the box; stopping/restarting mysqld or ntpd doesn't make any difference.
- bifrost 14y agoNo burps from my BSD boxes either, although they're all in UTC so the leap second hasn't happened for them yet.
- MrUnderhill 14y agoThe leap second is added at the same point in time regardless the timezone your server is configured to use. So if you're GMT+3, the leap second will be inserted at 03:00 local time. From the answer: "The reason this is occurring before the leap second is actually scheduled to occur is that ntpd lets the kernel handle the leap second at midnight, but needs to alert the kernel to insert the leap second before midnight. ntpd therefore calls adjtimex sometime during the day of the leap second, at which point this bug is triggered."
- sohn5 14y agoThat wouldn't happen if servers were Macs
- bifrost 14y agoThats relevant how?
- deleted 14y ago[deleted]
- TazeTSchnitzel 14y agoMac OS X is a BSD variant, so there's every chance.
- daeken 14y agoHow does that follow? OS X runs an odd hybrid kernel (XNU) which is Mach and parts of BSD, but... this is a Linux kernel bug. There's an effectively zero chance of this impacting anything but Linux.
- TazeTSchnitzel 14y agoThe kernel is not the only OS component relying on time that may have not considered this.
- chc 14y agoThis is evidently a kernel bug. The fact that both operating systems rely on time isn't particularly relevant. Could there be time bugs in OS X? Certainly. But it wouldn't be this one. Windows relies on time too, so I don't see why you bring up the fact that OS X is a BSD variant.
- tedunangst 14y agoBSD doesn't have an adjtimex syscall, so it's very unlikely for there to be a spinlock bug in the adjtimex syscall that doesn't exist.
- duiker101 14y ago2012. and we still have problems keeping track of time. This is both fascinating and scary. P.S. for people wanting to know more this video is simple to understand but really amazing http://www.youtube.com/watch?v=xX96xng7sAE http://www.youtube.com/watch?v=xX96xng7sAE
- crazygringo 14y agoMaybe we always will have problems? It seems to be the unique class of bug that not only is it easy to forget to test, and won't ever show up until a particular date... but then affects everyone! I can't think of any other kind of bug that never shows up ever, but then affects everyone. Rare bugs tend to stay rare, common bugs tend to get caught before they affect everyone... this is the exception.
- mbq 14y agoIts... worse. We can track time so easily and so well that we decided to screw it up.
- thaumasiotes 14y agoFrom discussion of this same issue in prior threads, my takeaway was (a) it's really not at all difficult to handle leap seconds, but (b) the POSIX standard specifically disallows them, by specifying that a day must contain exactly 86400 seconds. (Analogously, imagine if leap days occurred as normal, but a "year" by definition contained exactly 365 days.) The existence of leap seconds means that it's not possible to simultaneously have (1) system time representing the number of seconds since the epoch, and (2) system time equal to (86400 * number_of_days_since_epoch) + seconds_elapsed_today, and all the proposed methods of dealing with the problem involve preserving (2), which seems worthless to me, and throwing away (1), which I would have thought was a better model. edit: actual system times may be in units other than seconds, but the point remains
- timr 14y agoIt's harder than leap days, because leap seconds aren't inserted on a regular schedule. Leap days follow a predictable pattern of insertion. Leap seconds are inserted whenever the IERS decides to insert them. The problem of leap seconds is therefore closer to that of time zone definitions -- which are a total mess, because they depend on keeping rapidly changing system tables up to date. I can see why people don't relish the idea of requiring similar tables just to keep system time accurate.
- aidanbrandt 14y agoRead that as "high rates of cash."
- deleted 14y ago[deleted]
- mkr-hn 14y agoIs this implementation-specific, or could the Windows equivalent to ntp cause the same problem?
- mjschultz 14y agoImplementation specific. It looks like it is a bug in the Linux kernel with how it adjusts the time. It is possible that Windows, OS X, and other BSDs will be affected by a similar bug, but that would be coincidental as the bug is not due to ntpd but rather how the kernel handles a request that ntpd generates. More specifically, there is a condition in which the kernel tries to insert a leap second and, in doing so, attempts to acquire the same lock twice causing the spinlock lockup and (effectively) halting the kernel.
- deleted 14y ago[deleted]
- MrUnderhill 14y agoNovell kb: http://www.novell.com/support/kb/doc.php?id=7001865 http://www.novell.com/support/kb/doc.php?id=7001865 SLE9 (kernel 2.6.5-7.325): NOT AFFECTED SLE10-SP1 (kernel 2.6.16.54-0.2.12): NOT AFFECTED SLE10-SP2 (kernel 2.6.16.60-0.42.54.1): NOT AFFECTED SLE10-SP3 (kernel 2.6.16.60-0.83.2): NOT AFFECTED SLE10-SP4 (kernel 2.6.16.60-0.97.1): NOT AFFECTED SLE11-GA (kernel 2.6.27.54-0.2.1): VERY UNLIKELY SLE11-SP1 (kernel 2.6.32.59-0.3.1): VERY UNLIKELY SLE11-SP2 (kernel 3.0.31-0.9.1): VERY UNLIKELY Update (06/26/2012): after thorough code review -> SLE9 and SLE10 not affected at all.
- ChuckMcM 14y agoNot surprising. In spite of all press that Y2K was just a silly waste of money, its events like these that makes me suspect it would have been a much bigger deal if everyone had ignored it and fixed it after things where shown to break.
- pud 14y agoA lot of engineers[1] spent a lot of time successfully fixing Y2K bugs. Because nothing well known blew up, many people wrongly assumed that Y2K was never a real problem to begin with. [1] I moved a Fortune 100 manufacturing company's database off an ancient mainframe that would've been disastrous come Y2K. It went smoothly and was thus a thankless job. They paid well though (mid six figures - those were the days).
- infinite8s 14y agoWhen people say mid-six figures, do they mean 500k? Or 150k?
- reitzensteinm 14y agoI'd usually say just Google it, but coincidentally enough one of the front page results is someone asking the same question on this very site: http://news.ycombinator.com/item?id=2261637 http://news.ycombinator.com/item?id=2261637 It means 500k, or there abouts; 400-600k would probably qualify, with anything higher or lower being mid to high or low to mid six figures, respectively.
- joelhooks 14y ago$100/h approaches $200k, and consultants can definitely make $250/h+.
- infinite8s 14y agoYeah, but that rate usually assumes they aren't billing 40hr/weeks for 50 weeks a year.
- icefox 14y agoOddly netflix went down for me at 12:01 last night... I assumed some cronjob or something similar was to blame.
- brongondwana 14y agoFYI: I've updated the post with details of the workaround as implemented on our servers.
- dfc 14y agoGoogle uses a "leap smear" and slowly accounts for the leap second before it happens.[1] As long as you are not doing any astronomical calculations or constrained by regulatory requirements I think google has the right idea. [1] http://googleblog.blogspot.com/2011/09/time-technology-and-leaping-seconds.html http://googleblog.blogspot.com/2011/09/time-technology-and-l...
- jbeda 14y agoAs part of Google Compute Engine we provide an NTP server to the guest which is based on Google Production time. As such our VMs get to take advantage of this leap second smearing implementation. I was going to mention this at my talk at IO but forgot.
- brongondwana 14y agoMarco's blog post (linked) had a similar idea - running ntp with -x for a day so it smears time.
- dfc 14y agoIn case anyone is looking for the actual link to marco's posts on ntp: http://my.opera.com/marcomarongiu/blog/index.dml/tag/ntp http://my.opera.com/marcomarongiu/blog/index.dml/tag/ntp
- ralph 14y agoSo a VM on G's Compute Engine could in turn run an NTP server that exported G's Production Time? Do I also see GPT on App Engine? Any chance Google could just make a GPT NTP server available as a public service anyway, just as 8.8.8.8 is their public ping responder. ;-)
- enneff 14y ago8.8.8.8 and 8.8.4.4 are Google Public DNS, not a "ping responder." https://developers.google.com/speed/public-dns/ https://developers.google.com/speed/public-dns/ Google does provide time servers, although I'm not sure whether they are officially supported. The addresses are: time1.google.com time2.google.com time3.google.com time4.google.com
- kzk_mover 14y agoNow facing this issue... By using 'adjtimex' command, you can clear the problematic INS bit. At first, you can confirm the status flag like this. $ ./adjtimex --print | grep status status: 8209 8209's binary representation is like this. This surely have INS bit "100000000[1]0001" (5th LSB). $ ruby -e 'p 8209.to_s(2)' "10000000010001" 8193 is the value after the clearance of the INS big. $ ruby -e 'p 8193.to_s(2)' "10000000000001" Then, let's set it as a current value. Please ensure your ntpd is not running. $ adjtimex --status 8193
- x3c 14y agoHey, I'm running Ubuntu 12.04 . Could someone guide me through what I can do to detect/prevent this from crippling my server? Thanks.
- klodolph 14y agoRead the linked article. > The work-around is to just turn off ntpd. If ntpd already issued the adjtimex(2) call, you may need to disable ntpd and reboot to be 100% safe.
- __david__ 14y agoIt appears to be fixed in Linux 3.4 [1]. According to the original commit [2] it's been broken since 7dffa3c673fbcf835cd7be80bb4aec8ad3f51168 [3], which appeared in 2.6.26. So, kernels between 2.6.26 and 3.3 (inclusive) are vulnerable. [1] https://git.kernel.org/?p=linux/kernel/git/torvalds/linux-2.6.git;a=commit;h=bcd550745fc54f789c14e7526e0633222c505faa https://git.kernel.org/?p=linux/kernel/git/torvalds/linux-2.... [2] https://git.kernel.org/?p=linux/kernel/git/torvalds/linux-2.6.git;a=commit;h=6b43ae8a619d17c4935c3320d2ef9e92bdeed05d https://git.kernel.org/?p=linux/kernel/git/torvalds/linux-2.... [3] https://git.kernel.org/?p=linux/kernel/git/torvalds/linux-2.6.git;a=commit;h=7dffa3c673fbcf835cd7be80bb4aec8ad3f51168 https://git.kernel.org/?p=linux/kernel/git/torvalds/linux-2....
- moe 14y agoWhich, in summary, is pretty much every production kernel out there. Spent the last two hours recovering servers, tomorrow will be another interesting day. Whoever figured it'd be a good idea to INSERT[1] the leap-second instead of just slowing/accelerating time... <censored> [1] Clock: inserting leap second 23:59:60 UTC
- sharth 14y agoWell, except for RHEL 5. That runs 2.6.18.
- Ecio78 14y agoI'm still tryin to understand why all my servers seem to be ok even if they have kernel that should be affected and some of them are running mysql... For example one of them is a debian kernel 2.6.32 running mysql and ntpd, and i see in dmesg Clock: inserting leap second 23:59:60 UTC but the cpu load is ok...
- __alexs 14y ago> Whoever figured it'd be a good idea to INSERT[1] the leap-second instead of just slowing/accelerating time... <censored> That would be the IERS organisation. There's going to be a vote in 2015 to abolish them entirely.
- 14y ago
- Monotoko 14y agoPirate Bay has also been crashed by this: "TPB crashed just after midnight June 30th GMT (5.5 hrs ago) The crash appears to have been caused by the leap second that was issued at midnight." https://forum.suprbay.org/showthread.php?tid=125071 https://forum.suprbay.org/showthread.php?tid=125071
- piggity 14y agoWe just had 100s of EC2 instances generate high (alleged) load. Instances had load averages of 90+ but were responsive. Running on a 3.2 kernel Rebooted them all and they're fine.
- sehugg 14y agoWhat he said.
- mootothemax 14y agoI was logged on to a couple of CentOS 6 servers when I saw this happen, and on each one the Java processes went absolutely haywire. Everything else seemed to work fine. I attempted to fix with adjtimex and the script in the linked question, but to no avail, in the end having to restart them all instead. After that, all was good again.
- cagenut 14y agoI just had the exact same experience.
- scottbruin 14y agoHad the same issue across all our VMs running Java/Tomcat applications.
- shaggy 14y agoPardon the ignorance if this is a stupid question. I've been looking at some of my hosts and have noticed a message "Clock: inserting leap second 23:59:60 UTC" in dmesg output but each of the hosts is in the EDT timezone so the I was under the impression that the leap second hadn't been applied yet. So what does that mean? That the systems have applied the leap second successfully or have only received it from their NTP servers?
- deleted 14y ago[deleted]
- DEinspanjer 14y agoThe leap second is applied at midnight UTC time, regardless of what timezone the server is in.
- shaggy 14y agoOkay, so does that mean that the various bugs that have been circulating can still hit as it hasn't hit midnight in EDT yet or can I exhale?
- politician 14y agoAfter reading these tales of woe, all I can say is that I hope the criminal element doesn't start assaulting NTP servers.
- sayeed 14y agoOur Linux instances running on Amazon EC2 had no issues since we are not running ntpd on these servers and adjtimex returns status as 64 (clock unsynchronized). I think the Xen host takes care of the synchronization and we need not do it in the guest OS. (see http://serverfault.com/questions/100978/do-i-need-to-run-ntpd-in-my-ec2-instance http://serverfault.com/questions/100978/do-i-need-to-run-ntp...). Is this fine or should we run ntpd for better accuracy?
- csarva 14y agoYes. This issue notwithstanding, you should be running ntpd.
- arohner 14y agoStupid question: Why was this not caught? Seems pretty easy to test. Just set the clock to today (or any day with a leap second), and watch what happens.
- duskwuff 14y ago> Just set the clock to today (or any day with a leap second), and watch what happens. That won't work. The bug is only triggered when an upstream NTP server reports that a leap second was scheduled. Since leap seconds aren't predictable (and aren't even scheduled very far in advance), just setting the time back to the date of a previous leap second won't do anything.
- eadvgf 14y agoTrue, but the question still stands, since you can still test it by just telling the kernel to insert a (fake) leap second.
- Someone 14y agoIt also should not be that hard to provide your own upstream ntp server, and have that generate leap seconds at will. Both machines could be VMs, too.
- glawatscheck 14y agoPOSTMORTEM fix for CPU eating softirqd threads without rebooting: stop ntpd, run ntpdate or sntp, start ntpd /etc/init.d/ntp stop; sntp -s <ntpserver>; /etc/init.d/ntp start Unfortunately sntp / ntpdate wrapper is not shipped with squeeze for example. I've used the binary from SuSE 11.4 just fine on squeeze.
- glawatscheck 14y agoOK this is how it works on squeeze etc.: apt-get install ntpdate; /etc/init.d/ntp stop; ntpdate pool.ntp.org; /etc/init.d/ntp start
- glawatscheck 14y agoor easier still just date -s "`date`" without ntpd restart
- kabdib 14y agoFear the Unix 32-bit time-becomes-negative bugs, in 2037. We have 25 years to get ready. I still think we'll be patching at the last minute. (Yeah, lots of systems will be 64-bit by then, but there will still be a lot of embedded crackerbox systems running 32-bit timestamps. It's all the embedded stuff I'm worried about).
- bcantrill 14y agoIt's 2038, not 2037.[1] (Specifically, January 19th, 2038 at 3:14:08am.) And while lots of systems will be 64-bit, many programs still won't be -- and it seems highly likely that this will be a significantly more serious and widespread problem than, say, Y2K or DST. (And certainly more serious than leap seconds, which happen relatively frequently.) Then again, I might be biased: perhaps I'm secretly hoping to spend the years leading up to 2038 paying for my retirement with high-priced consulting gigs to fix it... [1] http://en.wikipedia.org/wiki/Year_2038_problem http://en.wikipedia.org/wiki/Year_2038_problem
- el_presidente 14y agohttp://article.gmane.org/gmane.linux.kernel/1184914 http://article.gmane.org/gmane.linux.kernel/1184914 Less than a year ago there were already people thinking about your job security. (It's a better explanation than "the glibc maintainers are insane".)
- kabdib 14y agoBut MUCH less than a year ago, many more people were still writing 32-bit-dirty time_t based code. It's gonna be a fun one.
- btilly 14y agoIf you think that being 64-bit protects you, then you do not understand the problem. The problem is that 32-bit time is embedded in filesystem representations and related protocols. (eg the POSIX specification for file times in tar.) Therefore even if your machine is 64-bit, it still needs to use 32-bit time for many purposes. To name a random example, the POSIX specification for times in the tar format is 32-bit. GNU tar has a non-standard extension that already takes care of it. But will everything else that expects to read/write tar files that a GNU tar program implement the same non-standard extension to the format in the same non-standard way? Almost certainly not. And there will be no sign of disaster until the second that we need to start relying on that more precise representation.
- yaix 14y agoTwo days ago while booting, the BIOS time on my eeepc was suddenly reset, with an error message on boot to adjust the time manually. Was just thinking that it may be related?
- kristopher 14y agoFYI: Our Debian servers did not kernel panic but system CPU load went through the roof; A quick restart brought levels back to normal.
- wiredfool 14y agoMy Ubuntu 10.04 desktop went to 100% proc and load avg of 20, none of my 10.04 servers or Debian stable servers were affected. This fixed it: date; sudo date `date +"%m%d%H%M%C%y.%S"`; date;
- ajays 14y agoYou are a lifesaver. All morning my desktop's load has been pegged at 20. I upgraded FF, Chrome, etc. and no impact. I was dreading a full re-start, as I have lots of windows, tabs, etc. open. The above command knocked the load down to almost nothing in seconds.
- chmod775 14y agoIf really all of the Linux where affected more than half of the Internet would be still down by now. Could be only a specific combination of kernel/userspace bugs that only exists in some systems. What a bit sucks is that my VPN was affected to (openvpn) causing my computer to do a poweroff. I replaced the poweroff with ip route add to 192.168.1.0/24 dev lo hope that saves me when the next leap second occurs.
- cullenking 14y agoOn debian, I was able to fix the issue (fix the load issue specifically) with this command /etc/init.d/ntp stop; date; date `date +"%m%d%H%M%C%y.%S"`; date;
- ernestipark 14y agoMy AWS EC2 instances got spun up to 100% cpu and have been like that for a day. Basically saw a step function from 0 to 100 in the CPU graph. Just had to reboot them.