7 ms·
Man. Time based bugs are some of the worst. I spent weeks trying to figure out what was wrong with some of my scripts running under WSL. Apparently wsl Linux ke
by saltedonion 5y ago
Man. Time based bugs are some of the worst. I spent weeks trying to figure out what was wrong with some of my scripts running under WSL. Apparently wsl Linux kernel had a bug that could cause time to drift by minutes.
- dmos62 5y agoMy most memorable bug (because of how long debugging took) was inconsistent use of local and UTC timezones. Learned to localize timedates as late as possible (or preferably never) and delocalize as early as possible.
- znpy 5y agoas if it is a rite of passage to become an adult engineer, i'm now facing timezone issues at work. no real problems yet, but the utc vs local timezone has to be handled. Could you please share what did you read on this topic? Thank you!
- dmos62 5y agoI don't think I read anything specific on the topic. I'd simply suggest only using local time in the views. Hopefully that's possible in your use case. Be defensive in your design.
- treeman79 5y agoThis is why mission critical systems like Boeing airliners require a reboot. At some point it’s just a lot more practical and safer.
- sokoloff 5y agoThat was done via Airworthiness Directive (a mandatory prod patch for aviation), not by original engineering design intent. Similar story with the Patriot missile defense battery. * - https://www.federalregister.gov/documents/2020/03/23/2020-06092/airworthiness-directives-the-boeing-company-airplanes https://www.federalregister.gov/documents/2020/03/23/2020-06... * - https://barrgroup.com/sites/default/files/case-study-patriot-missile-defects.pdf https://barrgroup.com/sites/default/files/case-study-patriot...
- donkarma 5y agoMy Windows 10 clock goes 2-5x faster when I hibernate and come back
- throwaway2048 5y agoThere is a funny bug with Windows Vista in KVM/QEMU where the clock ticks about 1000x faster than it should. You can see the hour hand moving in the clock panel, and animations play at warp speed, and media playback is extremely screwy.
- imglorp 5y agoOohh, are we doing time bugs? Sun T4 servers around early 2010s, and I forget which Solaris release this was, had a beaut of a random clock jump. Every now and then, certain apps on the a system would crash around the same time. We'd scrape through the logs and usually see our app cratered, some databases, ntpd, sshd naturally. Logging of timestamps was iffy, obviously. The ntpd was the obvious suspect because, despite keeping a nice low offset to its peers like 1ms or two, for months at a time, out of the blue it would confess something like "time offset is too large, I can't fix that so I'll exit!". After chasing Sun ntpd bug reports[1] for a while, we ruled it out when we saw a pattern in the undamaged logs that looked like 09:59:58.000 ... 09:59:59.123 ... 09:09:01.345 ... 09:09:02.123 ... Yep the system clock really had jumped back almost an hour. That explained everything about the userspace going nuts including ntpd exiting as a symptom and not a culprit. After some Sun support and some sunbugs searching [1 again] we found the T4 in that Solaris rev had a hardware RTC with separate registers for H, M, S etc and a write mutex protecting them, but no read mutex. It was possible to read the RTC while it was being updated, which happened when it was syncing the OS clock to the RTC, or something like that. Fixed in a later release. 1: RIP sunbugs database. It was such a mature relationship where Sun would let everyone see what they were working on and customers could participate or at least know about known issues. I would love to find an archive. Of course Oracle shut that off immediately so you had to open a ticket and ask.
- gregmac 5y agoMy favorite time bug: Someone on a team I was on spent probably weeks working on a way to process an inbound sensor data stream (which was generated by another bit of software also maintained by us). The data was akin to an odometer, increasing over time based on usage. The trouble was there appeared to be two parallel streams of the same name, which were always offset from each other and while the offset varied they were always roughly close. The algorithm this person came up with eventually sorted them into "thing_1" and "thing_2" streams, at which point it got turned over to me to display. It actually worked pretty well, for what it's worth. I started asking how am I supposed to show this to a user to make use of, what does this even represent?... but never got an adequate answer. So I started looking at the whole chain, and what I found was the piece of software generating the data had a small bug: it used "hh" instead of "HH" in the timestamp, but also no am/pm. The timestamps were supposed to be 24h and looked like it, but 9:02AM and 9:02PM both came in as "09:02". To confirm, we checked the database and, sure enough, every bit of data was between 1:00am and 12:59pm. In the end we fixed the timestamp bug and threw out the processing code. It's simultaneously a hilarious date bug, a face-palming colossal waste of time, and a lesson on how not to run technical teams (the manager had no technical skill, and the "team" of ~7 was very siloed and each operated more like teams of 1-2).
- throwaway2048 5y agoA very annoying DNS over HTTPS/TLS circular dependency bug manifests itself if your device doesn't have an RTC, or the battery is dead, or the clock is sufficiently skewed. Clock is fucked, so TLS certs don't verify due to validity times, so DNS is broken, so NTP cant look up domains, so the clock can't be set...
- exikyut 5y agoAt the end of the day since TLS depends on correct time to trust certificates, I guess the "everything is fine" solution is to fetch the DoH server's TLS cert, inspect the start and end dates, set the system time to the exact middle, then helicopter over NTP a bit to make sure it came up and changed the time to something hopefully correct. On the one hand there's not very much else you can do since you're pointing at bits of thin air and saying "there's the trust chain" in the first place, but on the other hand plaintext DNS is... not much better? Of course that's when the existential "why even DoH in the first place" starts (with side servings of "this feels so wrong putting it on the security report")... (...Why do I suddenly feel like disabling certificate verification is going to catch on in a big way in embedded ntpds, almost like a standard best practice... aaaaaaaaa)
- mrtesthah 5y agoThe best and simplest solution would just be to require an administrator to input the current date manually in that circumstance.