4 ms·
I was once (partially) responsible for the deaths of dozens of virtual machines at a distance of about three and a half years. Fun fact: none of these VMs had
by jsolson 7y ago
I was once (partially) responsible for the deaths of dozens of virtual machines at a distance of about three and a half years.
Fun fact: none of these VMs had rebooted in that time, or they wouldn't have crashed.
Anyway, back in 2014 or so I dropped a bunch of transmit packet completions. In most cases I also double completed packets which was immediately fatal. Kernels get mad about that sort of thing.
Turns out, not all of the affected VMs died. Some of them lived on with head indices forever unequal to tail indices (until they rebooted).
In 2018 a developer realized there was a potential bug in waiting for VMs entering a quiescent state -- a truly idle networking stack had retired all Tx packets that it had admitted. Having unequal indices was impossible under correct operating conditions. They fixed the glitch.
This change rolled out gradually.
Gradually, the kernel panics appeared.
The change rolled back, halting the impact, but then the analysis began. What had we broken?
Another fun fact: Linux often includes an uptime in dmesg logs.
Slowly a pattern appeared. The dmesg logs included unusually large numbers for uptimes. Plotting these, there was a clear cliff in terms of a minimum uptime. Historical deployment logs showed a noteworthy release at that date, years past. Noteworthy in that it was rolled back for my bug, years prior.
On the plus side, I realized this was almost certainly my years prior fuckup slightly sooner than anyone else, so at least I got to call myself out :)