3 ms·
This is why I love bugs, especially the sporadic ones--often you grow in understanding of the underlying system or code as you solve them.
by hawktheslayer 6y ago
This is why I love bugs, especially the sporadic ones--often you grow in understanding of the underlying system or code as you solve them.
- segfaultbuserr 6y ago> especially the sporadic ones Yes, but there's only a fine line between "sporadic but solvable by analysis" and "too sporadic to be reproduced or solved..." Device driver problems are often from the second category... Nevertheless, it can get interesting if the problem is reproducible like the "500 miles email" or this printer.
- marcan_42 6y agoNothing is too sporadic to be reproduced, you just haven't found how to reproduce it yet :) I once solved a once-in-a-petabyte bug at Google, when I was still working there. One of our systems ran in two data centers on the same data, and the cross-check was detecting a record offset mismatch but not a data mismatch (?!); retrying the process almost always fixed it. Long story short, there was a buffer overread in Google's bespoke zlib compressor that caused it to emit slightly less efficient compressed data (one or two bytes longer) nondeterministically, even though the data was perfectly well-formed. We were detecting compressed records being at different offsets, even though the uncompressed data matched. Another Google one involved doing a postmortem on a single instance of bad data ending up on a system. I traced that back purely from logs to having been caused by a kernel panic two systems upstream, that caused disk data to not be flushed while metadata was, so on reboot the system picked up garbage from the disk sectors in the file, which turned out to be well formed data from a different file and it trickled down the layers. Also, this one has made the rounds on HN a couple times (second time I fix a golang runtime heisenbug too; first one was also fun): https://marcan.st/2017/12/debugging-an-evil-go-runtime-bug/ https://marcan.st/2017/12/debugging-an-evil-go-runtime-bug/ And I recently fixed an OBS bug that has been randomly crashing audio for streamers for at least 3 years. I found a reliable repro that involved scripting monitoring mode toggles dozens of times per second; along the way I found and fixed 3 other bugs. That one was exacerbated by some of the OBS developers having built a rhetoric that those failures were caused by user configuration mistakes, because they often went away after changes - when what was actually going on was that users were blindly trying things to solve the problem, and sometimes stumbling on a somewhat stable workaround. (That one hasn't been reviewed/merged, still in the pipes). Then there was that bug in The Homebrew Channel that I introduced after adding freetype/TTF support... That was a heap corruption that was crashing things way later (embedded baremetal code, so no memory protection or debugging tools beyond remote gdb). Finding that one ended up involving dumping memory around the corrupted heap and following pointers to what eventually turned out to be graphics queue entries, then following texture pointers and teaching myself how to read swizzled character textures from hexdumps... I eventually found I was not allocating entries for space characters in strings, but was incrementing the index on spaces anyway. Lots of fun bug hunting stories...