3 ms·
Thanks so much for this answer, I just learned a lot. > There is no guarantee that the attempt to measure or detect faults won't hide them... I got completely
by i336_ 9y ago
Thanks so much for this answer, I just learned a lot.
> There is no guarantee that the attempt to measure or detect faults won't hide them...
I got completely stuck on this in my original ponderings. I totally didn't think of sprinkling instrumentation instructions into the code and seeing if the bug still fired. If it did, the approach you described would certainly work very well (and, indeed, you describe it being widely used).
Major TIL with the address-based tracing hardware idea. That's an awesome approach, to do it that way... wow. :)
Building something like this would actually be a really cool challenge in designing a really fast piece of hardware. Considering the kind of access speed needed, though (particularly with the memory bus approach)... would a custom ASIC be required? :/ Or could I get away with using a (perhaps decent/pricey) FPGA?
I say this because it would be awesome to make something like this inexpensively available for people to put in their workstations. I can totally see a device like this also having some fast, nonvolatile* memory-mapped storage for things like infinite logging, as well. For example, the way Linux handles crashes is to kexec into a new kernel that hopefully fishes the log out of RAM and saves it. Very clunky. This approach also does not handle early kernel bringup - or even BIOS/EFI bringup, the libreboot folks would probably love something like this.
(*By "fast, nonvolatile" I mean something that writes straight to a large DDR-backed buffer and is then quickly yanked onto something like an NVMe disk.)