24 ms·
Part 1 was interesting; it isn't clear why he split that into a Part 2 since it adds little to the story and is a paragraph long.
by Someone1234 3mo ago
Part 1 was interesting; it isn't clear why he split that into a Part 2 since it adds little to the story and is a paragraph long.
- taneq 3mo agoMight have been an “I need to look into this” segueing into “ never mind”?
- londons_explore 3mo agoI assume the fact it is a third party application means debugging gets harder, and the business case for doing so is weaker/none. But I would hope that some kind of reverse debugger triggered on one of these crashes would make it pretty simple to say "who wrote this 01".
- microgpt 3mo agoYou could also look at modules loaded into all of those processes that crashed this way.
- garaetjjte 3mo agoYou usually hope that TTD points to the culprit in such situations. But once I encountered single-byte corruption that didn't make any sense in TTD trace, there was good value at write and next read was garbage. I never discovered whether that was CPU bug, corruption by GPU shaders, stray kernel writes, or whatever.(I think it's unlikely that CPU bug would manifest with both native and TTD-instrumented runs. Corrupted byte was inside heap allocated memory so it shouldn't be in GPU pagetables at all. Kernel writes wouldn't appear in TTD trace, so really I think that was most likely issue, but how to debug that...)
- nianderwallace 3mo agoFor specific cases, you'd convert your memory allocator - hopefully you could reduce the need to just certain mem allocs - and write-protect (aka read-only memory) those mem allocs except for the situations where your code is purposely writing to those areas of memory. Yes, it'll be slow and use p lots of memory pages, unless you can reduce the mem allocs to a certain small set of allocs. And you'll have to have code to write-enable/write-disable those mem allocs. But if it catches the culprit writing bytes where they shouldn't, it'll be worth it. The one time I did this for a buffer passed to a HW device, I could prove that the hardware was doing DMA-writes where it shouldn't. Had to bring a HW logic analyzer to verify.
- rramadass 3mo agoPart-2 is more than a paragraph and is logically distinct from Part-1. In this, Raymond actually gets the crucial clue from another colleague's debugging efforts which leads him to identify that the bottom byte of HMODULE of the DLL gets overwritten by <something> which is the root cause of the bug; viz. The “DLL unmapped from memory” crash is just an alternate manifestation of the “somebody is writing 01 bytes to places they shouldn’t” bug. The original bug had a larger bucket spray than we initially thought. Part-2 is the essence of the solution while Part-1 is a series of investigations and inferences.