3 ms·
I basically was using the Debug Watchpoint and Trace (DWT) block of the armv7m architecture (the ARMv7-M architecture reference manual specifies it). DWT inclu
by codys 10y ago
I basically was using the Debug Watchpoint and Trace (DWT) block of the armv7m architecture (the ARMv7-M architecture reference manual specifies it).
DWT includes "comparators" typically used as hardware watch points. My methodology was to use one of the DWT comparators to watch for something reading/writing the last u32 of the current task's stack. Note that this might not be perfect as it could be possible for the stack to somehow skip over the last u32.
Then, install a handler for the Debug Monitor exception (which triggers when a watch point is hit) to disable the DWT comparator, print/store the appropriate info, and finally reset the chip.
Essentially, we're behaving like a simple "debug monitor" which only is capable of setting a single hardware watch point automatically.
Note that DWT is technically an optional component in armv7m, so it may not be functional on every chip. That said, most folks like having debug capabilities (this is the same hardware that jtag debuggers use to set watchpoints), so most seem to include them. On that note: doing this can (to some extent) interfere with external debuggers. To mitigate some of that, I did try to use the last DWT comparator instead of the first, but this will really depend on how your debugger works.
- TickleSteve 10y agogood idea about using the DWT.... you're right, I can imagine that it would interfere with debuggers but still a useful technique. I've got hopes for the v8-m arch as it modifies the MPU scheme to make it much more flexible than v7-m. I've had issues with fitting code around the power-of-2 size windows.... it makes things really clumsy. I'm currently using code instrumentation to get our required level of robustness but I may have a play around with the DWT now.
- codys 10y agoDWT for armv7m also uses base+power-of-2-mask. I was in the process of creating some code [1] for calculating the best possible power-of-2-mask (or set of masks + bases) given a limit, but never completed it. 1: https://github.com/jmesmon/ccan/blob/maskn/ccan/maskn/maskn.c https://github.com/jmesmon/ccan/blob/maskn/ccan/maskn/maskn....
- _fs 10y agoI implemented something similar on a coldfire cpu. I was able to protect against the skip over problem you describe by combining the hardware debug point with -fstack-check to create a "canary" zone at the bottom of every stack. This prevents a huge buffer that goes past the bottom of the stack (which would not normally hit your single watchpoint) from overwriting code far outside of the stack space.
- TickleSteve 10y agoyes, that would work for that. In my case, with my domain protection system I had well defined domain-crossing points across which I could guarantee that no data was being passed-by-reference, At these points I was trying to fit the Cortex MPU windows as closely as possible over the 'used' stack area so I could prevent access to it by any potentially malicious code. I would poison the return address to cause a known fault when the code returned and move the MPU windows accordingly. Although it worked, due to the limitations of the v7-M arch MPU windows the protection was never complete and proved to be time consuming to implement due to the size of the application. Maybe in future it will become more practical.