5 ms·
It's really nice to have these things written out, the interfaces used by gcc options that insert extra code/calls are typically undocumented (except for their
by codys 10y ago
It's really nice to have these things written out, the interfaces used by gcc options that insert extra code/calls are typically undocumented (except for their source code).
kasan, asan, and ubsan are similarly lacking in docs (but are also composed in much the same way). I imagine some of the profiling hooks are done similarly, but I have not yet looked into how to use them.
It'd be really nice if gcc considered these semi-public interfaces that they should document.
> I plan to add other other debugging features in the following weeks. I have an idea on how to use the MPU to prevent thread stack overflows (different from the buffer stack overflow we explored today).
On that note, one can actually abuse the debug hardware on cortex-m3 chips without an MPU to add stack overflow protection. Simply need some hooks into the scheduler around where it switches tasks (some already poison the top of their stacks & check for the poison on switching, the keep-out can typically be applied very close to these locations).
The debug hardware takes a base+mask, so one can theoretically block access to wider spans of memory that a single u32, but at that point calculating the best way to cover that memory range becomes a bit difficult.
- TickleSteve 10y agoCurious to know about the debug-abuse trick, I've implemented an MPU-based stack protection scheme that involves poisoning the return address to trap returns and implement a 'sliding' stack window. It does work, but due to the limitations of Cortex MPU windows I've moved over to code instrumentation for our protection which involves a very much beefed-up version of the GCC stack protector.
- codys 10y agoI basically was using the Debug Watchpoint and Trace (DWT) block of the armv7m architecture (the ARMv7-M architecture reference manual specifies it). DWT includes "comparators" typically used as hardware watch points. My methodology was to use one of the DWT comparators to watch for something reading/writing the last u32 of the current task's stack. Note that this might not be perfect as it could be possible for the stack to somehow skip over the last u32. Then, install a handler for the Debug Monitor exception (which triggers when a watch point is hit) to disable the DWT comparator, print/store the appropriate info, and finally reset the chip. Essentially, we're behaving like a simple "debug monitor" which only is capable of setting a single hardware watch point automatically. Note that DWT is technically an optional component in armv7m, so it may not be functional on every chip. That said, most folks like having debug capabilities (this is the same hardware that jtag debuggers use to set watchpoints), so most seem to include them. On that note: doing this can (to some extent) interfere with external debuggers. To mitigate some of that, I did try to use the last DWT comparator instead of the first, but this will really depend on how your debugger works.
- TickleSteve 10y agogood idea about using the DWT.... you're right, I can imagine that it would interfere with debuggers but still a useful technique. I've got hopes for the v8-m arch as it modifies the MPU scheme to make it much more flexible than v7-m. I've had issues with fitting code around the power-of-2 size windows.... it makes things really clumsy. I'm currently using code instrumentation to get our required level of robustness but I may have a play around with the DWT now.
- codys 10y agoDWT for armv7m also uses base+power-of-2-mask. I was in the process of creating some code [1] for calculating the best possible power-of-2-mask (or set of masks + bases) given a limit, but never completed it. 1: https://github.com/jmesmon/ccan/blob/maskn/ccan/maskn/maskn.c https://github.com/jmesmon/ccan/blob/maskn/ccan/maskn/maskn....
- _fs 10y agoI implemented something similar on a coldfire cpu. I was able to protect against the skip over problem you describe by combining the hardware debug point with -fstack-check to create a "canary" zone at the bottom of every stack. This prevents a huge buffer that goes past the bottom of the stack (which would not normally hit your single watchpoint) from overwriting code far outside of the stack space.
- TickleSteve 10y agoyes, that would work for that. In my case, with my domain protection system I had well defined domain-crossing points across which I could guarantee that no data was being passed-by-reference, At these points I was trying to fit the Cortex MPU windows as closely as possible over the 'used' stack area so I could prevent access to it by any potentially malicious code. I would poison the return address to cause a known fault when the code returned and move the MPU windows accordingly. Although it worked, due to the limitations of the v7-M arch MPU windows the protection was never complete and proved to be time consuming to implement due to the size of the application. Maybe in future it will become more practical.
- brandmeyer 10y ago> It'd be really nice if gcc considered these semi-public interfaces that they should document. Once upon a time, the LSB documented the libssp interface, but I cannot find it now.