3 ms·
I‘ve recently tried implementing the Noise protocol on a 7.8 MHz 68000, while it runs in 10 seconds I don’t think there is a way to make it constant time withou
by Herdinger 2mo ago
I‘ve recently tried implementing the Noise protocol on a 7.8 MHz 68000, while it runs in 10 seconds I don’t think there is a way to make it constant time without using addition and running an order of magnitude more slowly :(
MULU operations are not constant time but depend on popcount. That has an easy work around just also do the MULU of the bitwise inverse, both instructions sum to a constant cycle count.
The brick wall I hit is related to the Macintosh SE ram being shared between CPU and video system. Every few cycles the CPU stalls instruction fetching, which turns the both MULUs together take the same amount of time each time into there is slowdown depending on the popcount of the first instruction.
I don’t know if anyone has any solution for this, there is prefetching of one instruction so all should be well, but it seems like the cpu stalls the in progress instruction if the prefetch is stalled.
- boomlinde 2mo agoIf there is a portion of the frame where the cycle stealing doesn't occur, e.g. during vertical blanking, maybe you can schedule your multiplications to only happen there. It will not strictly be constant time of course, and it will be slower, but timing would be invariant of input. Timing characteristics would instead reveal to the attacker where the raster beam was when the calculation started :)
- Herdinger 1mo agoThat’s a great idea but that takes away a lot of compute! I didn’t pursue this further since at that point we might as well do the adds vs muls. I thought about calculating the worst case execution time by hand and than just doing a literal report solution on a timer after time has passed The issue with that strategy is that I would like the OS to stay responsive (and there is interrupts that fall into that time frame) so could overshoot the budget and then we’re back to square zero. I think at this point I probably have no choice other than unrolling it into addition :( There is mitigations for those kind of things that make delta analysis impossible but I would REALLY just like the perfect solution that doesn’t depend on the secret for timing at all.