3 ms·
It would just be slow for the caller to have to push and pop (say 30) registers in general that the specific callee (and transitive callees) may not even use.
by dannymi 4y ago
It would just be slow for the caller to have to push and pop (say 30) registers in general that the specific callee (and transitive callees) may not even use.
Most ABI specify some registers caller-saved, some registers callee-saved (retained unchanged from the perspective of the caller) and some registers scratch (not-saved).
In the end it depends on the architecture and on typical workload which are the fastest--and measurements can be made and it can be found out which combination is the fastest on average.
- dzaima 4y agoYou usually wouldn't need to push&pop all registers, just ones that you want to preserve across the call. Regardless, yeah, non-volatile aka callee-saved registers are extremely important for good performance of code that calls functions (esp. loops - without callee-saved registers, you'd have to store the loop counter & length on the stack!)
- dreamcompiler 4y ago...which is exactly what the callee within the loop would have to do if it used those registers. But I guess your point is that in the callee-save case it only happens when it needs to happen, while in the caller-save case it happens every time whether it needs to or not.
- dzaima 4y agoyep, hence why calling conventions usually have both callee-saved and caller-saved registers, so that you only have "unnecessary" stack usage when you need to use more than roughly half of the registers.
- not2b 4y agoProgramming defensively in this way is possible but has a huge cost (extra saves and restores). The caller is supposed to be able to trust that certain registers persist across the call. If it can't, it has to save everything to memory before the call and restore it. I suppose a compiler could add a special annotation for "this is an assembler routine and I don't trust that they know the rules", which would generate extra saves and restores, but presumably the routine was coded in assembly for extra speed, so 6 to 8 extra saves and restores would cancel that out. If someone is trying to use the same assembler routine on Windows as on Linux, it's likely to be wrong for one of them.