3 ms·
I'm looking at the C maxArray example and see that the for loop: for (i = 0; i < 65536; i++) gets compiled as .L2: add rax, 8 cmp
by bori 4y ago
I'm looking at the C maxArray example and see that the for loop:
for (i = 0; i < 65536; i++)
gets compiled as
.L2:
add rax, 8
cmp rax, 524288
jne .L4
Is there any fundamental reason why the loop increments with 8 instead of 1? Or is this completely arbitrary? Thanks.
- kazinator 4y agoBecause it's working with arrays of double which is 8 bytes wide, and the optimized loop doesn't use any scaled indexing; basically the loop index is pre-scaled to the element size. Perhaps those MMX instructions don't support addressing modes with index scaling (wild guess). The unoptimized code increments by 1 to 65535, but the memory accesses use scaling. Well, not exactly. We see this: lea rdx, [0+rax*8] mov rax, QWORD PTR [rsp-120] add rax, rdx movsd xmm0, QWORD PTR [rax] This LEA here, though it means "load effective address" is not actually an effective address calculation. The base address is zero, and so this is just LEA being exploited to multiply RAX by 8, and get that into RDX. RAX is then clobbered with the base address of an array, to which the scaled displacement is added and then finally used to make an access.