4 ms·
> for arr[i] = 2, you'll still need to subtract 1 You can, but don't need to. Just have compiler store array pointer constant as (arr-1) instead of arr, et voi
by nousermane 4y ago
> for arr[i] = 2, you'll still need to subtract 1
You can, but don't need to. Just have compiler store array pointer constant as (arr-1) instead of arr, et voila, zero runtime overhead.
Or, for modern(-ish) ISAs, often you can add/substract a small constant at runtime, with no extra cycles taken. For example, for x86_64:
# rbp contains "true" pointer to arr
# rax contains 1-based array index
mov rax, QWORD PTR [rbp-8+rax*8]
- shpx 4y agoWhat if the variable is used for array indexing and something else, like displayed to the user?
- murderfs 4y agoIt's basically only x86 among modern ISAs that lets you do base + literal + register * literal, aarch64 only gives you base + register shifted by a literal, and I believe RISC-V is similar: https://gcc.godbolt.org/z/dz768768z https://gcc.godbolt.org/z/dz768768z
- snvzz 4y ago>and I believe RISC-V is similar Yes it is. It was evaluated, carefully weighted and discarded, as it was not worth it.
- kaba0 4y agoBut it will still execute with likely no extra time at all due to OOE and how fast arithmetics are.
- NohatCoder 4y agoIf it is not on the hot path, it is likely free, but not guaranteed. If it is on the hot path then it is wasting a whole cycle. And of course in highly ALU-dependent code it is another instruction, so a fraction of a clock.
- kaba0 4y agoWhat do you mean it wastes a whole cycle? It may indeed have worse performance due to blowing the instruction cache, but I don’t see why would out-of-order execution be slower on the hot path - I doubt there would be too many hot paths without any dependence on memory fetches outside specific benchmarks - the memory loads will take significantly more time even if they hit cache.
- murderfs 4y agoOOE doesn't necessarily save you if you end up with a hard dependency on the value of the read (and even if it did, the little cores on ARM SoCs are in-order). This is a pretty obvious candidate for macro-op fusion, but I'm not sure whether this actually happens (and if it happens on ARM little cores, etc.)
- astrobe_ 4y agoAnother option for users is to "sacrifice" the 0th entry of an array. Depending on the size(as in sizeof) of the entry, it can be worth it. A benefit is that you can then use 0 as a sentinel value; for instance if you have a find() routine that surely can fail, it can just return 0 instead of having e.g. -1 (which can introduce minor issues). In my experience, though, I am so used to 0-based index that switching schemes can cause stupid off-by-one bugs. I guess that's the main reason behind complains about Lua. It's not that "natural" arrays are thought of as bad, but mixing both schemes (often C an Lua) is error-prone.