4 ms·
Value types also give you more control over cache locality. Not having to travel that extra pointer redirection can really add up depending on what you're doing
by hermitdev 8y ago
Value types also give you more control over cache locality. Not having to travel that extra pointer redirection can really add up depending on what you're doing.
- namibj 8y agoThere is at least half a magnitude in speed benefit from a cache-line-size/alignment tuned B-tree compared to e.g. a red/black tree or similar, if used to 'only' store pointer-size value types, and even more if the value type takes up even less space. A case would be e.g. a float/int pair ordered by floats and used as a priority queue or such. That's an 8-byte-large value type. Though, it might, for that case, need to be doubled up if there is a uniqueness constraint on the integers associated with the floats. LLC cache pressure can be very real. I highly recommend perf stat -dd and perf top -g -e uops_executed.stall_cycles as well as the -e cycle_activity.stalls_ldm_pending. The former counts how long stalls happen, including those where instructions like DIV just take long to execute, the second counts stall events, i.e. one stall per load/store, not weighted for how long this stall event actually stalled the core. I'm pretty sure the former ignores cases where just one hyperthread stalls, but I'm not as sure for the latter. On Haswell and newer, unless using branch taken/not-taken profiling, I recommend setting perf config --system call-graph.record-mode=lbr before using perf top -g and perf record -g, as -fomit-frame-pointer binaries cause issues otherwise.