5 ms·
From the pseudo code, what is disadvantage that making get_current_cpu_cache_line_size() always get called?
by xxie24 10y ago
From the pseudo code, what is disadvantage that making get_current_cpu_cache_line_size() always get called?
- JoshTriplett 10y agoPerformance. get_current_cpu_cache_line_size would need to run some code to determine the cache line size, and that code takes longer to run than using a cached value. Along similar lines, if you have an optimized routine using specific CPU instructions, you don't want to call CPUID (or equivalent) on every call to find out if you have those instructions; you want to call it once and cache the answer. If it can return different answers on different CPUs in the same system, you need to use a different mechanism instead, such as asking the OS for the least common denominator features of available CPUs, or notifying the OS that you need to run on one of the more capable CPUs that has the feature you need.
- cm2187 10y agoStupid question but does that work in a virtualised environment where your program can be live-migrated to another physical machine with a different CPU?
- mjg59 10y agoNope. There's not really any alternative other than "Don't do that", or limit migration to machines that have a superset of the instructions on the original machine.
- cm2187 10y agoI presume that AWS or Azure wouldn't do that?
- pm215 10y agoUsually you configure the VM to only report CPUID values corresponding to lowest-common-denominator features on everything you might want to migrate to. Then as long as the guest code plays nicely and looks at the CPUID feature flags to see what it can use, it'll migrate happily. QEMU has support for this, for instance.
- tcas 10y agoI believe virtual machines will typically alter the CPUID a bit: https://tech.mendix.com/linux/2016/08/18/xen-cpuid-masking/ https://tech.mendix.com/linux/2016/08/18/xen-cpuid-masking/ You'll choose a lowest common denominator of features sets.
- Someone 10y agoAs others said, one typically pretends to run a fixed CPU on all CPUs. Also, for this specific case, I doubt migration will keep around the contents of the cache lines and their 'dirty' bits (corollary: it will be possible to reliably detect a move, if one is willing to continuously run code that detects cache-line timing differences)
- rodrigokumpera 10y agoThat's a question for the libgcc team. I seriously doubt it to be slow enough to matter.
- dmitrygr 10y agoit reads some coproc regs, which are instructions that cannot be reordered. they slow down everything on an OOO core. after that just some bitmasking (not slow)
- pm215 10y agoFor 64-bit ARMv8 (ie AArch64) system registers are in general reorderable; software must provide explicit synchronization (typically via barrier instructions) where it does not want the reordering, except for a few registers which have implicit synchronization. Since CTR_EL0 is entirely constant there's no inherent reason why it shouldn't be reorderable pretty freely, though it's an implementation detail how fast or otherwise it is in practice. (Benchmark if it matters to you!) (This is all documented in the v8 ARM ARM section "Synchronization requirements for AArch64 System Registers".)
- _ihaque 10y agoThat would create a race condition addressed at the bottom of the article: the process can get switched onto another CPU between the invocation of get_current_cpu_cache_line_size() and the invalidation. An astute reader might realize that computing the cache line on every invocation is not enough for user space code: It can happen that a process gets scheduled on a different CPU while executing the __clear_cache function with a certain cache line size, where it might not be valid anymore.
- K0nserv 10y agoThe follow up doesn't make sense to me Therefore, we have to try to figure out a global minimum of the cache line sizes across all CPUs. Wouldn't this mean they'd always just end up clearing half the cache line for larger core anyway?
- tveita 10y agoNo, you just sometimes issue twice as many flush requests as necessary. You can't flush or invalidate half a cache line, since the data is stored in units of cache lines. In theory I think you could just invalidate the addresses byte for byte, ignoring the cache line size, but I assume the performance hit would be noticeable.
- Tuna-Fish 10y agoIt's slow, and also it doesn't fix the bug, as the code could get migrated immediately after calling it. The only real fix is to do what they are doing -- always use the smallest line size of the system, regardless of which core you are running on.