10 ms·
A tale of an impossible bug: big.LITTLE and caching
- xxie24 10y agoFrom the pseudo code, what is disadvantage that making get_current_cpu_cache_line_size() always get called?
- JoshTriplett 10y agoPerformance. get_current_cpu_cache_line_size would need to run some code to determine the cache line size, and that code takes longer to run than using a cached value. Along similar lines, if you have an optimized routine using specific CPU instructions, you don't want to call CPUID (or equivalent) on every call to find out if you have those instructions; you want to call it once and cache the answer. If it can return different answers on different CPUs in the same system, you need to use a different mechanism instead, such as asking the OS for the least common denominator features of available CPUs, or notifying the OS that you need to run on one of the more capable CPUs that has the feature you need.
- cm2187 10y agoStupid question but does that work in a virtualised environment where your program can be live-migrated to another physical machine with a different CPU?
- mjg59 10y agoNope. There's not really any alternative other than "Don't do that", or limit migration to machines that have a superset of the instructions on the original machine.
- cm2187 10y agoI presume that AWS or Azure wouldn't do that?
- pm215 10y agoUsually you configure the VM to only report CPUID values corresponding to lowest-common-denominator features on everything you might want to migrate to. Then as long as the guest code plays nicely and looks at the CPUID feature flags to see what it can use, it'll migrate happily. QEMU has support for this, for instance.
- tcas 10y agoI believe virtual machines will typically alter the CPUID a bit: https://tech.mendix.com/linux/2016/08/18/xen-cpuid-masking/ https://tech.mendix.com/linux/2016/08/18/xen-cpuid-masking/ You'll choose a lowest common denominator of features sets.
- Someone 10y agoAs others said, one typically pretends to run a fixed CPU on all CPUs. Also, for this specific case, I doubt migration will keep around the contents of the cache lines and their 'dirty' bits (corollary: it will be possible to reliably detect a move, if one is willing to continuously run code that detects cache-line timing differences)
- rodrigokumpera 10y agoThat's a question for the libgcc team. I seriously doubt it to be slow enough to matter.
- dmitrygr 10y agoit reads some coproc regs, which are instructions that cannot be reordered. they slow down everything on an OOO core. after that just some bitmasking (not slow)
- pm215 10y agoFor 64-bit ARMv8 (ie AArch64) system registers are in general reorderable; software must provide explicit synchronization (typically via barrier instructions) where it does not want the reordering, except for a few registers which have implicit synchronization. Since CTR_EL0 is entirely constant there's no inherent reason why it shouldn't be reorderable pretty freely, though it's an implementation detail how fast or otherwise it is in practice. (Benchmark if it matters to you!) (This is all documented in the v8 ARM ARM section "Synchronization requirements for AArch64 System Registers".)
- _ihaque 10y agoThat would create a race condition addressed at the bottom of the article: the process can get switched onto another CPU between the invocation of get_current_cpu_cache_line_size() and the invalidation. An astute reader might realize that computing the cache line on every invocation is not enough for user space code: It can happen that a process gets scheduled on a different CPU while executing the __clear_cache function with a certain cache line size, where it might not be valid anymore.
- K0nserv 10y agoThe follow up doesn't make sense to me Therefore, we have to try to figure out a global minimum of the cache line sizes across all CPUs. Wouldn't this mean they'd always just end up clearing half the cache line for larger core anyway?
- tveita 10y agoNo, you just sometimes issue twice as many flush requests as necessary. You can't flush or invalidate half a cache line, since the data is stored in units of cache lines. In theory I think you could just invalidate the addresses byte for byte, ignoring the cache line size, but I assume the performance hit would be noticeable.
- Tuna-Fish 10y agoIt's slow, and also it doesn't fix the bug, as the code could get migrated immediately after calling it. The only real fix is to do what they are doing -- always use the smallest line size of the system, regardless of which core you are running on.
- anarazel 10y agoDifferent cacheline sizes for the different cores seems like an absurdly bad idea. One because it opens one up to bugs like these, but also because it makes optimization a lot harder. I have a hard time believing the savings due to a larger line size are worth it.
- digi_owl 10y agoBest i can tell, the whole thing is designed so that you can go from a a single "little" all the way to two sets of four. Thus each core is designed to be used both independently and in larger configurations.
- gcp 10y agoARM's own designs (A53, A57, A72, A73) all have 64-byte cache line sizes and avoid the problem entirely. The one at fault appears to be Samsung, who designed M1 Mongoose with 128 byte lines and packed it together with A53 cores in their SoC.
- pawadu 10y agoYou could also say that the ARM implementation is not optimal since it uses the same cache size for vastly different designs (different pipeline length and memory access characteristics). Samsung tried to improve performance, I don't think they are at fault at all.
- gcp 10y agoSamsung's design breaks correctly written user-land code. Hence, they're 100% at fault. Going to 128-byte cache lines was fine, but they should have made the lower-power part of the chip match (which in this case they likely couldn't because they licensed that design), or they should have made sure 128-byte lines were reported in all circumstances (which requires a hack similar to the workaround done in the kernel because again they can't make the A53 report 64 bytes).
- sangnoir 10y ago
- pm215 10y agoProperly configured big.LITTLE clusters should be set up so that all CPUs report the same cache line size (which might be smaller than the true cache line size for some of the CPUs), to avoid exactly this kind of problem. The libgcc code assumes the hardware is correctly put together. There is a Linux kernel patchset currently going through review which provides a workaround for this kind of erratum by trapping the CTR_EL0 accesses to the kernel so they can be emulated with the safe correct value: http://www.mail-archive.com/linux-kernel@vger.kernel.org/msg1227904.html http://www.mail-archive.com/linux-kernel@vger.kernel.org/msg... and it seems to me that that's really the right way to deal with this.
- pm215 10y agoAlso, if I'm reading the proposed fix in the mono pull request correctly, it doesn't deal with the problem entirely because there's a race condition where the code might start execution on the core with the larger cache line size, and then get context-switched to the core with the smaller cache line size midway through executing its cache-maintenance loop. The chances of things going wrong are much smaller, but they're still there... (Edit: rereading the blog post, they say they need to figure out the global minimum, but I can't see how their code actually does that, since there's nothing that guarantees that the icache flush code gets run on every cpu before it's needed in anger.)
- deleted 10y ago[deleted]
- hrydgard 10y agoIt's fine if core migration happens during the invalidation loop - the core migration itself surely must wipe the non-shared cache levels thoroughly, otherwise nothing would work. EDIT: Actually if the big and little cores are used together, and not exclusively, then this might still be an issue, yeah.
- rodrigokumpera 10y agoYes and no. Core migration don't need to reach a global synchronization point, just enough so that the 2 cores in question agree with each other. This can be done without requiring global visibility of all operations of the source core.
- deleted 10y ago[deleted]
- dsp1234 10y agoIt appears the that the caching code was added in this patch: https://gcc.gnu.org/ml/gcc-patches/2012-09/msg00076.html https://gcc.gnu.org/ml/gcc-patches/2012-09/msg00076.html Prior to that, the call: asm volatile ("mrs\t%0, ctr_el0":"=r" (cache_info)); was always made.
- willvarfar 10y agoBut the task can be rescheduled on a little core part-way through the execution... The big cores should report the smaller cache line size always.
- dsp1234 10y agoI was just pointing out where the caching of the size was added. I'm not making a comment on where the fix should be.
- sjmulder 10y agoHuge props to the team for finding this out. What a nasty issue. One question about the intro, which states that this is the first mass produced AMP architecture, but isn't the PlayStation 3's Cell CPU one?
- wmf 10y agoThe difference is that big.LITTLE ARMs use the same instruction set on all cores. In general I wouldn't expect a lot of terminological rigor from blog posts.
- masklinn 10y agoCell behaves more like a CPU + GPGPU system, big.LITTLE can schedule an instruction stream on any core of the cluster (depending on the configuration setup), on Cell you'd primary use the general-purpose PPE from which you'd start (and chain) vector-based threads on SPEs. PPE and SPE don't even share an ISA, SPE have a custom-built SIMD-oriented ISA.
- GrumpyNl 10y agoI don't understand that level of coding, what i do understand is the great way of debugging. Its all about deduction mr Watson.
- dsp1234 10y agoWatson was actually Dr. Watson. Which is/was also the name of a debugger in Windows[0]. [0] - https://en.wikipedia.org/wiki/Dr._Watson_(debugger) https://en.wikipedia.org/wiki/Dr._Watson_(debugger)
- wolfgke 10y ago> Worse, not even the ARM ISA is ready for this. An astute reader might realize that computing the cache line on every invocation is not enough for user space code: It can happen that a process gets scheduled on a different CPU while executing the __clear_cache function with a certain cache line size, where it might not be valid anymore. I rather see the problem in the fact that there seems to be no possibility to say to the Linux scheduler: Only schedule this process/thread between cores that have the same cache line size. Or add an attribute when some thread is created on the cache line size of the cores it is allowed to run. Or an attribute when some thread is created for "allow arbitrary cache line size but don't let it run on cores with a different size". This way it would suffice to check for the cache size on program or thread start.
- monocasa 10y agoIn that case, you could read /proc/cpuinfo and set your affinity mask.
- wolfgke 10y agoSuggest this to the authors of the original article. :-)
- tveita 10y agoThat's too specific a feature to expose to developers - most people wouldn't even know about the feature, and the ones that did might needlessly enable it just to be conservative. The intention of the big.LITTLE architecture is to let processes be migrated seamlessly between the small and big core and let the unused core be turned off to save power. The kernel and the hardware should work together to make the core switching transparent and expose a safe way to invalidate the cache independent of the current processor.
- wolfgke 10y ago> The intention of the big.LITTLE architecture is to let processes be migrated seamlessly between the small and big core and let the unused core be turned off to save power. On the other hand flushing specific cache lines is a rather special feature. If the code uses such obscure low-level features (much more than the "typical" application) such as the invariant that the size of a cache-line stays constant over the execution (which most applications really don't care about) one can at least expect from the developer to pass this information to the scheduler so that the scheduler can take that this invariant will indeed be satisfied.
- ChuckMcM 10y agowow, just wow. That is a really awesome bug (and like the authors I have issues with trying to sleep when that sort of puzzle is sitting there :-) Still a bit hazy on why they manually flush the cache for a given block of memory (presumably for protecting disclosure?) but I'm also a bit curious how it works if you get the sequence big fetches a cache line, switches to little which fetches a line (half as long and changing half the bytes in the cache) and then you switch back to big and its thinking it has a full cache line? Presumably there is some mechanism that invalidates cache lines?
- pm215 10y agoHandling of the case where big and little both want the same thing in their cache should be dealt with by the usual cache-coherency traffic between the CPUs that ensures they don't disagree about what's in their L1 caches (very handwaved because I don't know the details). The reason for the manual cache operations is because they're generating JITted code -- on ARM to ensure that what you execute is the same thing you just wrote you have to (1) clean the data cache, so your changes get out to main memory[.] and then (2) invalidate the icache, so that execution will fetch the fresh data from memory rather than using stale info. This clean-and-invalidate operation is usually informally called a flush, though it isn't really one in ARM terminology. [.] not actually main memory, usually: only has to go out to the "point of unification" where the iside and dside come together, which is probably the L2 cache.
- ChuckMcM 10y agoThanks for the update on the manual cache management! I'll admit that I find the question of dissimilar cache line sizes into the same cache intriguing from an architecture point of view. It has me doodling all sorts of questions into my notebook.
- brandmeyer 10y agoI don't see why they have to do this in userspace at all. If they did: * allocate read/write buffer * JIT instructions into it * change mapping to read/execute * run the JITted code Then the kernel manages flushing the data caches on the mapping change, and Mono gets to wrap a Somebody Else's Problem field around it. It sounds like they are instead: * allocate read/write/execute buffer * JIT instructions into it * manually flush relevant data caches (with an assumption that the cache line size is constant) * run the JITted code
- AceJohnny2 10y agoAs an embedded developer also working on big.LITTLE arm64... I would do nasty, painful things to the designer of such a system.
- Mizza 10y agoThis is an excellent bug journey, and I'm even more impressed that the resulting discovery has already been used to improve Dolphin. A testament to the quality of both projects.
- datenwolf 10y agoI'm so sorry, but I simply cannot resist: Yo Dawg, I heard you like cache invalidation problems! So I put a cache invalidation problem in your cache invalidation problem.
- SixSigma 10y agoThere's no such thing as a simple cache bug. - Rob Pike Caches are bugs waiting to happen. Rob Pike @rob_pike 21 Mar 2014
- akavel 10y ago"There are only two hard things in Computer Science: cache invalidation and naming things" ― Phil Karlton; not sure to what extent this quote is compatible with the second one by Rob. Or does it mean by implication that simply "Computer Science is bugs waiting to happen"?...
- mikeash 10y agoIt's a great quote, but it's wrong. There are actually two hard things in CS: cache invalidation, naming, and off-by-one errors.
- deathanatos 10y agoThe parent actually appears to have the quote both correct in content and attribution. Supposedly, someone else added the "off by one"[1][2][3]. That seems to be the extent of the Internet's knowledge on the quote, though the Skeptics link notes that there's nothing direct to the supposed originator. This is one of those quotes where I feel there's more than one right answer. I like the addition of "off-by-one", and to make it a nice round three things, I usually use this version: There are three hard things in computer science: 1. Naming things 2. 3. Concurrency Cache Invalidation 4. Off-by-one errors [1]: https://twitter.com/timbray/status/506146595650699264 https://twitter.com/timbray/status/506146595650699264 [2]: https://skeptics.stackexchange.com/questions/19836/has-phil-karlton-ever-said-there-are-only-two-hard-things-in-computer-science https://skeptics.stackexchange.com/questions/19836/has-phil-... [3]: http://martinfowler.com/bliki/TwoHardThings.html http://martinfowler.com/bliki/TwoHardThings.html
- mikeash 10y agoThe other one is the original, I just like the "off by one" version better. I like the addition of "concurrency" too, but I'm not quite sure how to make it flow....
- flamedoge 10y agoI wonder why no one tried to validate Asymmetric MultiProcessing by first validating all cases with either little or big first. And then bisect further down when both are enabled.
- PDoyle 10y agoBecause that's the sort of thing you think of only when you already know the answer.
- Animats 10y ago"first mass produced AMP architecture" Nope. Remember the Cell? The processor in the Playstation 3? One main CPU with 8 little CPUs and no shared memory, just channels. The Playstation 4 isn't a AMP machine because programming the Cell was so hard.
- maximilianburke 10y ago> The Playstation 4 isn't a AMP machine because programming the Cell was so hard. The PS4 could be an AMP design if you consider the (closely coupled) GPU a processor. It doesn't require the same gymnastics as the PS3 though because both the main processor and GPU are more capable. The SPUs were required to perform computation that the anemic PPU could not do as well as fill in where the pre-unified shader model GPU was unable to keep the pace. Memory was shared on the PS3 but, from the SPUs, required explicit put and fetch operations.
- dfox 10y agoCell is essentially distributed memory cluster on single chip, because each SPU has it's own address space and cannot directly access main memory. I'm not sure about what the exact definition of AMP is, but it does not exactly match my feeling of what AMP should be. In this regard Wii seems more like AMP systems with two completely different CPUs (PPC and ARM) sharing what essentially amounts to be same address space (and in WiiU there are 3 PPC cores where one of them is slightly different than other two and cache coherency between them can only be described as broken). There is no question of hardness of programming for Cell, but I think it's mostly about middleware support (probably because the platform is so different from PC and xbox360).
- Animats 10y agoThe real problem with the Cell was that each Cell SPE processor only has 256K of local memory. It has bulk DMA access to main memory, but that's more like I/O. 256K is too small for a video frame, a game level, or much else in a modern game. So everything has to be done on an assembly line basis, where data is pumped into a Cell processor, processed, and pumped out. Great for audio, terrible for everything else. In comparison, the main processor had access to 256MB of RAM. If they'd had, say, 16MB per processor, it might have worked out. One CPU for collision detection and physics, one for NPC management and AI, etc. But giving each SPE processor only 0.1% of the total memory space was too constraining.
- randyrand 10y agoWhen would a programmer explicitly need to invalidate the CPU cache? Does this not happen automatically on a context switch?
- JoeAltmaier 10y agoOne example is when doing a dma operation to memory. That often bypasses the cache. If the new data is to 'stick' the cache needs to be convinced it doesn't know what's in that buffer any more. This is an ARM thing; Intel architectures integrate DMA with the cache.
- dragon_ninja 10y agoThis is so stupid.
- userbinator 10y agoSeeing bugs like this reminds me of how much nicer things are on x86 where JITs do not need to flush caches. You can actually modify the instruction immediately ahead of the currently executing one, and the CPU will naturally "do the right thing"[1] --- it does slow down execution, as the CPU is essentially automatically detecting and flushing its cache/pipeline, but used sparingly can be a great optimisation. The write can even come from another core ("cross-modifying code") and everything will still work. Someone I know used this to great effect in squeezing out the last bits of performance from an application by eliminating checks on a few flag variables and the associated branching in a tight loop --- it simply "poked" instruction bytes from another core into the loop when it was time for that core to do something else. [1] With the exception of pre-Pentium CPUs, where modifying at various forward offsets from the locus of execution could give insight into how big the prefetch queue is. With the Pentium it was fully detected and the later multithreaded/multicored react similarly to cross-modifying code as described above, which leads me to believe that Intel is very much supportive of these things as otherwise they could've just told programmers to do as ARM does. Maybe what ARM needs, short of doing it the Intel way, is a "flush region" instruction which takes both the address and size, so it can automatically flush the appropriate cache lines based on the current hardware's cacheline size.
- pm215 10y agox86 is really the oddity here -- it has its no-explicit-cache-maintenance design because of wanting to maintain backwards-compatibility with self-modifying code that was written for x86 cores that had no caches at all. Almost all other architectures have explicit cache maintenance because it's more efficient (and requires less hardware), at the minor cost of requiring the very few bits of software which do odd things like JITting to explicitly tell the CPU what they're doing. An instruction for flushing an entire region would potentially have a very long execution time, which is awkward because you would want to be able to interrupt and resume it. So it would need "how far have I got" state stored somewhere. The obvious observation from a RISC-architecture point of view is that you can get the equivalent effect without the pain of making a long-running interruptible instruction, by having an "invalidate one cache line" instruction plus an explicit loop in the code, and that's what most architectures do.
- betolink 10y agoCan someone explain why cache flush is used for on ARM or in general low level programming?
- wolf550e 10y agoOverwriting executable machine code in memory is a rare case (only JITs need it, only for emitting code). Some CPUs do wire the memory write operation to the instruction cache to make sure the instruction cache knows some code changed, but ARM chose a simpler design where the instruction cache assumes the code never changes in memory and if this assumption is wrong, the programmer has to use a special instruction to tell the instruction cache to throw something out of cache. The cache clearing instruction throws out either 64 bytes aligned to multiple of 64 or 128 bytes aligned to multiple of 128, depending on the core.
- hexa00 10y agoI had this problem too with GDB on the Odroid UX4 big.LITTLE SoC. Since GDB is patching the instrution with ptrace to insert a breakpoint for example. See my blog post about it: https://www.kayaksoft.com/blog/2016/05/11/random-sigill-on-arm-board-odroid-ux4-with-gdbgdbserver/ https://www.kayaksoft.com/blog/2016/05/11/random-sigill-on-a... Or the post on the GDB mailling list: https://www.sourceware.org/ml/gdb/2015-11/msg00030.html https://www.sourceware.org/ml/gdb/2015-11/msg00030.html Too bad however that the kernel patchset mentionned in a previous post only covers arm64.. So it's still a problem from arm32.
- funny_falcon 10y agohttps://gcc.gnu.org/ml/gcc-patches/2012-09/msg00076.html https://gcc.gnu.org/ml/gcc-patches/2012-09/msg00076.html
- pawadu 10y agoWell, that is the patch that causes this problem... At the time this patch was submitted, ARM was running a program that awarded engineers who could improve performance of aarch64.