11 ms·
Arm AArch64 Adds Memcpy() Instructions
- cjensen 5y agoInstructions like this need to be interruptible since they take longer than standard instructions. I assume the ARM designers have thought about this?
- addaon 5y agoThe old ARM<=7 load multiple / store multiple instructions were interruptible on most implementations. My recollection is that some implementations checkpointed and resumed, but at least the smaller cores tended to do a full-restart (so no guarantee of forward progress when approaching livelock). I'd expect the same here, with perhaps more designs leaning towards checkpointing.
- baybal2 5y agoIt's well known in the ARM world, and it's the reason we were complaining for years about impossibility of using DMA controller from userspace to do large memcpys. More importantly today, using DMA to do large memcpy for non-latency-sensitive tasks allows cores to sleep more often, and it's a godsend for I/O intensive stuff like modern Java apps on Android which are full of giant bitmaps.
- brandmeyer 5y agoARMv7-M squirrels away the progress of the ldm/stm in the program status register to avoid restarting it completely.
- brandmeyer 5y agoARM has been managing interruptible instructions with partial execution state for a long time. In ARM assembly syntax, the exclamation point in an addressing mode indicates writeback. Its difficult to be certain without seeing the architecture reference manual, but it would be consistent for instruction to be writing back all three of the source pointer, destination pointer, and length registers. A memcpy is interruptible without replaying the entire instruction (say, because it hit a page that needed to be faulted-in by the operating system) if it wrote back a consistent view of all three registers prior to transferring control to an interrupt handler.
- nneonneo 5y agoThese could be really great if they get optimized well in hardware - as single instructions, they’d be easy to inline, reducing both function call overhead and code size all at once. I do wish they’d included some documentation with this update so it’d be clearer how these instructions can be used, though.
- addaon 5y agoThus begins the slide from RISC to (what POWER/PowerPC ended up calling) FISC. It's not about reducing the instruction set, it's about designing a fast instruction set with easy-to-generate, generalizable instructions. Even more than PowerPC (which generally added interesting but less primitive register-to-register ops), this is going straight to richer memory-to-memory ops.
- marcodiego 5y ago> Thus begins the slide from RISC to (what POWER/PowerPC ended up calling) FISC. You mean from RISC to CISC, right?
- jamesfinlayson 5y agoLooks like FISC is Fast Instruction Set Computing (maybe - all I could find was a Medium article that says that).
- galdosdi 5y agoFair, haha. But I think the distinction intended lies in that the old CISC ISAs were complex out of a desire to provide the assembly programmer ergonomic creature comforts, backwards compatibility, etc. Today's instruction sets are designed for a world where the vast majority of machine code is generated by an optimizing compiler, not hand crafted through an assembler, and I think that was part of what the RISC revolution was about.
- addaon 5y agoNo, although one could make that argument. RISC (reduced instruction set) has a few characteristics besides just the number of instructions -- most "working" instructions are register-to-register, with load/store instructions being the main memory-touching instructions; instructions are of a fixed size with a handful of simple encodings; instructions tend to be of low and similar latency. CISC starts at the other side -- memory-to-register and memory-to-register "working" instructions, variable length encodings, instructions of arbitrary latency, etc. FISC ("fast instruction set") was a term used for POWER/PowerPC to describe a philosophy that started very much with the RISC world, but considered the actual number of instructions to /not/ be a priority. Instructions were freely added when one instruction would take the place of several others, allowing higher code density and performance while staying more-or-less in line with the "core" RISC principles. None of the RISC principles are widely held by ARM today -- this thread is an example of non-trivial memory operations, Thumb adds many additional instruction encodings of variable length, load/store multiple already had pretty arbitrary latency (not to mention things like division)... but ARM still feels more RISC-like than CISC-like. In my mind, the fundamental reason for this is that ARM feels like it's intended to be the target of a compiler, not the target of a programmer writing assembly code. And, of the many ways we've described instruction sets, in my mind FISC is the best fit for this philosophy.
- geerlingguy 5y agoMaybe someone at Arm has sympathy on my plight to get graphics cards running on a Pi—I've had to replace memcpy calls (and memset) to get many parts of the drivers to work at all on arm64. Note that the Pi also has a not-fully-standard PCIe bus implementation, so that doesn't really help things either.
- Ballas 5y agoI have watched most of your videos on the subject but cannot recall - have you tried running an external GPU on a Nvidia Jetson? Perhaps that is a place to start? (or perhaps I am just letting my ignorance on the matter show)
- gradschoolfail 5y agoYou mean using the M.2 Key E Slot? Been trying to find a good adaptor for that, any recommendations? (Including nonstandard carrier boards with 8x pcie slots, etc)
- consp 5y agoIf you want to do a bit DIY and go cheap you can but the m.2 to SSF-8643 cards and Dremel a proper keying into it and get a linkreal LRFC6911 for the pcie slot side.
- cbm-vic-20 5y agoSounds like a job for Red Shirt Jeff.
- Ballas 5y agoThat is a possibility, but if you like to do things the easy way there is a full-size PCIe slot on the Xavier AGX dev kit.
- girvo 5y agoAm I nuts or is that like $5000 AUD? The only seller I found with a quick search anyway Ninja edit: Apparently its $699 USD but it's been scalped by third party sellers. A shame, I'd love to pick one up for some of the work I'm doing!
- userbinator 5y ago...so it only took them over three decades to realise the power of REP MOVS/STOS? ;-) On x86, it's been there since the 8086, and can do cacheline-sized pieces at a time on the newer CPUs. This behaviour is detectable in certain edge-cases: https://repzret.org/p/rep-prefix-and-detecting-valgrind/ https://repzret.org/p/rep-prefix-and-detecting-valgrind/
- gatronicus 5y agoExcept that for decades REP MOVS/STOS were avoided on x86 because they were much slower than hand written assembly. This only changed recently.
- userbinator 5y agoThat was really only in the 286-486 era. On the 8086 it was the fastest, and since the Pentium II, which introduced cacheline-sized moves, it's basically nearly the same as the huge unrolled SIMD implementations that are marginally faster in microbenchmarks. Linus Torvalds has some good comments on that here: https://www.realworldtech.com/forum/?threadid=196054&curpostid=196566 https://www.realworldtech.com/forum/?threadid=196054&curpost...
- josefx 5y agoLinus seems to consider rep mov still too slow for small copies: https://www.realworldtech.com/forum/?threadid=196054&curpostid=196611 https://www.realworldtech.com/forum/?threadid=196054&curpost... https://www.realworldtech.com/forum/?threadid=196054&curpostid=196616 https://www.realworldtech.com/forum/?threadid=196054&curpost... It seems to me that rep move is so bad that you want to avoid it, but trying to write a fast generic memcpy results in so much bloat to handle edge cases that rep move remains competitive in the generic case.
- wbsun 5y agoA reminder that ARM is short for Advanced RISC Machines or previously Acorn RISC Machine[1]. [1]: https://en.wikipedia.org/wiki/ARM_architecture https://en.wikipedia.org/wiki/ARM_architecture
- zibzab 5y agoCISC jokes aside, this an interesting turn of events. Classic ARM had LDM/STM which could load/store from a list of registeres. While very handy, it was a nightmare from a hardware POV. For example, it made error handling and rollback much much more complex in out-of-order implementations. ARMv8 removed those in aarch64 and introduced LDP/STP which only handled two registers at a time (the P is for Pair, M for multiple). This made things much easier but it seems the performance hit was not negligible. Now with v8.8 and v9.3 we get this, which looks much nicer than intels ancient string functions that have been around since 8086. But I am curious how it affects other aspects of the CPU, specially those with very long and wide pipelines.
- dvdkhlng 5y agoNote that in ARM-based controllers, LDM/STM also have a non-negligible impact on interrupt latency. These are defined in a way that they cannot be interrupted mid-instruction, so worst-case interrupt latency is higher that would be expected with a RISC CPU (especially if LDM/STM happen to run on a somewhat slower memory region) AFAICS x86 "rep" prefixed instructions are defined so that they can in fact be interrupted without problems. The remaining count is kept in (e)cx, so just doing an iret into "rep stosb" etc. will continue its operation. I think VIA's hash/aes instruction set extension also made use of the "rep" prefix and kept all encryption/hash state in the x86 register set, so that they could in fact hash large memory regions on a single opcode without hampering interrupts.
- duskwuff 5y ago> These are defined in a way that they cannot be interrupted mid-instruction... Usually. Cortex-M3 and M4 cores allow LDM/STM to be interrupted by default, and offer a flag to disable that (SCB->ACTLR.DISMCYCINT). https://developer.arm.com/documentation/ddi0439/b/System-Control/Register-descriptions/Auxiliary-Control-Register--ACTLR https://developer.arm.com/documentation/ddi0439/b/System-Con...
- dvdkhlng 5y agoYes, looks like they added a few bits to the PSR register to capture the internal state of LDM/STM . Like a small version of x86's CS register.
- colonwqbang 5y agoIt's strange that such features seem to not be standard in CPUs. I wonder why? Copy-based APIs are not ideal but they seem to be hard to avoid. In those ARM cores I've programmed, the core has a few extra DMA channels which can be used for such things. However, using them from userspace has always seemed a bit of a hassle.
- hannob 5y agoI haven't done assembler for a long time, but if my memory serves me well on x86 there's the rep movsb commands that will do effectively a memcpy-like operation.
- hyperman1 5y agoCorrect. There is the whole rep family doing all kinds of fun stuff. You can add the rep/repnz/repz prefixes to at least: movs[b|w|d]: move data in bytes/words/doublewords aka memcpy stos[b|w|d]: put a value in bytes/words/dwords aka memset cmps[b|w|d]: compare values aka memcmp scas[b|w|d]: scan for a value aka memchr ins[b|w|maybe d]: read from IO port outs[b|w|maybe d]: write to IO port. lods[b|w|d] : read from memory was probably not meant to be combined with rep as it would just throw everything but the last byte away. I once saw a rep lodsb to do repeated reads from EGA video ram. The video card saw which bytes were touched and did something to them based on plane mask. This way touching 1 bit changed the color of a whole 4 bit pixel, speeding up things with a factor 4. Then one day, someone found that rep movs was not the fastest way to copy data on an x86 and they all went out of vogue. I think rep stos recently came back as fastest memset, as it had a very specific CPU optimization applied. Update: See https://stackoverflow.com/questions/33480999/how-can-the-rep-stosb-instruction-execute-faster-than-the-equivalent-loop https://stackoverflow.com/questions/33480999/how-can-the-rep...
- vardump 5y ago"Copy-based APIs are not ideal but they seem to be hard to avoid." If everything resides in CPU L1 cache, it hardly matters at all. Other than L1 cache pressure, of course. Other example is copying DMA transferred data followed by immediately consuming said data. Also in this case, the copy often effectively just brings the data to the CPU cache and the consuming code reads from cache. Of course it does increase overall memory write bandwidth use when the cache line(s) are eventually evicted, but total performance degradation can be pretty minimal for anything that fits in L1.
- forty 5y agoExtended Zawinski's Law: "Every Instruction Set attempts to expand until it can read mail. Those Instruction Set which cannot so expand are replaced by ones which can." ;)
- cmrdporcupine 5y agoI feel like every instruction set eventually becomes VAX.
- CoastalCoder 5y agoIs it true that VAX allowed customers to extend the ISA themselves? I think I learned about that many years ago, but I couldn't find anything about that recently when skimming the 11/780 user manual.
- jhgb 5y agoWas this about the KU780 option?
- GeorgeTirebiter 5y agoAbsolutely correct, KU780 was the Writeable Control Store described here http://bitsavers.trailing-edge.com/pdf/dec/vax/handbook/VAX_Hardware_Handbook_Volume_1_1986.pdf http://bitsavers.trailing-edge.com/pdf/dec/vax/handbook/VAX_... Sophisticated customers hacking the instruction sets of their machines goes back pretty much to the beginning. The earliest I personally know of is Prof Jack Dennis hacking MIT's PDP-1 to support timesharing, sometime in 1961. Commercial machines like the Burroughs B1700 had a WCS that was designed so various compiled languages could be optimized - e.g. a FORTRAN instruction set, a COBOL instruction set, etc https://en.wikipedia.org/wiki/Burroughs_B1700 https://en.wikipedia.org/wiki/Burroughs_B1700 It was also in the IBM360s because they had to emulate the IBM1401 software (although I don't know if the capability was open to users to modify). Today of course you have the various optional features of the RISC-V ecosystem --- easy to load up on an FPGA. Perhaps we should remember that we are in the very very early days of Computers, and we should expect continued modification / experimentation.
- ncmncm 5y agoHow/when will we ever be able to confidently tell Gcc to generate these instructions, when we generally only know the code will be expected to run on some or other Aaargh64? It is the same problem as POPCNT on Amd64, and practically everything on RISC-V. Checking some status flag at program start is OK for choosing computation kernels that will run for microseconds or longer, but for things that take only a few cycles anyway, at best, checking first makes them take much longer. I imagine monkeypatching at startup, the way link relocations used to get patched in the days before we had ISAs that didn't support PIC. But that is miserable.
- floatboth 5y agoAll implementation selection should be done with ifuncs. Sadly lots of programs still do it with just function pointers.
- th3typh00n 5y agoifuncs is a non-standard compiler extension that only works on certain operating systems. Developers that cares about portability are obviously going to stay far away from such things.
- Unklejoe 5y agoI guess the fist step could be to handle it in the C library using some capability check and function pointers, then perhaps later on in the compiler if some mcpu flag or something is provided.
- ndesaulniers 5y agoGood questions. For computer support, generally you would pass a -mcpu= flag (or maybe -mattr=, but that might be a compiler internal flag, I forget). Obviously then that's not portable and has implications on the ABI. I didn't read the article but I suspect they might be in ARMv9.0, hopefully, otherwise "better luck next major revision." For monkey patching, the Linux kernel already does this aggressively since it generally has permission to read the relevant machine specific registers (MSRs). Doesn't help userspace, but userspace can do something similar with hwcaps and ifuncs.
- truth_seeker 5y agoHow efficient it would be from performance and security point of view ?
- Aissen 5y agoI saw this the other day, wanted to read it and failed; and again today. Luckily Google has it in cache: http://webcache.googleusercontent.com/search?q=cache%3Ahttps%3A%2F%2Fcommunity.arm.com%2Fdeveloper%2Fip-products%2Fprocessors%2Fb%2Fprocessors-ip-blog%2Fposts%2Farm-a-profile-architecture-developments-2021 http://webcache.googleusercontent.com/search?q=cache%3Ahttps...
- Aissen 5y agoI'm wondering if this isn't solving a problem only with a local optimum. How much better would be to have a standard way (i.e, not device-specific) to memzero (or memset) directly into the DRAM chips ? Or to use DMA for memcpy, while the CPU does other things ? Now of course, this could be a nightmare for cache coherency, but I've seen worse things done for performance.
- dvdkhlng 5y agoIn fact the Cell CPU [1] had a DMA facility accessible from the SPU cores by non-privileged software [2]. This worked cleanly, as all DMA operations were subject to normal virtual memory paging rules. But then the SPU did not have direct RAM access (only 256 kB of local S-RAM addressible from the CPU instructions), so DMA was something that followed naturally from the general design. Also not having any cache meant there were none of the usual cache coherency problems (though you may run into coherency problems during concurrent DMA to shared memory from multiple SPUs). [edit] note also that the SPUs did not usually do any multitasking / multi-threading, which also simplified handling of DMA. Otherwise task switches would have to capture and restore the whole DMA unit's state (and also potentially all 256 kB of local storage as these cannot be paged). [1] https://en.wikipedia.org/wiki/Cell_(microprocessor) https://en.wikipedia.org/wiki/Cell_(microprocessor) [2] https://arcb.csc.ncsu.edu/~mueller/cluster/ps3/SDK3.0/docs/accessibility/sdkpt/cbet_3compintrs.html https://arcb.csc.ncsu.edu/~mueller/cluster/ps3/SDK3.0/docs/a...
- Unklejoe 5y ago> all DMA operations were subject to normal virtual memory paging rules. That's the key right there. Many embedded SoC's I've worked with have DMA engines, but they are all behind the MMU and only work with physical addresses. It makes using them for something like "accelerated memcpy" kind of cumbersome and usually not even worth it unless it's moving HUGE chunks of memory (to overcome the page table walk that you have to do first).
- monocasa 5y agoThankfully we're starting to get IO-MMUs on larger systems with DMA controllers like that. Much easier to pass around.
- mackman 5y agoI remember implementing memcpy for a PS3 game. If you were doing a lot of copying (which we were for some streaming systems) it was hugely beneficial to add some explicit memory prefetching with a handful of compiler intrinsics. I think the PPC processor on that lacked out of order execution so you would stall a thread waiting for memory all too easily.
- mhh__ 5y ago23-stage in-order pipeline according to Wikipedia https://en.wikipedia.org/wiki/Cell_(microprocessor)#Power_Processor_Element_(PPE) https://en.wikipedia.org/wiki/Cell_(microprocessor)#Power_Pr...
- dvdkhlng 5y agoWell, the Cell CPU also had DMA engines that were fully integrated into the MMU memory-mapping, so you would have been able to asynchronously do a memcpy() while the CPU's execution resources were busy running computations in parallel.
- JoeAltmaier 5y agoI'm concerned this is another patch on a very difficult problem. There are something like 16 different combinations of source alignment, destination alignment, partial starting word, partial ending word memory-move operations. What's needed is an efficient move that does the right thing at runtime, which is to fetch the largest bus-limited chunks and align as it goes. This includes a pipeline to re-align from source to destination; partial-fill of the pipe line at the start and partial dump at the end; and page-sensitive fault and restart logic throughout. Multiple versions of memcpy is suspicious to start with: is the compiler expected to know the alignment statically at code generation time? It might be from arbitrary pointers. Alignment is best determined at runtime. Each pass through the same memcpy code may have different aligment and so on. Years ago I debugged the standard linux copy on a RISC machine. It has a dozen bugs related to this. I remember thinking at the time, this should all be resolved at runtime by microcode inside the processor. It's been years now, and we get this. Sigh. It's a step anyway.
- ruslan 5y agoA completely useless use of silicon. The fastest way of copying memory block is to offload it to DMA or some other dedicated hardware. Using CPU to copy blocks is just a stall. And please, do not call ARM a RISC!
- branson2 5y agoDMA is cache oblivious, so you will have to clear all CPU caches before every memcpy. Very bad idea.