13 ms·
Surprising new feature in AMD Ryzen 3000
- gigatexal 6y agoI wonder if it’s worth the effort to add functionality to take advantage of this in GCC, clang for just these CPUs?
- pja 6y agoI think the point is that most ordinary C-code should already conform to this access pattern. The side comment about aliasing is a great big hint in that direction. (This article is grist to the mill of my "every CPU eventually evolves to become an interpreter for its dominant programming language" thesis.)
- convFixb 6y agoI wonder if this causes a performance penalty on position independent code, which seems to be used a lot on non-Windows platforms []. [] It always struck me as one of those propeller-head features that GCC & co. love but which the MS and Intel compilers avoid 'just to be on the safe side' :]
- zbjornson 6y agoHe mentions this doesn't work with RIP-relative addressing, but I don't see how it would cause a penalty otherwise.
- qppo 6y agore your thesis, can't wait for a WASM ISA on a commercial CPU and canvas instructions for its graphics coprocessor
- PaulDavisThe1st 6y agoyou'll need to wait for WASM to become "the dominant language" on a commercial CPU. definitely not next year.
- bryanlarsen 6y agoYes, AMD probably added the optimization because it was a common pattern emitted by compilers.
- wyldfire 6y agoAFAIK there's existing microarchitectural optimizations in gcc and clang already, so it's reasonable to assume that AMD would add just such a feature. But most people don't enable these features when building because they often make the resulting executable/library unusable on general CPUs. However, businesses/individuals with more specialized needs and better control over their targets often do capitalize on these.
- Hello71 6y agothis is the purpose of the gcc -mtune flag, which does have an actual purpose beyond being used in the utterly redundant -march=native -mtune=native.
- sroussey 6y agoBack in the days of compiling MySQL myself, it was worth doing. But at some point They hardware cluster was not identical, and so I scripted builds and each machine built its own.
- segfaultbuserr 6y ago-march=arch -mtune=arch has existed for decades, and the performance improvement is measurable. But mainstream distributions often cannot take advantage, since they need to support all CPUs, and optimization for one is often a deoptimization for another. It's also a reason why Gentoo exists. Interestingly, Intel's Clear Linux - a optimization-oriented distribution - uses -march=westmere -mtune=haswell https://docs.01.org/clearlinux/latest/guides/clear/performance.html https://docs.01.org/clearlinux/latest/guides/clear/performan...
- marmaduke 6y agoWestmere is pretty old but significantly is the first generation to support virtual machines with little overhead. If clear Linux allowed something newer then it would rule itself out for a lot of existing machines (such as the 20+ I have running HPC jobs with surprising reliability)
- segfaultbuserr 6y agoYeah, it's a reasonable tradeoff - be compatible with 1st Gen (basically all post-2010 modern x86_64 Intel processors), but try fine-tuning it a little bit for Haswell.
- gpderetta 6y agoprobably yes. It means that spills are cheaper.
- Exorus18 6y agoWaiting for next side-channel attack..
- flipgimble 6y agoCould you explain the connection between this feature and a possible side channel attack?
- kmeisthax 6y agoA thread in the same process could observe the timing change and infer certain things about the memory address. This is mitigated by the fact that intra-thread and host-guest communication is already a side channel nightmare and everyone now only considers process boundaries to be defensible. It's sort of like worrying about someone being able to cut keys from a picture you posted on Facebook when all your house locks can be raked open with a $5 tool from Amazon.
- gliptic 6y agoAgner Fog mentions that the CPU assumes different registers have different values and the calculation has to be rerun if it turns out a write with one register invalidated a cached read with another register. That means you can measure this penalty and infer that the registers had the same or overlapping addresses.
- gpderetta 6y agoThis is no different than the existing memory aliasing predictor that has existed for a long time already though, so it doesn't add any new speculation opportunity as far as I can tell.
- ljhsiung 6y agoIt might also depend on how AMD clears this temporal register across contexts, and the method it's updated. The first I could think of (in the 5 minutes I've read this) is that it could be potentially ASLR breaking. As Agner says, it's highly useful for stack operations, but that also can help an adversary derive where the stack is. Second, I could see an attack where the CPU predicts a sequence will use this register (since the address in the register could be speculative) and prefetch (since it kind of looks like it's prefetching) but it ends up being wrong and not killing/clearing the register, or not preventing its usage speculatively. But it all depends heavily on lots of things. I feel like it's pretty similar to spec v4 in that exploitation would depend on the memory disambiguator, but we'll see.
- eganist 6y agoI wonder how many little wins like this contribute to AMD's immense efficiency over Intel's current chips?
- rurban 6y agoLooks like the single greatest win to me. This explains now a lot.
- rob74 6y agoI think the area where Intel is currently lagging is not so much the internal architecture of the chips, but the fabrication process, where they have fallen behind AMD (who is using 3rd party fabs): https://www.forbes.com/sites/linleygwennap/2020/07/28/intel-fab-troubles-problem-for-pc-servers/#6d3eb334408f https://www.forbes.com/sites/linleygwennap/2020/07/28/intel-...
- qppo 6y agoBigger (or rather, smaller) than the fab process was the clever use of chiplets to improve yields.
- spockz 6y agoI love this kind of technical analysis. Is there any repository containing more of these kind of analysis?
- formerly_proven 6y agohttps://www.agner.org/optimize/#manuals https://www.agner.org/optimize/#manuals
- radres 6y agoIt's interesting that the author uses a forum as a personal blog
- brudgers 6y agoSeems like a reasonable approach to allow and manage comments...and build and manage a community. It's a nice hack.
- segfaultbuserr 6y agoIt works the best if you already run a forum and you want to write some related articles. Just open a subforum, and you get a blog with community traffic.
- magicalhippo 6y agoKISS at its finest. You can inline images and code easily. Easy to self-host, so you can own all the data, and you get solid support for comments so you can avoid data-peddling companies like Disqus.
- jayflux 6y ago> KISS at its finest. It’s certainly an interesting way to do it. I’m not sure id set up a forum just to write notes though, there’s far simpler methods. I think in this case theres already a community on this forum and it’s easier to just write there than anywhere else.
- toyg 6y ago> Easy to self-host I wouldn't go as far. Forum software is notoriously annoying to configure securely and to maintain. > you get solid support for comments ... at the price of people having to create Yet Another User/Password Set, and likely bother you at some point for user-admin tasks. You definitely keep control of the data though, and some people really like forums.
- penagwin 6y ago
- dis-sys 6y agoThe author wrote in another thread that "If anybody has access to the new Chinese Zhaoxin processor, I would very much like to test it." Will be very interesting to see how much actual changes Zhaoxin made to the VIA cores. I'd expect it to be minimum.
- deleted 6y ago[deleted]
- userbinator 6y agoSurprising... and a little scary. This is not something I would've expected to be done in the current world of multiple cores. I wonder if things like volatile and lock-free algorithms would behave any differently or even break.
- chrisseaton 6y ago> I wonder if things like volatile and lock-free algorithms would behave any differently or even break. I don't understand how do you think it could change or break a volatile or lock-free algorithm? It's a transparent optimisation - it doesn't change observable behaviour. Just like how register renaming, top-of-stack-caching, and so on and so on don't change observable behaviour. The only difference it would make is timing.
- userbinator 6y agoSo there is logic to determine if a different thread has modified the value between those instructions?
- aclindsa 6y agoNo - there doesn't have to be. As long as there are no barriers or other instructions which restrict memory ordering relative to particular instructions in the dynamic instruction stream, it doesn't matter what the actual interleaving is between different cores as long as there is a given execution order which is valid for the observable side effects.
- cma 6y agoX64 has stronger guarantees than that without barriers doesn't it?
- PixelOfDeath 6y agox86 guarants that memory writes to different locations are seen in order by other cores. If you write memory first to address A and then to B. Another core never can see the B change without also seeing the A change at any moment in time. But in this case there is only a single memory location. So there is no ordering that could be violated in the first place.
- josmala 6y agoAs a Zen2 owner I'm very disappointed in VPGATHERDD througput, that's so 2013. On the other hand I like the loop and call instruction performance a lot.
- ajross 6y agoThat gather needs to issue 8 independent loads. It's never going to be fast, and I think there's a strong argument that you don't even want to spend the transistors on all the extra load/store units required. The goal of scatter/gather instructions is that they should be demonstrably faster than assembling the values in scalar code, and beyond that... meh. If you're doing random access to memory like that, you're probably out of the realm of what is appropriate in vector code and should be looking at other hardware (c.f. a GPU's texture units) to manage your memory access.
- brandmeyer 6y agoIts theoretically possible to run gather as fast as one cache line per cycle instead of one SIMD lane per cycle. I don't think anyone has thrown that much permute hardware at the problem, though. Its only profitable if you believe that scatter and gather do have cache locality even when they don't have regularity.
- PixelOfDeath 6y agoIsn't AVX512 basically cacheline-instructions?
- ajross 6y agoThat's the way normal SIMD loads work, yeah. But the scatter/gather instructions do random access memory operations. You have one SIMD register with a 8 (or whatever the width is) indexes to be applied to a base address in a scalar register, and the hardware then goes and does 8 separate memory operations on your behalf, packing the results into a SIMD register at the end. That has to hit the cache 8 times in the general case. It's extremely expensive as a single instruction, though faster than running scalar code to do the same thing.
- Waterluvian 6y agoI sense this question is pretty elementary, but maybe someone can point me in the right direction for reading: "When the CPU recognizes that the address [rsi] is the same in all three instructions..." Is there another abstraction layer like some CPU code that runs that would do the "recognition" or is this "recognition" happening as a result of logic gates connected in a certain static way? To put more broadly: I'm really interested in understanding where the rubber meets the road. What "code" or "language" is being run directly on the hardware logic encoded as connections of transistors?
- ch_123 6y agoOn most x86 chips since the Pentium Pro (and the K6 for AMD, IIRC) the machine code instructions which are visible to the programmer are turned into "micro operations" by the CPU's instruction decoder. See here for a introduction: https://en.wikipedia.org/wiki/Micro-operation https://en.wikipedia.org/wiki/Micro-operation
- Waterluvian 6y agoThanks for the link! So on older chips, like maybe an 8080, I could expect to see literal hardware implementations of instructions?
- klelatti 6y agoKen Shirriff has an excellent series of blog posts on early Intel chips (running from 4004 to 8086 I think). Don't believe he's done one on the 8080 but the post on the 8008 is great [1] and I'd expect that the 8080 (which followed it and was designed by the same team) is very similar. In short no microcode but there is a PLA (Programmable Logic Array) which helps to decode the instructions. [1] http://www.righto.com/2016/12/die-photos-and-analysis-of_24.html http://www.righto.com/2016/12/die-photos-and-analysis-of_24....
- kens 6y agoThanks for the nice comment! I haven't looked into the 8080 because other people have examined it in detail. See https://news.ycombinator.com/item?id=24101956 https://news.ycombinator.com/item?id=24101956 for an exact Verilog representation of the 8080.
- whizzter 6y agoInteresting, L1 caches are fast and even if compilers do register allocation they kinda rely on it being not-too-shitty so when spilling (and many compilers for higher level languages doesn't always invest too much time in reg-alloc since they might need to de-opt soon). I'm curious if this change is an effect of more transistors (more space for a bigger register file) or if they're taking advantage with the microcode translation of the fact that most code doesn't use the SIMD vector registers and re-use unused parts of the register file for these memory aliases.
- wtallis 6y agoI highly doubt the chips are using vacant portions of the SIMD register file to hold scalar values, especially memory addresses. The SIMD registers probably have no easy connection to ports where load/store units ingest addresses.
- Tuna-Fish 6y agoI'd bet dollars to peanuts that the temporaries are stored in the gpr prf. It's big, and it's sized so it's not usually a constraint, so they can just push memory operands in there when there is free space and when it would be useful. And most importantly, unlike the vector registers (or any other structure already in the CPU), it can directly feed the ALUs without any new datapaths.
- Sharlin 6y agogpr prf? General purpose register something register file? Private?
- otherjason 6y agoProbably physical register file as referenced by https://en.wikipedia.org/wiki/Register_file https://en.wikipedia.org/wiki/Register_file. PRF in this context refers to the actual memory that contains all of the register data, which can be much larger than the number of architectural registers to provide for register renaming, which can help performance (basically, the mapping from architectural register names to actual physical registers continually changes so you relieve the bottleneck of waiting for a particular physical register to become available).
- whereistimbo 6y agoAnother great thread from the same author: https://www.agner.org/forum/viewtopic.php?f=1&t=6 https://www.agner.org/forum/viewtopic.php?f=1&t=6
- cwt137 6y agoThis hidden copy feature, do you think someone can exploit it in a similar way as recent exploits like meltdown and spectre?
- andy_ppp 6y agoThe attacks for meltdown and spectre are completely different to this, I’m fairly certain that explicit features that could break memory protection would be tested. The point of the other attacks is to trick branch prediction to circumvent these protections. It is of course possible this could be used in a way I haven’t foreseen...
- abainbridge 6y agoTrying to understand this. Using latencies from Zen 1 instruction table (see https://www.agner.org/optimize/instruction_tables.pdf https://www.agner.org/optimize/instruction_tables.pdf): mov dword [rsi], eax ; MOV m,r latency is 4 add dword [rsi], 5 ; ADD m,i latency is 6 mov ebx, dword [rsi] ; MOV r,m latency is 4 Total = 14 Each instruction depends on the result of the previous, so we need to sum all the latency figures to get the total cycle count. Is this right? How does Agner make it add up to 15? Then for Zen 2: mov dword [rsi], eax ; MOV m,r latency is 0 (rather than 4, ; because it is mirrored) add dword [rsi], 5 ; ADD m,i cannot find an entry for this. ; Looks like there's a typo in the doc. ; I guess the latency is 1. mov ebx, dword [rsi] ; MOV r,m latency is 0 Total = 1 Again, how does Agner make it add up to 2? And for Intel Skylake: mov dword [rsi], eax ; MOV m,r latency is 2 add dword [rsi], 5 ; ADD m,i - latency is 5 mov ebx, dword [rsi] ; MOV r,m latency is 2 Total = 9
- CalChris 6y agoI think the Zen 2 reasons about the operations and operands and converts them into something like these micro-ops and then schedules them for a functional unit and the renamer: functional unit renamer add tmp_reg, eax, 5 st [rsi], tmp_reg mov ebx, temp_reg Consequently, I don't think you can use the latency tables directly to get "2". I think Agner probably got those 15 and 2 aggregate numbers from measuring them directly rather than calculating them from the latency tables. BTW, I'm hand waving on the micro-ops and FUs. There's probably some address generation, ... going on that I'm leaving out. I don't even think the renamer requires a translated micro-op. You could replace tmp_reg with renamed(ebx). Now the downside. What happens at an exception? You have to back all of this optimization out.
- vardump 6y ago> Now the downside. What happens at an exception? You have to back all of this optimization out. Just what always happens at an exception. You just replay everything up until the exception occurred.
- gpderetta 6y agoThere have been rumors that Zen could do memory renaming [1], this pretty much confirms it. [1] Basically the same as register renaming, but instead of using the register file to rename architectural registers, it can rename memory instead.
- tus88 6y agoWonder what side channel attacks this will result in.
- Razengan 6y agoA problem for every solution.
- saagarjha 6y agoIf the memory renaming cache isn't cleared, I can foresee Spectre-like attacks against that micro-architectural buffer. However, for a new feature like this hopefully they've figured out how to not leave that open…
- gpderetta 6y agoThe memory renaming cache is almost surely just the register file.
- saagarjha 6y agoFigured I might as well mention it in case AMD was doing something totally wild, but I would agree that using the register file sounds likely :) I wonder how they manage the pressure on it between renaming and memory mirroring…
- BeeOnRope 6y agoIn principle if you are using the GP PRF it is possible to implement it so that there is little additional pressure: the store instruction already has it's data input a register which has been renamed: now you just need to organize it so that the subsequent load that targets the same location is a no-go: simply alias the arch reg targeted by the load onto the existing register used by the store (much like mov-elimination) So except for a small window in the middle you have the same pressure on the PRF. I don't know if that is how it is actually implemented, of course!
- Paul-ish 6y agoIt looks like there is a significant miss penalty for aliasing. Does anyone know if Rust's ownership rules would help avoid these penalties.
- thelazydogsback 6y agoOk, so most my asm coding and knowledge of exactly what the CPU was doing ended sometime between the Z-80/68K/8086 timeframes. Are there any good books/resources on all the modern trickery that CPUs now utilize?
- Tuna-Fish 6y agoThe best resource for details of any specific x86 cpu is Agner Fog's 3rd manual at: https://www.agner.org/optimize/#manuals https://www.agner.org/optimize/#manuals
- ritter2a 6y agoWell, there is the Hennessy&Patterson book, "Computer Architecture - A Quantitative Approach". The newer editions include modern features and follow industry's developments. This will of course not be able to tell you all the black magic that is happening inside AMD's and Intel's newest designs. The software optimization manuals for these processors do include some more interesting insights, but especially the Intel ones are not the most entertaining read...