12 ms·
AMD Disables Zen 4's Loop Buffer
- jb1991 2y ago[flagged]
- olejorgenb 2y agoThat was not my takeaway from (admittingly) skimming the article. (that it's "huge")
- Shalah 2y agoWould people please stop hiding posts they disagree with. I'm a grown-up and can make my own mind up as to the veracity of any stated opinions.
- gjs278 2y ago[dead]
- School-Cotton 2y agoHN has a “showdead” option in the user preferences that will make it so you can see everything, no matter how downvoted.
- Shalah 2y agogjs278: “yeah and they make sure to put it in super light color so you have to highlight as well” And unregistered readers can't see these posts :o
- kergonath 2y agoThe problem is not agreeing or not. The problem is posting uninformed opinions without reading the story, which explains in detail that it is both true and not that important (i.e., not "huge").
- throwuxiytayq 2y agoNeither huge, nor if.
- moffkalast 2y agoSo just true?
- syntaxing 2y agoInteresting read, one thing I don’t understand is how much space does loop buffer take on the die? I’m curious with it removed, on future chips could you use the space for something more useful like a bigger L2 cache?
- progbits 2y agoIt says 144 micro-op entries per core. Not sure how many bytes that is, but L2 caches these days are around 1MB per core, so assuming the loop buffer die space is mostly storage (sounds like it) then it wouldn't make a notable difference.
- Remnant44 2y agoMy understanding is that it's a pretty small optimization on the front end. It doesn't have a lot of entries to begin with (144) so the amount of space saved is probably negligible. Theoretically, the loop buffer would let you save power or improve performance in a tight loop. In practice, it doesn't seem to do either, and AMD removed it completely for Zen 5.
- akira2501 2y agoI think most modern chips are routing constrained and not floorspace constrained. You can build tons of features but getting them all power and normalized signals is an absolute chore.
- deleted 2y ago[deleted]
- atq2119 2y agoJudging from the diagrams, the loop buffer is using the same storage as the micro-op queue that's there anyway. If that is accurate (and it does seem plausible), then the area cost is just some additional control logic. I suspect the most expensive part is detecting a loop in the first place, but that's probably quite small compared to the size of the queue.
- eqvinox 2y ago> Strangely, the game sees a 5% performance loss with the loop buffer disabled when pinned to the non-VCache die. I have no explanation for this, […] With more detailed power measurements, it could be possible to determine if this is thermal/power budget related? It does sound like the feature was intended to conserve power…
- eek2121 2y agoHe didn’t provide enough detail here. The second CCD on a Ryzen chip is not as well binned as the first one even on. non-X3D chips. Also, EVERY chip is different. Most of the cores on CCD0 of my non-X3D chip hit 5.6-5.75ghz. CCD 1 has cores topping out at 5.4-5.5ghz. V-Cache chips for Zen 4 have a huge clock penalty, however the Cache more than makes up for it. Did he test CCD1 on the same chip with both the feature disabled and enabled? Did he attempt to isolate other changes like security fixes as well? He admitted “no” in his article. The only proper way to test would be to find a way to disable the feature on a bios that has it enabled and test both scenarios across the same chip, and even then the result may still not be accurate due to other possible branch conditions. A full performance profile could bring accuracy, but I suspect only an AMD engineer could do that…
- clamchowder 2y agoYes, I tested on CCD1 (the non-vcache CCD) on both BIOS versions.
- ryao 2y agoHe mentioned that it was disabled somewhere between the two UEFI versions he tested. Presumably there are other changes included, so his measurements are not strict A/B testing.
- Pannoniae 2y agoFrom another article: "Both the fetch+decode and op cache pipelines can be active at the same time, and both feed into the in-order micro-op queue. Zen 4 could use its micro-op queue as a loop buffer, but Zen 5 does not. I asked why the loop buffer was gone in Zen 5 in side conversations. They quickly pointed out that the loop buffer wasn’t deleted. Rather, Zen 5’s frontend was a new design and the loop buffer never got added back. As to why, they said the loop buffer was primarily a power optimization. It could help IPC in some cases, but the primary goal was to let Zen 4 shut off much of the frontend in small loops. Adding any feature has an engineering cost, which has to be balanced against potential benefits. Just as with having dual decode clusters service a single thread, whether the loop buffer was worth engineer time was apparently “no”."
- deleted 2y ago[deleted]
- londons_explore 2y agoThe article seems to suggest that the loop buffer provides no performance benefit and no power benefit. If so, it might be a classic case of "Team of engineers spent months working on new shiny feature which turned out to not actually have any benefit, but was shipped anyway, possibly so someone could save face". I see this in software teams when someone suggests it's time to rewrite the codebase to get rid of legacy bloat and increase performance. Yet, when the project is done, there are more lines of code and performance is worse. In both cases, the project shouldn't have shipped.
- adgjlsfhk1 2y ago> but was shipped anyway, possibly so someone could save face no. once the core has it and you realize it doesn't help much, it absolutely is a risk to remove it.
- glzone1 2y agoNo kidding. I was adjacent to a tape out w some last minute tweaks - ugh. The problem is the current cycle time is very slow and costly and u spend as much time validating things as you do designing. It’s not programming.
- hajile 2y agoIf you work on a critical piece of software (especially one you can't update later), you absolutely can spend way more time validating than you do writing code. The ease of pushing updates encourages lazy coding.
- chefandy 2y ago> The ease of pushing updates encourages lazy coding. Certainly in some cases, but in others, it just shifts the economics: Obviously, fault tolerance can be laborious and time consuming, and that time and labor is taken from something else. When the natures of your dev and distribution pipelines render faults less disruptive, and you have a good foundational codebase and code review process that pay attention to security and core stability, quickly creating 3 working features can be much, much more valuable than making sure 1 working feature will never ever generate a support ticket.
- londons_explore 2y agoIn the "power" section, it seems the analysis doesn't divide by the number of instructions executed per second. Energy used per instruction is almost certainly the metric that should be considered to see the benefits of this loop buffer, not energy used per second (power, watts).
- eek2121 2y agoEvery instruction takes a different amount of clock cycles (and this varies between architectures or iterations of an architecture such as Zen 4-Zen 5), so that is not feasible unless running the workload produced the exact same instructions per cycle, which is impossible due to multi threading/tasking. Even order and the contents of RAM matters since both can change everything. While you can somewhat isolate for this by doing hundreds of runs for both on and off, that takes tons of time and still won’t be 100% accurate. Even disabling the feature can cause the code to use a different branch which may shift everything around. I am not specifically familiar with this issue, but I have seen cases where disabling a feature shifted the load from integer units to the FPU or the GPU as an example, or added 2 additional instructions while taking away 5.
- rasz 2y agoAnecdotally one of very few differences between 1979 68000 and 1982 68010 was addition of "loop mode", a 6 byte Loop Buffer :)
- crest 2y agoMuch more importantly they fixed the MMU support. The original 68000 lost some state required to recover from a page fault the workaround was ugly and expensive: run two CPUs "time shifted" by one cycle and inject a recoverable interrupt on the second CPU. Apparently it was still cheaper than the alternatives at the time if you wanted a CPU with MMU, a 32 bit ISA and a 24 bit address bus. Must have been a wild time.
- phire 2y ago> run two CPUs "time shifted" by one cycle and inject a recoverable interrupt on the second CPU. That's not quite how it was implemented. Instead, the second 68000 was halted and disconnected from the bus until the first 68000 (the executor) trigged a fault. Then the first 68000 would be held in halt, disconnected from the bus and the second 68000 (the fixer) would take over the bus to run the fault handler code. After the fault had been handled, the first 68000 could be released from halt and it would resume execution of the instruction, with all state intact. As for the cost of a second 68000, extra logic and larger PCBs? Well, the of the Motorola 68451 MMU (or equivalent) absolutely dwarfed the cost of everything else, so adding a second CPU really wasn't a big deal. Technically it didn't need to be another 68000, any CPU would do. But it's simpler to use a single ISA. For more details, see Motorola's application note here: http://marc.retronik.fr/motorola/68K/68000/Application%20Notes/DC001_Virtual_Memory_Using_The_MC68000_and_the_MC68451_MMU_%5BMotorola_1982_9p%5D.pdf http://marc.retronik.fr/motorola/68K/68000/Application%20Not...
- phire 2y agoFurther thoughts: While this executor + fixer setup does work for most usecases, it's still impossible to recover the state. The relevant state is simply held in the halted 68000. Which means, the only thing you can do is handle the fault and resume. If you need to page something in from disk, userspace is entirely blocked until the IO request completes. You can't go and run another process that isn't waiting for IO. I suspect it also makes it impossible to correctly implement POSIX segfault signal handlers. If you try to run it on the executor, then the state is cleared and it's not valid to return from the signal handler anymore. If you run the handler on the fixer instead, then you are running in a context without pagefaults, which would be disastrous if the segfault handler access code or data that has been paged out. And the now segfault handler wouldn't have access to any of the executor's CPUs state. ------ So there is merit to the idea of running two 68000s in lockstep. That would theoretically allow you to recover the full state. But there is a problem: It's not enough to run the second 68000 one cycle behind. You need to run it one instruction behind, putting all memory read data and wait-states into a FIFO for the second 68000 to consume. And 68000 instructions have variable execution time, so I guess the delay needs to be the length of the longest possible instruction (which is something like 60 cycles). But what about pipelining? That's the whole reason why can't recover the state in the first place. I'm not sure, but it might be necessary to run an entire 4 instructions behind, which would mean something like 240 cycles buffered in that FIFO. This also means your fault handler is now running way too soon. You will need to emulate 240 cycles worth of instructions in software until you find the one which triggered the page fault. I think such an approach is possible, but it really doesn't seem sane. -------- I might need to do a deeper dive into this later, but I suspect all these early dual 68000 Unix workstations simply dealt with the issues of the executor/fixer setup and didn't implement proper segfault signal handlers. It's reasonably rare for programs to do anything in a segfault handler other than print a nice crash message. Any unix program that did fancy things in segfault handlers weren't portable, as many unix systems didn't have paging at all. It was enough to have a memory mapper with a few segments (base, size, and physical offset).
- eek2121 2y agoIt sounds to me like it was too small to make any real difference except in very specific scenarios and a larger one would have been too expensive to implement compared to the benefit. That being said, some workloads will see a small regression, however AMD has made some small performance improvements since launch. They should have just made it a BIOS option for Zen 4. The fact they do not appear to have done so does indicate the possibility of a bug or security issue.
- crest 2y agoThem *quietly* disabling a feature that few users will notice yet complicates the frontend suggests they pulled this chicken bit because they wanted to avoid or delay disclosing a hardware bug to the general public, but already push the mitigation. Fucking vendors! Will they ever learn? sigh
- whaleofatw2022 2y agoDevils advocate... if this is being actively exploited or is easily exploitable, the delay in announcement can prevent other actions.
- dannyw 2y agoEvery modern CPU has dozens of hardware bugs that aren’t disclosed and quietly patched away or not mentioned.
- BartjeD 2y agoQuitely disabling it is also a big risk. Because you're signalling that in all probablity you were aware of the severity of the issue; Enough so that you took steps to patch it. If you don't disclose the vulnerability then affected parties cannot start taking countermeasures, except out of sheer paranoia. Disclosing a vulnerability is a way shift liability onto the end user. You didn't update? Then don't complain. Only rarely do disclosures lead to product liability. I don't remember this (liability) happening with Meltdown and Spectre either. So wouldn't assume this is AMD being secretive.
- shantara 2y agoThis is a wild guess, but could this feature be disabled in an attempt at preventing some publicly undisclosed hardware vulnerability?
- throw_away_x1y2 2y agoBingo. I can't say more. :(
- pdimitar 2y agoHave we learned nothing from Spectre and Meltdown?... :(
- aseipp 2y agoThis might come as a shock, but I can assure you that the designing high end microprocessors have probably forgotten more about these topics than most of the people here have ever known.
- pdimitar 2y agoHuh?
- gpderetta 2y agoComplex systems are complex?
- pdimitar 2y agoSadly you're right. And obviously we're not about to give up on high IPC. I get it and I'm not judging -- it's just a bit saddening.
- StressedDev 2y agoA lot has been learned. Unfortunately, people still make mistakes and hardware will continue to have security vulnerabilities.
- deleted 2y ago[deleted]
- ksec 2y agoWondering if Loop Buffer is still there with Zen 5? ( Idly waiting for x86 to try and compete with ARM on efficiency. Unfortunately I dont see Zen 6 or Panther Lake getting close. )
- monocasa 2y agoIt is not.
- CalChris 2y agoIf it saved power wouldn’t that lead to less thermal throttling and thus improved performance? That power had to matter in the first place or it wouldn’t have been worth it in the first place.
- kllrnohj 2y agoNot necessarily. Let's say this optimization can save 0.1w in certain situations. If one of those situations is common when the chip is idle just keeping wifi alive, well hey that's 0.1w in a ~1w total draw scenario, that's 10% that's huge! But when the CPU is pulling 100w under load? Well now we're talking an amount so small it's irrelevant. Maybe with a well calibrated scope you could figure out if it was on or not. Since this is in the micro-op queue in the front end, it's going to be more about that very low total power draw side of things where this comes into play. So this would have been something they were doing to see if it helped for the laptop skus, not for the desktop ones.
- Out_of_Characte 2y agoYou're probaly right on the mark with this. Though even desktops and servers can benefit from lower idle power draw. So there is a chance that it might have been moved to a different c-state.
- mleonhard 2y agoIt looks like they disabled a feature flag. I didn't expect to see such things in CPUs.
- astrange 2y agoThey have lots of them (called "chicken bits"). Some of them have BIOS flags, some don't. It's very very expensive to fix a bug in a CPU, so it's easier to expose control flags or microcode so you can patch it out.
- rasz 2y agoThere were CPUs with whole plethora of optional optimizations. For example Cyrix packed their CPUs with goodies, but had no money to test so made it all optional. https://www.ardent-tool.com/CPU/Cyrix_Cx486.html#soft https://www.ardent-tool.com/CPU/Cyrix_Cx486.html#soft https://www.vogons.org/viewtopic.php?t=45756 https://www.vogons.org/viewtopic.php?t=45756 Register settings for various CPUs https://www.vogons.org/viewtopic.php?t=30607 https://www.vogons.org/viewtopic.php?t=30607 Cyrix 5x86 Register Enhancements Revealed L1, Branch Target Buffer, LSSER (load/store reordering), Loop Buffer, Memory Type Range Registers (Write Combining, Cacheability), all controlled using client side software. Cyrix 5x86 testing of Loop Buffer showed 0.2% average boost and 2.7% maximum observable speed boost.
- fulafel 2y agoInteresting that in the Cortex-A15 this is a "key design feature". Are there any numbers about its effect other chips? I guess this could also be used as an optimization target at least on devices that are more long lived designs (eg consoles).
- nwallin 2y agoI'm curious about this too. I would expect any RISC architecture to gain relatively little from a loop buffer. The point of RISC is that instruction fetch/decode is substantially easier, if not trivial.
- Loic 2y agoFor me the most interesting paragraph in the article is: > Perhaps the best way of looking at Zen 4's loop buffer is that it signals the company has engineering bandwidth to go try things. Maybe it didn't go anywhere this time. But letting engineers experiment with a low risk, low impact feature is a great way to build confidence. I look forward to seeing more of that confidence in the future.
- jjjleooeoeoeo 2y ago[flagged]
- Neywiny 2y agoI have a 7950x3d. It's my upgrade from.... Skylake's 6700k. I guess I'm subconsciously drawn to chips with hardware loop buffers disabled by software.
- saghm 2y agoIf you're going to buy a new machine at some point, definitely let us know in advance so we can avoid it!