7 ms·
I remember when Buildzoid of AHOC did the mobo breakdowns of the TR4 boards and thinking that while some of the super high end boards were probably good for thi
by leeter 7y ago
I remember when Buildzoid of AHOC did the mobo breakdowns of the TR4 boards and thinking that while some of the super high end boards were probably good for this sort of beating, the mid and low range might struggle hard with the 64/128 part if it ever came into being (3990X was just a rumor at that time). But it looks like they need more caps to handle the transient response time, and probably also some firmware fixes to slow ramp because I don't think all the SMD caps in the world are going to handle that sort of ramp. It's just not possible to get them close enough to the actual CPU without literally putting them under the IHS.
- franzb 7y agoHere's BuildZoid review of the motherboard I'm using (GIGABYTE TRX40 Aorus Xtreme): https://www.youtube.com/watch?v=HMUWzDSAS9c https://www.youtube.com/watch?v=HMUWzDSAS9c A very interesting watch if you're interested in electronics in general and in power delivery in particular (the whole YouTube channel is awesome to be honest).
- leeter 7y agoThat was one of the few that I thought could handle the 64/128 part. However look at the output filtering: It's roughly the same if not a cap or two larger as what you'd find on x299... which is a higher voltage and thus has lower amperage requirements. I have yet to find a back of board shot but unless it has a ton of SMD AL-poly caps back there that board would still struggle with the ramp described in the thread. Even then I'm not sure the socket resistance wouldn't cause enough V-droop to cause a crash anyway.
- close04 7y agosTRX4 can take one of three CPUs ranging from 24 to 64 cores. The TRX40 Aorus Xtreme that's mentioned in the thread should be the absolute top of the line even if it was launched before the 64 core monster was available. So I'd expect it to work just fine with the 3970X which is a 32 core part. But I wonder if this is a Gigabyte issue who have a history of playing around with the power delivery and using "fake phases" (to the point where they now have to advertise their boards as having 16 "real phases"). As far as I can tell many (most?) reviewers benchmarked the board with 24 core CPUs and most likely skipped on the power intensive tests.
- franzb 7y agoThis mobo has 16 real phases, according to BuildZoid's analysis.
- close04 7y agoI imagine it is. I meant they had such issues in the past (hence the "this time for real" approach) so it may be that they cut different corners in order to meet some other targets that are marketable.
- leeter 7y agoTo me this smells 100% like a transient issue, Gigabyte probably assumed that more phases would wipe the transient response time out. The issue is that if you're shutting down phases as Buildzoid mentions you'd have to do to get efficiency; you then lose that response time advantage until the phase is spun back up which is the longest part of the power system cycle. Normally output filtering is designed to make up that gap, but I think in the case of these chips that's not happening. I suspect it's a lot more complicated because I have a sneaky suspicion that adding more caps wouldn't solve the problem (completely).
- derefr 7y agoTangent: is there a reason that CPUs’ instruction sets aren’t designed with explicit “hint”-ops in them to let a compiler assert “I’m going to execute some instructions that’re going to draw a lot of power about 1000 cycles from now, so start ramping up for it now”? That’d basically eliminate what Intel calls “license-switching” costs. Do they just not believe in compiler authors to be able to emit these kinds of hints? Is there too much legacy code that would come without the hints for it to be worthwhile? Would it just not be worth it given how often OS context switches could drop the CPU directly from regular code in one process to AVX2 code in another process?
- gruez 7y agoProbably because the better approach would be to stall the processor (or at least stall on power intensive instructions) until the requisite power has arrived. 1000 instructions is probably too far out (in terms of instructions) for compilers to accurately add the requisite instructions. Also, it requires code change, unlike the stalling approach. The only disadvantage to the stalling approach would be if the workload is constantly switching between low power and high power, but that can be solved by making the switching interval longer/shorter in the firmware.
- londons_explore 7y agoStalling is the way to go. I have my doubts about this being a power issue, because localized power analysis is pretty advanced in IC design and they'd totally have tested for this. Assuming it is a power issue though, nearly all power is used in dynamic switching, so simply inserting stall cycles will resolve it. Power rails on silicon chips have quite a lot of capacitance, at least enough for many clock cycles, so you don't even need to make every cycle in every bit of hardware a possible stall cycle - it would be enough to simply gate instruction issue I expect and let functional units drain.
- derefr 7y agoWhy is it “better” if the stall results in lower performance-per-watt? That’s what people buy these high-end CPU SKUs for, after all. They aren’t just concerned with CapEx (the cost of the CPU) but also OpEx (the aggregate cost of electricity for the CPUs, PSUs and cooling in their studio/data center.) You could also “fix” the problem by just allowing the customer to lock the CPU into AVX2 mode and thereby never switching out at all—but of course that’d be dumb; they’d be drawing more power, and also not going as fast as they could, when executing non-AVX2 code. Another approach I haven’t seen from either Intel or AMD yet, is to copy ARM’s big.LITTLE architecture: to set up separate “AVX2 cores” that can execute AVX instructions and most basic ALU ops but not e.g. branches, put those on the other side of the die with in-wafer thermal insulation between them and the regular cores; and then throw workloads between regular and AVX2 cores in a way where the AVX2 cores heating up doesn’t mean that the non-AVX2 cores are heating up, and the CPU can go back to full Turbo Boost as soon as the workload is thrown back to the regular cores, because the respective regular core is actually quite cool. (The performance effects of this are possible to loosely estimate, I think, by writing code that synchronously sets up a GPGPU pass on some data, executes it, retrieves the result, and then returns to executing CPU instructions for a while, in a loop; and then executing this code on an Intel CPU using its on-die IGPU as the GPU. The CPU and IGPU form somewhat-separate thermal domains already, though there’s no explicit insulation.)