8 ms·
Unexpected benefit with Ryzen – reducing power for build server
- ddorian43 8y agoAny server-hosting with Ryzen + ecc-ram ?
- Nux 8y agohttps://www.hetzner.com/dedicated-rootserver/matrix-ax https://www.hetzner.com/dedicated-rootserver/matrix-ax
- ddorian43 8y agoI know but it's not ecc-ram.
- Nux 8y ago"128 GB DDR4 ECC RAM" .. sounds pretty ECC to me.
- ddorian43 8y agoI was talkin about Ryzen 8core more actually. Those have low clock speed. And are called EPYC.
- Nux 8y agoGot it. For what it's worth I think EPYC is based on the same "Zen" arch as Ryzen.
- AstralStorm 8y agoWith a bit enhanced memory controller, Zen refresh (2xxx) should do a bit better still.
- BeeOnRope 8y agoPacket.net offers it: https://www.packet.net/bare-metal/servers/c2-medium-epyc/ https://www.packet.net/bare-metal/servers/c2-medium-epyc/
- ddorian43 8y agoPlease see other comments. It's EPYC and not RYZEN. 2.0 Ghz base clock speed is not nice for single-thread.
- dman 8y agoFor compilation it is a beast. Built a dual 7551 Epyc workstation for myself recently, it builds llvm in ~160 seconds from scratch. https://openbenchmarking.org/result/1809030-AR-DUAL7551867 https://openbenchmarking.org/result/1809030-AR-DUAL7551867
- morsma 8y agoGod damn that's a beast!
- paxswill 8y agoEpyc is the name for the server chips. Both Ryzen and Epyc are based on the same microarchitecture.
- ddorian43 8y agoBoth based on Zen. Still I wanted the high-clock-core and not the hundred-slow-cores.
- AstralStorm 8y agoIn that case, Threadripper is the core that is inbetween.
- afandian 8y agoDo you have experience with packet.net ? There's virtually no discussion on HN, which I find surprising.
- Jonnax 8y ago180w to 85w is pretty impressive. I didn't know that compilation was memory speed limited. Are there any good benchmarks on it? Anyone have any examples of getting faster memory boosting build speed? Over the last few years I'd settled into thinking that high speed ram barely did anything. I guess I was wrong!
- lykr0n 8y agoryzen loves memory speed, more so than Intel. phoronix has benchmarks that you're looking for. https://www.phoronix.com/scan.php?page=article&item=ryzen-ddr4-bios&num=1 https://www.phoronix.com/scan.php?page=article&item=ryzen-dd...
- gizmo686 8y agoMore specifically, the interprocessor interconnect (infinity fabric) system ryzen uses is tied to the RAM clock. Ryzen clumps there processosor in groups of 4, and uses infinity fabric as an interconnect between those; so I am not sure you will an effect larger then Intel on a quadcore ryzen. https://www.techpowerup.com/231585/amd-ryzen-infinity-fabric-ticks-at-memory-speed https://www.techpowerup.com/231585/amd-ryzen-infinity-fabric...
- _wmd 8y agoAnything involving larger-than-cache sized graph walks will usually be memory limited, be it compilation or iterating an XML document
- magila 8y agoCalling this "Unexpected" seems like a bit of a stretch. In particular this part: > Of course, in the server space, we've known for a long time that maximum efficiency occurs with a high number of cores running at lower frequencies, and that efficiency trumps performance on machines with high core counts. But I never considered that the consumer Ryzen CPUs could also benefit from the same thing until now. makes no sense. This principal applies to all CPUs from the smallest SoCs to the largest server CPUs, why on Earth would you not expect it to apply to desktop CPUs? You could do the same thing with a 6 core i7 and 2133 memory. Intel CPUs have long supported an adjustable power limit to constrain operating frequency based on power consumption just like he describes for Ryzen.
- ceratopisan 8y agoYou are confusing principle with implementation. Reducing clock speed reduces power usage and you can compensate with more cores is indeed a truth. However, finding that option in consumer hardware has been relatively difficult. That is the surprise indicated.
- BeeOnRope 8y agoNote that the author is not claiming that he compensated with more cores. He is claiming that the performance is roughly the same, at the same core count (8), regardless of frequency.
- magila 8y agoRyzen is known to be memory constrained even with much faster memory than he used. It is completely predicable that he found his CPU to be severely starved for memory bandwidth thus enabling him to reduce operating frequency without penalty. This is like putting an LS engine in an otherwise stock Miata and acting surprised that you can run the engine at lower RPM and still put in good lap times.
- AstralStorm 8y agoAre you kidding me? That's only true if the constraint can be removed in a hardware upgrade. Apparently latter day Xeons are not much better at hiding memory latencies than Zen and no longer outrun it as much like they did Bulldozer on other operations which made the latencies irrelevant. In other words, he's reached peak CPU. As in a faster unit will not speed it up, and more cores can only do that to a point. Amdahl law (power efficiency variant) and also memory controllers say hello.
- wolf550e 8y agoNot every workload is memory bandwidth bound like his "make -j16" compile. Some workloads need memory latency or fast inter-core (and inter-socket) operations (e.g. RDBMS OLTP), some need CPU throughput (e.g. HPC), some need best possible single thread CPU performance (e.g. some gaming). As he wrote, CPUs are most efficient (compute per Watt) at a specific frequency, and if his CPU mostly waits for RAM, this can be done at low power. It's probably possible to create x86-64 CPUs with narrower backends (fewer execution units) with microcode-emulated 128 and 256 bit registers/operations (and maybe even emulated FPU) and get a cheaper and faster build server, if it was economical to fab such narrow-use-case chips (those would be good for redis/memcached too I imagine).
- Dylan16807 8y agoZen is already light on vector units, and microcodes 256-bit operations. It's certainly possible to build a more lightweight core, but most of that work is reducing the complexity of the out-of-order machinery. The FPU+ALU is under a quarter of each Zen core. https://en.wikichip.org/w/images/c/cb/amd_zen_core_%28annotated%29.png https://en.wikichip.org/w/images/c/cb/amd_zen_core_%28annota...
- wolf550e 8y agoYou do want it to out-of-order and branch predict and speculate enough to issue speculative RAM reads as soon as a possibly needed address is available, to hide RAM latency (as long as rollbacks of speculatively executed operations hide the loaded values in the cache so the speculation leaves no side effects), that is important for performance. In that picture (thanks!) I see the FPU is big, the decoder is big, the branch predictor is big, the rest is probably needed. Maybe emulated FPU is good for some workloads, maybe ability to program in microinstructions instead of x86-64 is useful too. But maybe silicon area is not the expensive thing (dark silicon, etc.).
- fulafel 8y agoDo you have a link about microcoded avx256? I would think it would be way too slow.
- ebikelaw 8y agoHe doesn't seem to mention what the build times were with the Xeon.
- loeg 8y agoHere's the same author from pretty recently comparing the 2990WX to some Xeons (he says E5-2620 but doesn't mention which version — could be anything from Sandybridge to Broadwell): http://apollo.backplane.com/DFlyMisc/threadripper.txt http://apollo.backplane.com/DFlyMisc/threadripper.txt
- zrm 8y agoThe only available E5-2620 with that number of cores per socket is Broadwell.
- BeeOnRope 8y agoI think the claim that parallel compilation with gcc is memory bandwidth bound is unlikely. gcc is known to be a very pointer-chasy, branch-mispredicty load that is highly sensitive to memory latency - far from a streaming load that is sensitive to raw bandwidth. Still, the conclusion holds: if most of the time is spent waiting for values to come back from memory, a higher core frequency has strongly diminishing returns.
- AstralStorm 8y agoCloser to zero returns and you still get to improve latency hiding and memory controller design. (including bit width and block sizes) Even more cache ways won't help too much in this workload.
- BeeOnRope 8y agoUnless the AMD design is unusual it is not very close to zero return: a significant part of the "path to memory" involves things run at the core clock, in particular everything from the core to the L2 and probably some part of the coordination logic which communicates with the "uncore". I'm not sure about AMD chips, but on some chips there is a relationship between the uncore speed and the core speed: e.g., the uncore speed might often be the same as the maximum core speed for any core on the socket. Adding to that, there are other effects that allow core frequency to leak into the performance of memory-bound programs, such as a higher frequency allowing the core run ahead more quickly to get more memory requests in flight, recover more quickly after a branch misprediction, etc. Try it sometime: find something which is really memory bound and crank the frequency way down: there will probably be a significant effect, but not nearly in proportion to the frequency difference.
- yazr 8y agoI wonder if his results still hold for gcc -O3? That be much more CPU-bound. Yes - some optimization will do global traversals. I wonder if javac/clang has the same characteristics as gcc.
- 8y ago
- deleted 8y ago[deleted]
- mekpro 8y agoThis is to be expected. Since 180Watt is not the default TDP of Ryzen 2700X, the default TDP is 105Watt. https://www.amd.com/en/products/cpu/amd-ryzen-7-2700x https://www.amd.com/en/products/cpu/amd-ryzen-7-2700x Which mean the CPU is already shipped with the reasonable performance/watt TDP and over-TDP it will give diminish return in performance gain. However, It would be interesting to see benchmark in much lower TDP than 105Watt and see how far the TDP can go down before big drop in performance.
- manual 8y agoExcellent look into this by user "The Stilt" can be found here: https://forums.anandtech.com/threads/ryzen-strictly-technical.2500572/page-72#post-39391302 https://forums.anandtech.com/threads/ryzen-strictly-technica... It looks something like this: 4GHz 120W, 3.8 90W, 3.6 65W, 3.4 50W, 32 42W, 3.0 33W, 2.0 13W. This excludes the SOC.
- tgtweak 8y agoStilt is a magician in getting peak performance per watt out of everything. Down to tweaking individual straps for memory timing on binary firmwares for amd graphics cards. 3.6 @65W is impressive, almost stock speed at nearly half tdp.
- jandrese 8y agoBasically a 10% performance penalty for a 45% power savings. And you might not even see that 10% performance penalty if you're machine is bottlenecked on memory/storage.
- polskibus 8y agoHas anyone encountered webpack and C# compilation benchmarks that compare Ryzen and Intel?
- cestith 8y agoShort form and generalized: when one subsystem of a larger system is not your bottleneck, it's often possible to lower the resources for that part of the system without impacting overall performance.
- ulzeraj 8y agoIs this the amiga hacker/dragonflyBSD main dev? Why nobody has handed a 2990wx to this man?