30 ms·
GMP damaging Zen 5 CPUs?
- dkiebd 1y ago[flagged]
- internetter 1y agoPhotos taken by engineer, not photographer.
- dkiebd 1y agoWhat kind of camera is using the engineer that doesn't have autofocus and generates 378x283 photos?
- craftkiller 1y agoLooking at the AM5 pinout[0], it looks like those pins are VDDCR and VSS. There might be a little bit of PCIe sprinkled in towards the outer edges, but I'm not 100% on the orientation of this pinout vs the orientation of the CPU. I don't know anything about electricity so I've got nothing else to add. [0] https://upload.wikimedia.org/wikipedia/commons/2/2d/Socket_AM5_pinmap.svg https://upload.wikimedia.org/wikipedia/commons/2/2d/Socket_A...
- raverbashing 1y agoThis is a nice guess but the likelihood that actual silicon area is closely connected to the pins in that area is not so obvious
- nsteel 1y agoIsn't almost every other pin going to be power/ground on a high-power chip like this? On both the package and the die.
- topspin 1y agoAs there is ongoing drama with Zen 5 and power issues, there are people with the instruments and the motivation to investigate this. You should consider contacting Gamers Nexus, and help them to get your test suite running. They can measure power draw and do a thermal analysis of this CPU, and they'd likely be eager to do it, given the possibility of making a bunch of dramatic YouTube content about design flaws in widely used hardware. That's pretty much their whole schtick in recent years. > Modern CPUs measure their temperature and clock down if they get too hot, don't they? Yes. It's rather complex now and it involves the motherboard vendor's firmware. When (not if) they get that wrong CPUs burn up. You're going to need some expertise to analyze this.
- fxtentacle 1y agoHe's a bit sensationalist, yes, but I am thankful that he saved us from buying affected Intel CPUs.
- spookie 1y agoNot sure of sensationalist or just doing great reporting. I take him as one of the last good tech journalists on the platform.
- hnuser123456 1y agoGN wasn't the first to break the story the 13/14th gen was defective. The thousands and thousands of users experiencing the issues collectively noticed pretty quick. If anything, there was a period where he was saying "We've talked to Intel but we won't say anything yet until they do."
- BoorishBears 1y ago[flagged]
- wiredpancake 1y agoThe only real problem with GN is Steve is a bit of an egotist when it comes to content creators who do less technical analysis, like LTT or Jayz. He never really got over the stuff with Linus and doubled down on stupid things. I think they both have a great place in the tech scene and LTT's videos of recent have been a lot better quality and researched then yesteryear.
- tester756 1y agoMy Ryzen CPU recently died too! wtf
- FuriouslyAdrift 1y agoASRock motherboard?
- tester756 1y agoGigabyte
- LASR 1y agoZen5?
- tester756 1y agoRyzen 7
- fxtentacle 1y ago"We suspect that GMP's extremely tight loops around MULX make the Zen 5 cores use much more power than specified, making cooling solutions inadequate." I feel like if this was heat related, the overall CPU temperature should still somewhat slowly creep up, thereby giving everything enough time for thermal throttling. But their discoloration sure looks like a thermal issue, so I wonder why the safety features of the CPU didn't catch this...
- jeffbee 1y agoAre we talking "slowly" in a relative sense? A silicon die of this size has a thermal mass (guessing) around 10⁻³ J/K but a power dissipation rate over 200W, so it can rise from room temperature to junction temperature limits almost instantly.
- topspin 1y agoPeople without a background in electronics don't appreciate what modern CPUs and GPUs are doing: the amount of current flowing through these devices is just mind blowing. With adequate cooling, a Ryzen 9 9950X is handling somewhere in the neighborhood of 150-200 amps under high load.
- nisegami 1y agoI initially scoffed at the 150-200 amps. But I know core voltage is usually in the neighbourhood of 1V so to draw 200W, you really would have to basically be moving 200A of current. That's wild.
- mlyle 1y agoYup. P=IV is really surprising when you get to high power parts at low core voltages. Needless to say, you need lots of transistors and phases on voltage conversion, and you need lots and lots of plane area. (And,... 200A is the average when dissipating 200W. So how high are the switching currents? ;)
- 1y ago
- tux3 1y agoThe room temperature or precise way the paste was applied should not matter. Modern CPUs have very advanced dynamic voltage and frequency scaling (DVFS), which accounts for several sensors, including temperature. These big x86 CPUs in stock configuration can throttle down to speeds where they can function with entirely passive cooling, so even if the cooler was improperly mounted, they'd only throttle. All that to say, if GMP is causing the CPU to fry itself, something went very wrong, and it is not user error or the room being too hot.
- secabeen 1y agoI would be interested to see if they had the same result with PTM7950 thermal material instead of paste. I've seen significantly better temps with these modern phase-change compounds, and they essentially eliminate application errors.
- mk_stjames 1y agoThis was my first question as well- I thought it had been a long, long time since you could fry a CPU by taking away the heatsink. As in... what, AMD K6 / early Pentium 4 days was the last time I remember hearing about cpu cooler failing and frying a cpu?
- Twirrim 1y agoIt was some time around then. I remember AMD being late to it vs Intel.
- bob1029 1y agoCould be the power supply and load profile? I've heard some really wild noises coming out of my zen4 machine when I've had all cores loaded up with what is best described as "choppy" workloads where we are repeatedly doing something like a parallel.foreach into a single threaded hot path of equal or less duration as fast as possible. I've never had the machine survive this kind of workload for more than 48 hours without some kind of BSOD. I've not actually killed a cpu yet though.
- bee_rider 1y agoIs that, like, an intentional stress-test for the hardware that you’ve come up with?
- bob1029 1y agoNo. It is just how the algorithms play out: 1. Evaluate population of candidates in parallel 2. Perform ranking, mutation, crossover, and objective selection in serial 3. Go to 1. I can very accurately control the frequency of the audible PWM noise by adjusting the population size.
- userbinator 1y agoI've never had the machine survive this kind of workload for more than 48 hours without some kind of BSOD. Then you shouldn't trust the results of your work either, as that's indicative of a CPU that's producing incorrect results. I suggest lowering the frequency or even undervolting if necessary until you get a stable system. ...and yes, wildly fluctuating power consumption is even more challenging than steady-state high power, since the VRMs have to react precisely and not overshoot or undershoot, or even worse, hit a resonance point. LINPACK, one of the most demanding stress tests and benchmarks, is known for causing crashes on unstable systems not when it starts each round, but when it stops.
- bob1029 1y agoThe results might be invalid for one generation but the model is resilient to these kinds of events overall. Far more resilient than my operating system is. Randomly flipped genome bits could even be beneficial for escaping local minima and broken RNG in evolutionary algorithms. One bad evaluation won't throw the whole thing off. It's gotta be bad constantly.
- tw04 1y agoThat looks like a combination of improperly mounting the heatsink and noctuna being wrong in their recommendation to offset it. I’d imagine for gaming cooling one side more makes sense but my completely uneducated guess is that GMP is working a different part of the CPU than gaming does.
- jsheard 1y agoThis is what Zen5 looks like under the IHS: https://i.imgur.com/j85YUzX.jpeg https://i.imgur.com/j85YUzX.jpeg Everything is offset towards one side and the two CPU core clusters are way towards the edge, offset cooling makes sense regardless of usage.
- toast0 1y agoThey had failures with standard mounting and offset mounting. Also, take a look at a delidded 9950; the two cpu chiplets are to one side, the i/o chiplet is in the middle, and the other side is a handful of passives. Offsetting the heatsink moves the center of the heatsink 7mm towards the chiplets (the socket is 40mm x 40mm), but there's still plenty of heatsink over the top of the i/o chiplet. This article has some decent pictures of delidded processors https://www.tomshardware.com/pc-components/overclocking/delidded-amd-ryzen-9-9950x3d-runs-23-degrees-cooler https://www.tomshardware.com/pc-components/overclocking/deli...
- pharrington 1y agoI'd assume both GMP and any CPU intensive game just prefer the performance cores.
- jsheard 1y agoAMDs desktop chips don't have distinct P and E cores, they're all P cores. AMD do have an E core design but it's currently only used in mobile and server parts.
- pharrington 1y agoGotcha. Apparently Intel's marketing's gotten to me. I haven't really been keeping up with this stuff, so whenever I read about P & E cores in the past, I think I just assumed that was a thing both Intel & AMD were doing, without considering the source material too closely.
- nromiun 1y agoHow is that possible? Even if the chip did not get enough cooling it should have been just throttled heavily.
- jsheard 1y agoModern silicon is so dense and heats up so fast that throttling is easier said than done. I think they have to model and predict the thermals ahead of time nowadays, because by the time they could react to a temp sensor alone, the chip might already be toast.
- tliltocatl 1y agoMaybe the throttling circuitry/firmware simply doesn't have enough time to react.
- db48x 1y ago> The so-called TDP of the Ryzen 9950X is 170W. The used heat sinks are specified to dissipate 165W, so that seems tight. TDP numbers are completely made up. They don’t correspond to watts of heat, or of anything at all! They’re just a marketing number. You can't use them to choose the right cooling system at all. https://gamersnexus.net/guides/3525-amd-ryzen-tdp-explained-deep-dive-cooler-manufacturer-opinions https://gamersnexus.net/guides/3525-amd-ryzen-tdp-explained-...
- einpoklum 1y agoWow, I can't believe how BS this TDP is! I feel like a total idiot! I've always assumed it's sorta-kinda a tight upper bound on power consumption, perhaps with some allowance for "imperfections" in the dissipation properties of the CPU, and that I shouldn't sweat the details. Couldn't this count as false/misleading advertizing though?
- vel0city 1y agoIts pretty insane to see someone say something like: “TDP is about thermal watts, not electrical watts. These are not the same.” Watts are watts. But yeah, TDP means nothing. If you stick plenty of cooling and run the right motherboard board revision your "TDP" can be whatever you want it to be until the thing melts.
- o11c 1y ago"TDP is about average watts, not peak watts" would be an honest way of saying it.
- vel0city 1y agoBut in the end that's still not actually true in many modern desktop chips. You can take a 65W part, and with a "stock" motherboard firmware, good cooling, and the right workload end up averaging way more than 65W. Or if you have it in a hot room it just might end up using less than 65W. TDP is more of a rough idea of how much power the manufacturer wanted to classify the part as. It ultimately only loosely relates to the actual heat or electrical usage in practice.
- fithisux 1y ago[flagged]
- toast0 1y ago> The x86-64 ISA is a POS BTW. > It needs major refactoring. The backwards compatibility is killing the platform. People have been saying that for 30 years at this point. You can dislike the ISA all you want, but the platform doesn't seem to be dieing all that fast, and compatibility helps it live. Compatibility is how you sell chips: windows on arm doesn't sell because people are concerned about running their apps; android on x86 didn't sell well in part because many apps with native code didn't include x86 binaries; mac os on arm sells because Apple made things work and pushes hard. A better ISA doesn't prevent chips from burning up. Apple's chips don't burn up because they stay away from high clocks and high power, which is easy when they force vertical integration.
- fithisux 1y agoThe Intel ISA CPUs have too many features for backwards compatibility. This puts a lot of design strain, CPU have to support everything, have to emulate behaviors. Too much bloat on the CPU.
- account42 1y agoEmulating behaviors are not on the hot paths if they are even implemented in silicon at all rather than microcode that may as well not exist if you don't use those features. This is entirely irrelevant to the issue of CPUs burning up under normal use.
- 0x457 1y agoI have a pretty old atom based mini pc that is probably still significantly faster than any RPI today.
- fithisux 1y agoNot my experience.
- FuriouslyAdrift 1y agoMost likely it's the motherboard. ASRock is getting nailed right now for unstable XMP and CPU voltages (it's recommended to undervolt a little just in case). The Asus Prime B650M motherboards they are using aren't exactly high end.
- kvemkon 1y agoAnd the close-up photos of the socket with pins are missing.
- J_Shelby_J 1y agoMy friend just had an ASRock board cook his AMD CPU. Apparently a very common problem.
- wmf 1y agoYikes, this is the cheapest motherboard and failed Hardware Unboxed VRM tests. https://youtu.be/DTFUa60ozKY?t=744 https://youtu.be/DTFUa60ozKY?t=744
- caycep 1y agoconversely the asrocks actually did pretty good in that test...
- aidenn0 1y agoCan you link to a reputable source for what settings I should use on my asrock motherboard? I'd like to avoid this.
- FuriouslyAdrift 1y agoNo more than 1.2 volts on vsoc... but YMMV. "According to new details from Tech Yes City, the problem stems from the amperage (current) supplied to the processor under AMD's PBO technology. Precision Boost Overdrive employs an algorithm that dynamically adjusts clock speeds for peak performance, based on factors like temperature, power, current, and workload. The issue is reportedly confined to ASRock's high-end and mid-range boards, as they were tuned far too aggressively for Ryzen 9000 CPUs." https://www.tomshardware.com/pc-components/cpus/asrock-attributes-premature-ryzen-9000-cpu-failures-to-aggressive-pbo-settings-per-youtuber https://www.tomshardware.com/pc-components/cpus/asrock-attri...
- on_the_train 1y agoWhat is gmp?
- kgwgk 1y agoThe domain has the answer: https://gmplib.org/ https://gmplib.org/
- gus_massa 1y agoFrom https://gmplib.org/#WHAT https://gmplib.org/#WHAT > What is GMP? > The GNU Multiple Precision Arithmetic Library > GMP is a free library for arbitrary precision arithmetic, operating on signed integers, rational numbers, and floating-point numbers. There is no practical limit to the precision except the ones implied by the available memory in the machine GMP runs on. GMP has a rich set of functions, and the functions have a regular interface. Many languages use it to implement long integers. Under the hood, they just call GMP. IIUC the problem is related to the test suit, that is probably very handy if you ever want to fry an egg on top of your micro.
- beezle 1y agoAt first I thought it was Green Mountain Power ;)
- protomikron 1y agoValid question i think in this context. I knew about GNU multiprecision library, but thought that couldnt be it, as it's "just" a highly optimized low level bit fiddling lib (at least thats my expectation without looking into the source), so it's strange why it could be damaging Hardware ...
- account42 1y agoTFA is on gmplib.org which kind of answers the question though.
- wrs 1y agoNo actual die temperature measurements? That would seem a lot more relevant than the ambient temperature.
- wtallis 1y agoDie temperature readings aren't particularly helpful these days with desktop parts that will (depending on the power management settings) more or less keep increasing the clock speed until they reach ~90°C and just stay there. Upgrading from a bad/undersized heatsink can easily have only a tiny effect on temperature but have the effect of significantly increasing clock speed and power.
- mqus 1y agoAren't they at least useful for ruling out any anomalies there? Like the die temp being 110°C constantly? Imho the die temperature is very important here, even if not interesting.
- userbinator 1y agobut have the effect of significantly increasing clock speed and power. Ironically, if these failures are due to excessive automatic overvolting like what happened with Intel's a year ago), worse cooling would cause the CPU to hit thermal limits and slow down before harmful voltages are reached. Conversely, giving the CPU great cooling will make it think it can go faster (and with more voltage) since it's still not at the limit, and it ends up going too far and killing itself.
- mastax 1y agoEnthusiast-oriented motherboards often default enable Precision Boost Overdrive, causing higher power and temperature limits for longer periods. To run the CPU at “stock” you need to go in and disable that. Their default Load Line Calibration might be aggressive as well.
- gdwatson 1y agoWhich motherboards enable PBO out of the box? That’s crazy! I know that motherboard manufacturers set some sketchy default turbo durations for Intel CPUs back when Intel was cagey about the spec and let them get away with it, but I thought that AMD was stricter about such things.
- lloydatkinson 1y agoOne day I’ll understand why some websites refuse to have a way of navigating to the home page. I had to edit the URL in the address bar. I just wanted to find out what GMP is.
- caycep 1y agoI wonder if the risk is mitigated if you turn off PBO and turn on Eco Mode?
- gpapilion 1y agoGradual damage is consistent with over heating. I've seen racks of servers do the same thing. Overall, there is a continued challenge with CPU temperatures that requires much tighter tolerances both in the thermal solution. The torque specs need to be followed and verified that they were met correctly in manufacturing.
- giantg2 1y agoNot that it makes a huge difference since they are supposed to downclock when hot, but what was the actual cooler being used? It doesn't say in the article. My guess is that it's aircooled being only 165W max, but aircooled is not recommended for most newer high end CPUs.
- T-A 1y agoA quick search on the NH-U9S shows it's a compact cooler for small systems, rated for up to 140 W (see e.g. [1]). The 9950X's TDP (Thermal Design Power) is 170 W, its default socket power is 200 W [2], and with PBO (Precision Boost Overdrive) enabled it's been reported to hit 235 W [3]. [1] https://www.overclockersclub.com/reviews/noctua_nh_u9s_cpu_cooler/ https://www.overclockersclub.com/reviews/noctua_nh_u9s_cpu_c... [2] https://hwbusters.com/cpu/amd-ryzen-9-9950x-cpu-review-performance-thermals-power-analysis/ https://hwbusters.com/cpu/amd-ryzen-9-9950x-cpu-review-perfo... [3] https://www.tomshardware.com/pc-components/cpus/amd-ryzen-9-9950x-cpu-review/4 https://www.tomshardware.com/pc-components/cpus/amd-ryzen-9-...
- stouset 1y agoThat’s a good catch, but don’t modern CPUs thermally throttle, rather than risk damage? Not that you should rely on this with an underpowered cooling solution but I would expect worse performance, not a fried chip.
- spoaceman7777 1y agoNot really a lot it can do rapidly enough if there's only thermal paste on half the CPU. It sounds like the user likely did the opposite of the "offset seating" of the heatsink that Noctua recommended.
- account42 1y agoThere thermal paste on the whole CPU in TFA, it's just thinner on one side because there was more pressure there. Or are you looking at the pic of the heat sink, which is larger than the CPU heat spreader and thus only partially covered by paste?
- BugsJustFindMe 1y agoNoctua does not use TDP for their heatsinks and instead have CPU compatibility charts. They say it's fine, with "medium turbo/overclocking headroom". https://ncc.noctua.at/cpus/model/AMD-Ryzen-9-9950X-1831 https://ncc.noctua.at/cpus/model/AMD-Ryzen-9-9950X-1831
- codezero 1y agoI recently built a 9950x3d system and it definitely runs hotter/louder than my other builds. To be fair, it also has a 5090 in it, but I liquid cool the 9950 and over all the system is just hot. If I saw 165W rated on a 170W system, I think that kind of just answers it, I see no reason not to overcool a system with high end electronics on it, and there's no reason to toe the line so closely.
- chaoskitty 1y agoThis isn't good. Then again, the amount of power going in to these CPUs is way too high. Take the AlphaServer DS25. It has wires going from the power supply harness to the motherboard that are thick enough to jump a car. The traces on the motherboard are so thick that pictures of the light reflecting off of them are nothing like a modern motherboard. The two CPUs take 64 watts each. Now we have AMD CPUs that can take 170 watts? That's high, but if that's what the motherboards are supposed to be able to deliver, then the pins, socket and pads should have no problem with that. Where's AMD's testing? Have they learned nothing watching Intel (almost literally) melt down?
- wkat4242 1y ago> Take the AlphaServer DS25. It has wires going from the power supply harness to the motherboard that are thick enough to jump a car. The traces on the motherboard are so thick that pictures of the light reflecting off of them are nothing like a modern motherboard. The two CPUs take 64 watts each. I am not involved in power VRM for modern moderboards. But I can imagine they are some some smart stuff like compensating for transport losses by increasing the voltage somewhat at the VRM so the designed voltage still outputs at the CPU. Of course this will cause some heating in the motherboard but it's probably easily controlled. In the day of the alpha that kind of thing would have been science fiction so they had no alternative but to minimise losses. You can't use a static overvoltage because then when the load drops the voltage coming out will be too high (transport loss depends on current). Also, in those days copper cost a fraction of what it costs now so with any problem just doing 'moah copper' was an easy solution. Especially on server hardware like the Alpha with big markup. And server hardware is always overengineered of course. Precisely to prevent long-term load problems like this.
- hpcjoe 1y agoI noticed the comments pointing out that TDP is a marketing number, and max power draw for this part can be higher. The cooling seems to have been inadequate. A rule of thumb I use for cooling is, you can rarely have too much. You should over-engineer that aspect of your systems. That and the power supply. I have a 7950x, with a water block capable of sinking up to 300W. Under heavy load, I hear the radiator fans spinning up, and I see the cpu temp hover around 90-93 C. That is ok, though cooler would be better. My next build (this one is 2 years old) will also use a water block, but with a higher flow rate, and a better radiator system. I like silent systems, though I don't like the magic smoke being released from components.
- thway15269037 1y agoI don't know about GMP, but I recently built a PC with 9950X3D. As part of initial testing, I ran Prime95 for 48 hours. Everything ran stable, but I noticed that part of the tests, I think it was FFT or something like that, caused incredibly sharp increase in temp. We are talking 60C average in the rest of the test vs immediate (less than a 5 seconds) 95+ degrees when that FFT thingie started. It was very weird. That's when I discovered actually ancient term "power virus". Anyway, after talking to different people I dismissed this weird behavior and moved on. Reading this makes me worry I actually burned mobo in that testing.
- userbinator 1y agoTry LINPACK, it's even more stressful than Prime95.
- jmb99 1y agoDifferent use patterns will result in different temperatures. Very tight math loops (no memory/IO wait) will lead to higher temperatures than something that that relies on L2/3 cache or main memory, even though they’ll both report “100% CPU use” and probably use similar amounts of power. And, different operations will produce heat in different areas of the die; depending on physical layout, some operations might generate heat in a tiny cluster, whereas some others might generate heat in larger spread out areas. Even though both of those cases might use the same amount of power and generate the same amount of heat, the temperatures will be drastically different due to the heat concentration. Iirc the FFT step uses AVX, and on Zen 5 that’ll be AVX-512. It should keep 100% of the required data in L1 caches, so you’re keeping the AVX units busy literally 100% of the time if things are working right. The rest of the core will be cold/inactive, so if you’re dumping an entire core’s worth of power into a teeny tiny ALU, which is gonna result in high temps. Most (all?) processors downclock under heavy AVX load, sometimes by as much a 1GHz (compared to max boost), because a) the crazy high temperatures results in more instability at higher frequencies, and b) if the clocks were kept high, temperatures would get even higher.
- monster_truck 1y agoAll other potential causes aside, including the likely most-relevant of motherboard companies exceeding recommended defaults for power delivery: Running a cooling solution good for less than the TDP (which is NOT the max power, which tends to be about 30% higher than the TDP on these) is frankly extremely dumb. I've seen x950 processors of every generation pull at least double that on extreme workloads. I think it speaks to them being a bit clueless that they did not manually lower the thermal limits. You can cut the power and thermal limits by wild amounts and barely lose 15% multicore performance.
- userbinator 1y agoWe don't overclock or overvolt or play other teen games with our hardware. But doesn't the hardware "overclock" and "overvolt" automatically these days? This reminds me of the Intel CPUs with similar problems a year ago, and AFAIK it was caused by excessive voltage: https://news.ycombinator.com/item?id=41039708 https://news.ycombinator.com/item?id=41039708
- mrheosuper 1y agoAlso "play other teen games" should not damage your cpu.
- wkat4242 1y agoEhh I've seen some of these teen games involve complete immersion in oil or even water (can be done as distilled water doesn't conduct but if only a pinch of salt gets into it...). Or even more extreme things like liquid nitrogen. This can have all sorts of weird effects on CPUs not designed for that kinda stuff (e.g. thermal contraction to temperatures under low load way below spec, or cracking due to extreme thermal gradients).
- wkat4242 1y ago> But doesn't the hardware "overclock" and "overvolt" automatically these days? If it's done by the manufacturer it's within spec of course. As designed. The overclock game was all about running stuff out of spec and getting more performance out of it than it was supposed to create.
- account42 1y agoIf anything, what replaced overclocking is not PBO and similar features to dynamically clock the CPU but rather binning that lets better performing samples be sold with higher base frequencies than other samples of the same design.
- protocolture 1y agoI have a Ryzen and it ran fine until one day, after some load, it wont run at all with the virtualisation options turned on anymore. Having read all I can on the issue its largely been ignored by AMD. If its some kind of thermal runaway issue that would not surprise me.
- esseph 1y agoI had a bios reset itself to defaults before, and some AMD boards don't have all the required virtualization options on by default. I only realized this happened because every time I had ever upgraded firmware, I always had to go back and set XMP settings, one or two other things, and than cpu virt option.
- Seattle3503 1y agoSame with secure boot for me. Kinda makes sense that a BIOS upgrade would wipe the config. Its that or manage schema migrations.
- rurban 1y agoThe /r/asrock reddit is full of such stories. In my case I've burned two server motherboards with those watercooled 9950X chips. The CPU is still fine though. It's happening with all H100's on 100%. Too much power draw we assume.
- wkat4242 1y ago> Did the CPUs die of heat stroke? Modern CPUs measure their temperature and clock down if they get too hot, don't they? They do, but the thermal sensors are spread out a bit. It could be that there's a sudden spot heating happening that's not noticed by one of the sensors in time.
- Michvalwin 1y agoAbout two months ago My 9950X3D died. I presume it the memory controller died because my asus board reported a ram problem. It was with a 2x48GB 6000 ram. Which was not supported by the board (at least on the docs). I also limited the max PBO by 400Mhz and limited VDDIO to 1.3V. But it still died while shutting down cachyOS.
- edgineer 1y ago"We don't overclock or overvolt or play other teen games with our hardware." You overclock as a teen so that as an adult you know to verify your CPU's voltage, clock speeds, and temperature at a minimum when you build your own system. They made no mention of monitoring of CPU temperature, ECC corrected/detected errors, or throttling. They then ran CPU benchmark loads for several months on the system. "The so-called TDP of the Ryzen 9950X is 170W. The used heat sinks are specified to dissipate 165W, so that seems tight." Yikes. You need a heatsink rated much higher. These CPUs were overheated for months.
- Gracana 1y agoCPUs with stock cooling solutions will turbo boost up to max temp and stay there, that's completely normal and shouldn't cause a CPU to physically burn up, even if you do it for months.
- edgineer 1y ago"boost up to max temp and stay there" At stock settings CPUs will boost depending on many factors. Once one of several different limits is hit the CPU will not boost as high, trying to find a steady state where it stays below Tjmax. Note the following: the 9950x does not come with a stock cooler. AMD recommends water cooling for the 9950x. Transistor lifetime decreases exponentially with temperature. I'd expect that a 9950x under sustained load paired with a "165W" cooler would not only not boost, but would throttle to below base clocks. In the case of CPU cooling, I don't agree that relying on the CPU's thermal safety nets to continuously regulate the system to avoid damage is good practice. With additional cooling to ensure it never reaches Tjmax, this also will result in better CPU performance, a tangible benefit. Had the author monitored his systems, he would have observed high temperature and throttling. Yes, in 2025 it's arguable that a CPU's safety net should be reliable wrt temperature, even if you run with no heatsink at all. I also agree that TDP specifications are unclear. But the bottom line is that you should pay the extra $100 or so to cool your CPU properly. It will be faster and more reliable. Please take care of your equipment; do not take it for granted.
- rozab 1y agoI've not built a new PC in a while, is it normal for the cooler to only cover 2/3 of the IHS?? Seems like that leaves a lot of cooling performance on the table
- Ono-Sendai 1y agoThat's wild. My 9950X died for some reason, never overclocked.
- OhMeadhbh 1y agoWhen I worked at Linden Lab we had a deal going with IBM. Either as part of that deal or in an attempt to impress our larger partner, many of us got Thinkpads. I actually kind of like them, since the cost wasn't coming out of my budget. Inside Linden, about 90% of meetings were held in-world, so we constantly had the Second Life viewer up. About three months later our Thinkpads started failing. Apparently they thought people who would a) buy a thinkpad and b) use it to play video games wouldn't be playing video games 12 hours per day (though as many have pointed out, does one "play" Second Life? especially if you're using it for work.) After 3 months of use, the Second Life client had caused sufficient heating cycles so as to delaminate the PCB under the GPU. I'm sort of proud of this. Our software was dangerous.
- Gracana 1y agoDelaminating the PCB is absolutely wild. I'm really fascinated by the idea of doing meetings in Second Life, though. What did people use for avatars at work?
- OhMeadhbh 1y agoWe used whatever avatars we put together. Here's a link to Philip's avatar, which he intentionally kept basic for quite some time: https://community.secondlife.com/forums/topic/517881-help-me-with-my-new-avatar/ https://community.secondlife.com/forums/topic/517881-help-me... And here's a shot of mine: https://www.flickr.com/photos/opensourceobscure/2476204733/in/pool-dazzle/ https://www.flickr.com/photos/opensourceobscure/2476204733/i... And about in-world meetings. One thing Second Life did VERY WELL was it was always clear who was speaking. If someone was talking, there were green arrows sort of exploding out of their avatar's head. You couldn't miss it. Web-Ex at the time was HORRIBLE in this regard. Teams, Google Meet and Zoom are a little better than Web-Ex, but when meeting in Second Life, you could adjust your camera to get a good view of everyone in the meeting which also helped out.
- dukezzz 1y agoi think that if the cpu tdp is 170W and your cooling solution is rated for 165W and both cpu broke under sustained full load you replied yourself. Plus i don't see how reducing the contact ara with the skewed mounting can benefit in any way the heat transfer
- itvision 1y agoAMD's desktop CPU TDP numbers have been misleading for over a decade. I have no idea why they do it this way. Everything they list must be multiplied by about 1.35. For example, a 170W TDP CPU requires 230W of dissipation. Not 170W. AMD fans don't give a damn about that though. "It's just fine" (tm).
- ac130kz 1y agoASRock has Zen 5 CPUs dying with stock settings from brief core voltage spikes on idle (max core frequencies). I believe either PBO has to be enabled to allow undervolting headroom or VSOC has to be permanently fixed to a value lower than 1.2V.