13 ms·
July 2024 Update on Instability Reports on Intel Core 13th/14th Gen Desktop CPUs
- issafram 2y ago[flagged]
- sangeeth96 2y agoYeah sure, calling out Intel for lack of any good updates over the crashing laptop/desktop CPUs and demanding a recall after giving them such a long time to come up with a reasonable solution is definitely "weirdo" territory. FWIW, I have connections who splurged on these only to deal with BSODs all the friggin' time. Some of them even work at Intel.
- kjeldsendk 2y agoWithout those weirdos do you think Intel would be doing anything about this in public? And tell us how customers that bought the most expensive part from their lineup should feel about knowing that their cpu has been over voltaged from day one of operation..
- x3n0ph3n3 2y agoWe have yet to see - How much lifespan of these CPUs has already been lost and cannot be recovered by the microcode patch. - How much of a performance hit these CPUs will get after applying the patch.
- firebaze 2y agoNice that Intel acknowledges there are problems with that CPU generation. If I read this right, the CPUs have been supplied with a too-high voltage across the board, with some tolerating the higher voltages for longer, others not so much. Curious to see how this develops in terms of fixing defective silicon.
- tux3 2y agoRemains to be seen how the microcode patch affects performance, and how these CPUs that have been affected by over-voltage to the point of instability will have aged in 6 months, or a few years from now. More voltage generally improves stability, because there is more slack to close timing. Instability with high voltage suggests dangerous levels. A software patch can lower the voltage from this point on, but it can't take back any accumulated fatigue.
- giantg2 2y agoI was recently looking at building and buying a couple systems. I've always liked Intel. I went AMD this time. It seemed like the base frequencies vs boost frequencies were much farther apart on Intel than with most of the AMDs. This was especially true on the laptops were cooling is a larger concern. So I suspect they were pushing limits. Also, the performance core vs efficiency core stuff seemed kind of gimmicky with so few performance cores and so many efficiency cores. Like look at this 20 core processor! Oh wait, it's really an 8 core when it comes to performance. Hard to compare that to a 12 core 3D cached Ryzen with even higher clock... I will say, it seems intel might still have some advantages. It seems AMD had an issue supporting ECC with the current chipsets. I almost went Intel because of it. I ended up deciding that DDR5 built in error correction was enough for me. The performance graphs also seem to indicate a smoother throughput suggesting more efficient or elegant execution (less blocking?). But on the average the AMDs seem to be putting out similar end results even if the graph is a bit more "spikey".
- nullindividual 2y ago> It seems AMD had an issue supporting ECC with the current chipsets. AMD has the advantage with regards to ECC. Intel doesn't support ECC at all on consumer chips, you need to go Xeon. AMD supports it on all chips, but it is up to the motherboard vendor to (correctly) implement. You can get consumer-class AM4/5 boards that have ECC support.
- jwond 2y agoActually some of the 13th and 14th gen Intel Core processors support ECC.
- loufe 2y agoIntel cannot afford to be anything but outstanding in terms of customer experience right now. They are getting assaulted on all fronts and need to do a lot to improve their image to stay competitive.
- Joel_Mckay 2y agoTheir acquisition of Altera seemed to harm both companies irreparably. Any company can reach a state where the Process people take over, and the Product people end up at other firms. Intel could have grown a pair, and spun the 32 core RISC-V DSP SoC + gpu for mobile... but there is little business incentive to do so. Like any rotting whale, they will be stinking up the place for a long time yet. =)
- beacon294 2y agoCould you elaborate on the process people versus product people?
- basementcat 2y agoI would argue the fabrication process people at Intel are core to their business. Without the ability to reliably manufacture chips, they're dead in the water.
- Joel_Mckay 2y agoYou mean manufacturing "working chips" is supposed to be their business. It is just performance art with proofing wafers unless the designs work =3
- bgmeister 2y agoI assume they're referring to Steve Jobs' comments in this (Robert Cringely IIRC) interview: https://www.youtube.com/watch?v=l4dCJJFuMsE https://www.youtube.com/watch?v=l4dCJJFuMsE (not a great copy, but should be good enough)
- wnevets 2y agoAre the CPUs that received elevated operating voltage permanently damaged?
- Pet_Ant 2y agoThis is the most pressing question. If it was just a microcode issue a cooloff and power cycle ought to at least reset things but according to Wendel from Level 1 Tech, that doesn't seem to always be the case.
- kevingadd 2y agoThe problem is that running at too high of a voltage for sustained periods can cause physical degradation of the chip in some cases. Hopefully not here!
- chmod775 2y ago> can cause physical degradation of the chip in some cases. Not in some cases. Chips always physically degrade regardless of voltage. Higher voltages will make it happen faster.
- Pet_Ant 2y agoWhy do chis degrade? Is this due to the whiskers I’ve heard about?
- cesarb 2y ago> Why do chis degrade? Is this due to the whiskers I’ve heard about? No, tin whiskers are a separate issue, which happens mostly outside the chips. The keyword you're looking for is electromigration (https://en.wikipedia.org/wiki/Electromigration https://en.wikipedia.org/wiki/Electromigration).
- account42 2y agoYes, but usually this happens slowly enough that the chip will be long obsolete before the degredation becomes an issue.
- NBJack 2y agoI was concerned this would happen to them, given how much power was being pushed through their chips to keep them competitive. I get the impression their innovation has either truly slowed down, or AMD thought enough 'moves' ahead with their tech/marketing/patents to paint them into a corner. I don't think Intel is done though, at least not yet.
- magicalhippo 2y agoThere was recently[1] some talk about how the 13th/14th gen mobile chips also had similar issues, though Intel insisted it's something else. Will be interesting to see how that pans out. [1]: https://news.ycombinator.com/item?id=41026123 https://news.ycombinator.com/item?id=41026123
- tardy_one 2y agoFor server CPUs there's not a similar problem or they realize server purchasers may be less willing to tolerate it? I'm not all that thrilled with the prospect of buying Intels especially when wondering about waiting to 5 year out replacement compared to a few generations ago, but AMD server choices can be a bit limited and I'm not really sure how to evaluate if there may be increasing surprises more across the board.
- sirn 2y agoAre you talking about Xeon Scalable? Although they share the same core design as the desktop counterpart (Xeon Scalable 4th Gen shares the same Golden Cove as 12th Gen, Xeon Scalable 5th Gen shares the same Raptor Cove as 13th/14th Gen), they're very different from the desktop counterpart (monolithic vs tile/EMIB-based, ring bus vs mesh, power gate vs FIVR), and often running in a more conservative configuration (lower max clock, more conservative V/F curves, etc.). There has been a rumor about Xeon Scalable 5th Gen having the same issue, but it's more of a gossip rather than a data point. The issue does happen with desktop chips that are being used in a server context when pairing with workstation chipset such as W680. However, there haven't been any reports of Xeon E-2400/E-3400 (which is essentially a desktop chip repurposed as a server) with C266 having these issues, though it may be because there hasn't been a large deployment of these chips on the server just yet (or even if there are, it's still too early to tell). Do note that even without this particular issue, Xeon Scalable 4th Gen (Sapphire Rapids) is not a good chip (speaking from experience, I'm running w-3495x). It has plenty of issues such as slow clock ramp, high latency, high idle power draw, and the list goes on. While Xeon Scalable 5th Gen (Emerald Rapids) seems to have fixed most of these issues, Zen 4 EPYC is still a much better choice.
- tedunangst 2y ago
- christkv 2y agoThe amount of current their chips pull on full boost is pretty crazy. It would definitively not surprise me if some could get damaged by extensive boosting.
- brynet 2y agoCurious why Intel announced this on their community forums, rather than somewhere more official.
- samtheprogram 2y agoOptics / stock price
- guywithahat 2y agoThat’s probably where people are mostly likely to understand it. A lot of companies do this, especially while they’re still learning things.
- wmf 2y agoThese days people are more likely to see the announcement on YouTube, TikTok, or Twitter.
- slaymaker1907 2y agoThe first two require a lot more effort in video editing than creating a forum post. Plus, it’s just going to be digested and regurgitated for the masses by people much better at communicating technical information.
- cyanydeez 2y agoSobweird hearing high noise channels as the prefwrred distributiom
- paulmd 2y agothey did that too https://youtu.be/wkrOYfmXhIc https://youtu.be/wkrOYfmXhIc
- beart 2y agoBased on what I know about corporations, it's entirely plausible that the folks posting the information don't actually have access to the communication channels you are referring to. I don't even know how I would issue an official communication at my own company if the need ever came up... so you go with what you have.
- Covzire 2y agoJust want to say, I'm incredibly happy with my 7800X3D. It runs ~70C max like Intel chips used to and with a $35 air cooler and it's on average the fastest chip for gaming workloads right now.
- amiga-workbench 2y agoI'm also very happy with my 5800X3D, it was wonderful value back when AM5 had just released and DDR5/Motherboards still cost an arm and a leg. The energy efficiency is much appreciated in the UK with our absurd price of electricity.
- SushiHippie 2y agoSame, in my BIOS I can activate a "ECO Mode", which lets me decide if I want to run my 7950x on full 170W TDP, 105W TDP or 60W TDP. I benchmarked it, the difference between 170 and 105 is basically zero, and the difference to 60W is just a few percent of a performance hit, but way worth it, as it's ~0.3€/kWh over here.
- aruametello 2y ago(if you are running windows) you might want to check a tool called PBO2Tunner (https://www.cybermania.ws/apps/pbo2-tuner/ https://www.cybermania.ws/apps/pbo2-tuner/), you can tweak values like EDC,TDC and PPT (power limit) from the GUI, and it also accepts command line commands so you can automate those tasks. I made scripts that "cap" the power consumption of the cpu based on what applications are running. (i.e. only going all in on certain games, dynamically swaping between 65-90-120-180w handmade profiles) i made with power saving in mind given the idle power consumption is rather high on modern ryzens. edit: actually made a mistake given that PBO2Tunner is for Zen3 cpus, and you mentioned Zen4.
- fefe23 2y agoSo on one hand they are saying it's voltage (i.e. something external, not their fault, bad mainboard manufacturers!). On the other hand they are saying they will fix it in microcode. How is that even possible? Are they saying that their CPUs are signaling the mainboards to give them too much voltage? Can someone make sense of this? It reminds me of Steve Jobs' You Are Holding It Wrong moment.
- cqqxo4zV46cp 2y agoThe “you’re holding it wrong!”angle is all your take. They don’t make that claim.
- k12sosse 2y ago"OK, great, let’s give everybody a case" lives on
- ls612 2y agoThe claim seems to be that the microcode on the CPU is in certain circumstances requesting the wrong (presumably too high) voltage from the motherboard. If that is the case fixing the microcode will solve the issue going forward but won’t help people whose chips have already been damaged by excessive voltage.
- pitaj 2y ago> Are they saying that their CPUs are signaling the mainboards to give them too much voltage? Yes that's exactly what they said.
- TazeTSchnitzel 2y agoAfter watching https://youtube.com/watch?v=gTeubeCIwRw https://youtube.com/watch?v=gTeubeCIwRw and some related content, I personally don't believe it's an issue fixable with microcode. I guess we'll see.
- jpk 2y agoBecause HN doesn't provide link previews, I'd recommend adding some information about the content to your comment. Otherwise we have to click through to YouTube for the comment to make any sense. That said, the video is the GamersNexus one where they talk about an unverified claim that this is a fabrication process issue caused by oxidation between atomic deposition layers. If that's the case, then yeah, microcode can only do so much. But like Steve says in the video, the oxidation theory has yet to be proven and they're just reporting what they have so far ahead of the Zen 5 reviews coming soon.
- mjevans 2y agoHopefully Intel ships them, and allows them to, test and publish benchmarks with the current pre-release microcode revision for review comparison.
- mananaysiempre 2y agoGN mentioned shipping a few samples to a lab (number dependent on the price quote from said lab), so I hope we’ll have some closure regarding this hypothesis.
- acrispino 2y agoAn Intel employee is posting on reddit: https://www.reddit.com/r/intel/comments/1e9mf04/intel_core_13th14th_gen_desktop_processors/ https://www.reddit.com/r/intel/comments/1e9mf04/intel_core_1... A recent YouTube video by GamersNexus speculated the cause of instability might be a manufacturing issue. The employee's response follows. Questions about manufacturing or Via Oxidation as reported by Tech outlets: Short answer: We can confirm there was a via Oxidation manufacturing issue (addressed back in 2023) but it is not related to the instability issue. Long answer: We can confirm that the via Oxidation manufacturing issue affected some early Intel Core 13th Gen desktop processors. However, the issue was root caused and addressed with manufacturing improvements and screens in 2023. We have also looked at it from the instability reports on Intel Core 13th Gen desktop processors and the analysis to-date has determined that only a small number of instability reports can be connected to the manufacturing issue. For the Instability issue, we are delivering a microcode patch which addresses exposure to elevated voltages which is a key element of the Instability issue. We are currently validating the microcode patch to ensure the instability issues for 13th/14th Gen are addressed
- hsbauauvhabzb 2y agoSo they were producing defective CPUs, identified & addressed the issue but didn’t issue a recall, defect notice or public statement relating to the issue? Good to know.
- thelastparadise 2y agoDude's gonna be canned so hard.
- Dylan16807 2y agoIt sounds like their analysis is that the oxidation issue is comfortably below the level of "defective". No product will ever be perfect. You don't need to do a recall for a sufficiently rare problem. And in case anyone skims, I will be extra clear, this is based on the claim that the oxidation is separate from the real problem here.
- PedroBatista 2y agoGood for Intel to finally "figure it out" but I'm not 100% sure microcode is 100% of the problem. As in everything complex enough, the "problem" can actually be many compounded problems, MB vendors "special" tune comes to mind. But this is already a mess very hard to clean since I feel many of these CPUs will die in an year or 2 because of these problems today but by then nobody will remember this and an RMA will be "difficult" to say the least.
- johnklos 2y agoYou're right - at least partly. If the issue is that Intel was too aggressive with voltages, they can use microcode updates as 1) an excuse to rejigger the power levels and voltages the BIOS uses as part of the update, and 2) they can have the processor itself be more conservative with the voltages and clocking it calculates itself. Anything Intel announces, in my experience, is half true, so I'm interested to see what's actually true and what Intel will just forget to mention or will outright hide.
- nubinetwork 2y agoThey already tried bios updates when they pushed out the "intel defaults" a couple months ago...
- wmf 2y agoFirmware and microcode aren't the same thing.
- jeffbee 2y agoVery true and that's why it is odd that microcode has been mentioned here. Surely they mean PCU software (Pcode), or code for whatever they are calling the PCU these days.
- wmf 2y agoI assume Intel's "microcode" updates include the PCU code, maybe some ME code, and whatever other little cores are hiding in the chip.
- jeffbee 2y agoWell, do they? The operating system can provide microcode updates to a running CPU. Can the operating system patch the PCU, too? When I look at a "BIOS update" it usually seems to include UEFI, peripheral option ROMs, ME updates, and microcode. So if the PCU is getting patched I would think of it as a BIOS update. I think the ergonomics will be indistinguishable for end users.
- nicman23 2y agofirmware can include microcode though
- tedunangst 2y agoExcept they didn't. https://www.pcworld.com/article/2326812/intel-is-not-recommending-baseline-power-profiles-to-fix-crashing-cpus.html https://www.pcworld.com/article/2326812/intel-is-not-recomme...
- ChrisArchitect 2y ago(updated from other post about mobile crashes) Related: Complaints about crashing 13th,14th Gen Intel CPUs now have data to back them up https://news.ycombinator.com/item?id=40962736 https://news.ycombinator.com/item?id=40962736 Intel is selling defective 13-14th Gen CPUs https://news.ycombinator.com/item?id=40946644 https://news.ycombinator.com/item?id=40946644 Intel's woes with Core i9 CPUs crashing look worse than we thought https://news.ycombinator.com/item?id=40954500 https://news.ycombinator.com/item?id=40954500 Warframe devs report 80% of game crashes happen on Intel's Core i9 chips https://news.ycombinator.com/item?id=40961637 https://news.ycombinator.com/item?id=40961637
- tedunangst 2y agoNot a dupe.
- silisili 2y agoThat one is mobile, this one is desktop, which they claim are different causes.
- HeliumHydride 2y agohttps://scholar.harvard.edu/files/mickens/files/theslowwinter.pdf https://scholar.harvard.edu/files/mickens/files/theslowwinte... "Unfortunately for John, the branches made a pact with Satan and quantum mechanics [...] In exchange for their last remaining bits of entropy, the branches cast evil spells on future genera- tions of processors. Those evil spells had names like “scaling- induced voltage leaks” and “increasing levels of waste heat” [...] the branches, those vanquished foes from long ago, would have the last laugh." "John was terrified by the collapse of the parallelism bubble, and he quickly discarded his plans for a 743-core processor that was dubbed The Hydra of Destiny and whose abstract Platonic ideal was briefly the third-best chess player in Gary, Indiana. Clutching a bottle of whiskey in one hand and a shot- gun in the other, John scoured the research literature for ideas that might save his dreams of infinite scaling. He discovered several papers that described software-assisted hardware recovery. The basic idea was simple: if hardware suffers more transient failures as it gets smaller, why not allow software to detect erroneous computations and re-execute them? This idea seemed promising until John realized THAT IT WAS THE WORST IDEA EVER. Modern software barely works when the hardware is correct, so relying on software to correct hardware errors is like asking Godzilla to prevent Mega-Godzilla from terrorizing Japan. THIS DOES NOT LEAD TO RISING PROP- ERTY VALUES IN TOKYO. It’s better to stop scaling your transistors and avoid playing with monsters in the first place, instead of devising an elaborate series of monster checks- and-balances and then hoping that the monsters don’t do what monsters are always going to do because if they didn’t do those things, they’d be called dandelions or puppy hugs."
- mattnewton 2y agoI haven't read this piece before but I just knew it was going to be written by Mickens about halfway through your comment.
- throwup238 2y agoThe "mickens" in the URL on the first line was a dead giveaway :-)
- yieldcrv 2y ago
- tpurves 2y agoI think it's telling that they are delaying the microcode patch until after all the reviewers publish their Zen5 reviews and the comparisons of those chips against current Raptorlake performance.
- zenonu 2y agoWhy even publish a comparison? Raptor Lake processors aren't a functioning product to benchmark against.
- tankenmate 2y agoBecause if publishers don't publish then they don't make money.
- AnthonyMouse 2y agoBecause the benchmarks will still exist on the sites after the microcode is released and a lot of the sites won't bother to go back and update them with the accurate performance level.
- Night_Thastus 2y ago"Elevated operating voltage" my foot. We've already seen examples of this happening on non-OC'd server-style motherboards that perfectly adhere to the intel spec. This isn't like ASUS going 'hur dur 20% more voltage' and frying chips. If that's all it was it would be obvious. Lowering voltage may help mitigate the problem, but it sure as shit isn't the cause.
- dwattttt 2y agoThey also admit a microcode algorithm produces incorrect requests for voltages, it doesn't sound like they're trying to shift the blame; ASUS doesn't write that microcode
- sirn 2y agoIt's worth noting that W680 boards are not a server board, they're a workstation board, and often times they're overclockable (or even overclocked by default). Wendell actually showed the other day that the ASUS W680 board was feeding 253W into a 35W (106W boost) 13700T CPU by default[1]. Supermicro and ASRock Rack do sell W680 as a server (because it took Intel a really long time to release C266), but while they're strictly to the spec, some boards are really not meant for K CPUs. For example, the Supermicro MBI-311A-1T2N is only certified for a non-TVB E/T CPUs, and trying to run the K CPU on these can result in the board plumbing 1.55V into the CPU during the single core load (where 1.4V would already be on the higher side)[2]. In this particular case, the "non-OC'd server-style motherboard" doesn't really mean anything (even more so in the context of this announcement). [1]: https://x.com/tekwendell/status/1814329015773086069 https://x.com/tekwendell/status/1814329015773086069 [2]: https://x.com/Buildzoid1/status/1814520745810100666 https://x.com/Buildzoid1/status/1814520745810100666
- paulmd 2y agoSpecifically I think the concerns are around idle voltage and overshoot at this point, which is indeed something configured by OEMs. edit: BZ just put out a video talking about running Minecraft servers destroying CPUs reliably, topping out at 83C, normally in the 50s, running 3600 speeds. Which is a clear issue with low-thread loads. https://m.youtube.com/watch?v=yYfBxmBfq7k https://m.youtube.com/watch?v=yYfBxmBfq7k
- userbinator 2y agoReminds me of Sudden Northwood Death Syndrome, 2002. Looks like history may be repeating itself, or at least rhyming somewhat. Back then, CPUs ran on fixed voltages and frequencies and only overclockers discovered the limits. Even then, it was rare to find reports of CPUs killed via overvolting, unless it was to an extreme extent --- thermal throttling, instability, and shutdown (THERMTRIP) seemed to occur before actual damage, preventing the latter from happening. Now, with CPU manufacturers attempting to squeeze all the performance they can, they are essentially doing this overclocking/overvolting automatically and dynamically in firmware (microcode), and it's not surprising that some bug or (deliberate?) ignorance that overlooked reliability may have pushed things too far. Intel may have been more conservative with the absolute maximum voltages until recently, and of course small process sizes with higher potential for electromigration are a source of increased fragility. Also anecdotal, but I have an 8th-gen mobile CPU that has been running hard against the thermal limits (100C) 24/7 for over 5 years (stock voltage, but with power limits all unlocked), and it is still 100% stable. This and other stories of CPUs in use for many years with clogged or even detached heatsinks seem to contribute to the evidence that high voltage is what kills CPUs, and neither heat nor frequency. Edit: I just looked up the VCore maximum for the 13th/14th processors - the datasheet says 1.72V! That is far more than I expected for a 10nm process. For comparison, a 1st-gen i7 (45nm) was specified at 1.55V absolute maximum, and in the 32nm version they reduced that to 1.4V; then for the 22nm version it went up slightly to 1.52V.
- slaymaker1907 2y agoInteresting, I hadn’t heard about the Pentium overlocking issues. My theory on the current issue that running chips for long periods of time at 100C is not good for chip longevity, but voltages could also be an issue. I came up with this theory last summer when I built my rig with a 13900k, though I was doing it with the intention of trying to set things up so the CPU could last 10 years. Anecdotally, my CPU has been a champ and I haven’t noticed any stability issues despite doing both a lot of gaming and a lot of compiling on it. I lost a bit of performance but not much setting a power limit of 150W.
- cyanydeez 2y agoI believe the first round of Intel excuses here blamed the motherboard manufacturers for trying to "auto" overclock these CPUs.
- whalesalad 2y agoIf I didn’t just recently invest in 128gb of DDR4 I’d jump ship to AMD/AM5. My 13900k has been (knock on wood) solid though - with 24/7 uptime since July 2023.
- thangngoc89 2y agoI guess you’re lucky. I own 2 machines for small scale CNN training, one 13900k and one 14900k. I have to throttle the CPU performances to 90% for stable running. This cost me about 1 hour / 100 hours of training.
- whalesalad 2y agoAre you using any motherboard overclocking stuff? A lot of mobo’s are pushing these chips pretty hard right out of the box. I have mine at a factory setting that Intel would suggest, not the asus multi core enhancement crap. noctua dh15 cooler. It’s really been a stable setup.
- thangngoc89 2y agoI didn’t setup anything in BIOS. But my motherboard are from asus. I will look into this. Thanks for your suggestion.
- Dunati 2y agoMy 13900k has definitely degraded over time. I was running bus defaults for everything and the pc was fine for several months. When I started getting crashes it took me a long time to diagnose it as a CPU problem. Changing the mobo vdroop setting made the problem go away for a while, but it came back. I then got it stable again by dropping the core multipliers down to 54x, but then a couple months later I had to drop to 53x. I just got an rma replacement and it had made it 12 hours without issue.
- EricE 2y agoMake sure you update the BIOS and enable the Intel Baseline Profile - cleared up crashing issues I was having with my i914k https://www.pcgamer.com/hardware/motherboards/asus-adds-intel-baseline-profile-to-the-latest-bios-files-for-better-stability-but-the-tdp-is-still-higher-than-intels-stock-value/ https://www.pcgamer.com/hardware/motherboards/asus-adds-inte...
- phire 2y agoI find it hard to believe that it actually is a microcode issue. Mostly because Intel has way too much motivation to pass it off as a microcode issue, as they can fix a microcode issue for free, by pushing out a patch. If it's an actual hardware issue, then Intel will be forced to actually recall all the faulty CPUs, which could cost them billions. The other reason, is that it took them way too long to give details. If it's as simple as a buggy microcode requesting an out-of-spec voltage from the motherboard, they should have been able to diagnose the problem extremely quickly and fix it in just a few weeks. They would have detected the issue as soon as they put voltage logging on the motherboard's VRM. And according to some sources, Intel have apparently been shipping non-faulty CPUs for months now (since April, from memory), and those don't have an updated microcode. This long delay and silence feels like they spent months of R&D trying to create a workaround, create a new voltage spec to provide the lowest voltage possible. Low enough to work around a hardware fault on as many units as possible, without too large of a performance regression, or creating new errors on other CPUs because of undervolting. I suspect that this microcode update will only "fix" the crashes for some CPUs. My prediction is that in another month Intel will claim there are actually two completely independent issues, and reluctantly issue a recall for anything not fixed by the microcode.
- worthless-trash 2y agoI believe that the waters may be muddied enough that they wont have to do a full recall and only if you 'provide evidence' the system is still crashing.
- RedShift1 2y agoAs I understand it, there are multiple voltages inside the CPU, so just monitoring the motherboard VRM won't cut it. That said I too am very skeptical. I just issued a moratorium on the purchase of anything Intel 13th/14th gen in our company and waiting for some actual proof that the issue is fully resolved.
- phire 2y agoIt's complicated. On Raptor lake, there are a few integrated voltage regulators to which provide new voltages for specialised uses (like the E core's L2 cache, parts of DDR memory IO, PCI-E IO), but the current draw on those regulators is pretty low. The bulk of the power comes directly from motherboard VRMs on one of several rails with no internal regulation. Most of the power draw is grouped onto just two rails, VccGT for the GPU, and VccCore (also known as VccIA in other generations) which powers all the P-cores, all the E-cores and, the ring bus and the last-level cache. Which means all cores share the same voltage, and it's trivial to monitor externally. I guess it's possible the bug could be with only of the integrated voltage regulators, but those seem to only power various IO devices, and I struggle to see how they could trigger this type of instability.
- xyst 2y agoWonder what Linus has to say on this. Dude knows how to rip into crappy Intel products
- weberer 2y agoTorvalds or the Youtube guy?
- happosai 2y agoYes
- aruametello 2y agoI can imagine both will bash intel a bit. "Linus Tech Tips" for the gaming crowd situation (loss of "paid for" premium performance) and Torvalds for the hardware vendor lack of transparency with the community.
- deleted 2y ago[deleted]
- salamo 2y agoIs there any info on how to diagnose this problem? Having just put together a computer with the 14900KF, I really don't want to swap it out if not necessary.
- sudosysgen 2y agoRunning a full memtest overnight and a day of Prime95 with validation is the traditional way of sussing out instability.
- paulmd 2y agoit’s also a terrible stability test these days for the same reasons Wendell talks about with cinebench in his video with Ian (and Ian agrees too). Doesn’t work like 90% of the chip - it’s purely a cache/avx benchmark. You can have a completely unstable frontend and it’ll just work fine because prime95 fits in icache and doesn’t need the decoder, and it’s just vector op, vector op, vector op forever. You can have a system that’s 24/7 prime95 stable that crashes as soon as you exit out, because it tests so very little of it. That’s actually not uncommon due to the changes in frequency state that happen once the chip idles down… and it’s been this way for more than a decade, speedstep used to be one of the things overclockers would turn off because it posed so many problems vs just a stable constant frequency load.
- paulmd 2y agoPrime95 also completely ignores the high-clock domain btw so it can also be completely prime95 stable yet fail completely on a routine task that boosts up! So it’s technically not even a full test of core stability either.
- deleted 2y ago[deleted]
- J_Shelby_J 2y agoOCCP burn in test with AVX and XMP disabled. Tbh, XMP is probably the cause of most modern crashes on gaming rigs. It does not guarantee stability. After finding a stable cpu frequency, enable xmp and roll back the memory frequency until you have no errors in occp. The whole thing can be done in 20 minutes and your machine will have 24/7/365 uptime.
- Havoc 2y ago> Intel is delivering a microcode patch which addresses the root cause of exposure to elevated voltages. That’s great news for intel. If that’s correct. If not that’ll be a PR bloodbath
- eigenform 2y agoby "microcode" i assume they meant "pcode" for the PCU? (but they decided not to make that distinction here for whatever reason?)
- ChoGGi 2y agoHmm, mid August is after the new Ryzens are out, I wonder how bad of a performance hit this microcode update will bring? And will it actually fix the issue? https://www.youtube.com/watch?v=QzHcrbT5D_Y https://www.youtube.com/watch?v=QzHcrbT5D_Y
- uticus 2y agoDumb question: let’s say I am in charge of procurement for a significant amount of machines, do I not have the option of ordering machines from three generations back? Are older (proven reliable) processors just not available because they’re no longer made, like my 1989 Camry?
- wmf 2y agoYeah, 12th gen is probably still available.
- cdchn 2y agoI built a system last fall with an i9-13900K and have been having the weirdest crashing problems with certain games that I never had problems with before. NEVER been able to track it down, no thermal issues, no overclocking, all updated drivers and BIOS. Maybe this is finally the answer I've been looking for.
- EricE 2y agoIt was for me. Check for BIOS updates - most motherboard vendors have them. Look for and enable something labeled Intel Baseline Profile and then check. That cured it for me. For Asus: https://www.pcgamer.com/hardware/motherboards/asus-adds-intel-baseline-profile-to-the-latest-bios-files-for-better-stability-but-the-tdp-is-still-higher-than-intels-stock-value/ https://www.pcgamer.com/hardware/motherboards/asus-adds-inte...
- cdchn 2y agoI'll try that, thanks. Although the current cohort of games I play seems more stable now. If I ever go back to EVE Online then it'd be more of an issue - that thing crashed constantly.