6 ms·
This seems like deflection to me. There’s no reproducer, no technical info, no objective facts besides what is basically a meta study that lumps their code qual
by ComputerGuru 2y ago
This seems like deflection to me. There’s no reproducer, no technical info, no objective facts besides what is basically a meta study that lumps their code quality with intel’s hardware correctness as one metric.
If it’s as they say, switching to AMD shouldn’t lead to 100x fewer crashes, it should lead to no crashes.
Lumping both 13th and 14th gen hardware together makes it harder to take the claims as easily at face value. 13th gen have been around long enough that you’d think they’d be able to dig up a kernel mailing list bug report or some other corroboration of their accusation.
I clicked expecting basically an undocumented errata with specific, 100% reproducible instructions and a deep dive into the internal architecture changes that could have caused this. I’m tempted to flag the submission, honestly. Claims of such magnitude demand at least some baseline evidence.
- jandrese 2y agoThe strangest part of the whole thing is how they claim that the CPUs will work fine initially, but then degrade over time in a deterministic way. How is nobody else seeing this? Are these guys executing code paths that are normally not used? It sounds a bit like their code is buggy in a strange way that doesn't cause problems on AMD but does on later generation Intel, but that wouldn't explain why code would work initially but fail later unless maybe the problem was brought about by a microcode update from Intel? I fully agree that demanding Intel recall all of the CPUs at this point is premature. They've done nowhere near enough research to pin the problem on the processor and not their code.
- jchw 2y ago> How is nobody else seeing this? Well, ... they are. This has been being discussed for a while. It came up a bit ago, many people were seeing it in games with a weird error about low VRAM which apparently wound up being linked to Intel's 13th and 14th gen processors. Exactly why is anyone's guess. https://www.tomshardware.com/pc-components/cpus/nvidia-blames-intel-for-gpu-vram-errors-tells-geforce-gamers-experiencing-13th-or-14th-gen-cpu-instability-to-contact-intel-support https://www.tomshardware.com/pc-components/cpus/nvidia-blame... Level1Techs is relatively well-trusted, and they seem to have some contacts that can also corroborate these issues. https://www.youtube.com/watch?v=QzHcrbT5D_Y https://www.youtube.com/watch?v=QzHcrbT5D_Y
- willcipriano 2y agoOne idea: Get 500 watt power supply, it powers new chip fine with it's unlimited configuration. 500 watt power supply becomes less efficient over time, now it's not quite 500 watts, parts of the chip brown out. Undefined behavior.
- jchw 2y agoHmmm, I think this is unlikely. For one thing, most people's power supplies are quite wildly overspecced in gaming rigs since its genuinely pretty hard to find good small ATX PSUs these days. For another, even crummy low-quality PSUs tend to mostly work pretty well for a reasonably long time in most cases, so it'd be weird to see this so much in high-end gaming rigs with brand new CPUs. I also don't know exactly what failure mode you are suggesting with this (voltage rail out of spec?) but I think unless it starts feeding too much voltage into the CPU it is probably unlikely to be able to cause damage. (Instability, sure.) It also doesn't explain why Intel's 12th generation CPUs are still running strong, along with most of AMD's lineup, by comparison. Don't get me wrong, every SKU of CPU has its issues, but all of the major issues with other CPU lines we have explanations for; the issues with Intel Raptor Lake are thusfar unexplained. Intel has blamed motherboard manufacturers quite extensively, but so far the evidence of this is pretty uncompelling, because the failures can apparently also be seen in abnormally high numbers on motherboards that never had unreasonable configurations.
- eqvinox 2y agoHaving built a few power management circuits myself: this is not how PSUs degrade, at least not in the first 5 years, and generally not "ever" in most cases. The PSU (notably in this case the M/B VRM, not the ATX PSU) can't really trigger brownouts. 1.05V is 1.05V, and if asked to do so, the PSU will deliver 1.05V. It will just get hotter while doing so over time (and specifically: capacitors degrading.) The voltage references are quite stable over time, and so are the window comparators that ensure "1.05V" is between, say, 1.035V and 1.065V. (Voltage references have essentially continuous engineering history back to the 50s≈60s, it's a "solved" problem.)
- 2y ago
- deleted 2y ago[deleted]
- eqvinox 2y ago> Are these guys executing code paths that are normally not used? Yes: games are notorious for having poor multithreading and hitting one core with very high load. And unlike servers where you're rarely the only workload, the rest of the CPU will be rather idle on a desktop/gaming system. This pushes the CPU to a much more "imbalanced" mode (putting all the power and heat into a very small area) than is common elsewhere.
- hdhshdhshdjd 2y agoI have these issues on a dedicated Postgres server running Linux and literally nothing else. It’s not just games, it’s just a lot of gamers buy these chips.
- wmf 2y agoSilicon aging is a thing, especially at high voltage.
- forrestthewoods 2y ago> If it’s as they say, switching to AMD shouldn’t lead to 100x fewer crashes, it should lead to no crashes. No? There’s always multiple sources of crashes. Switching from Intel may eliminate 100% of one crash source, but not 100% of all crashes. One of my favorite stories is that ArenaNet once got Guild Wars servers so stable that any crash could be reliably attributed to hardware failure. Usually memory IIRC.
- ComputerGuru 2y agoThat’s kind of my point. They’re clearly running a buggy codebase to begin with that I wouldn’t trust a crash to mean hardware failure.
- eqvinox 2y agoCrashes caused by hardware failures look very different from crashes caused by a buggy codebase. 80% of code bugs will have very unisonous crash reports. The other 20% of code bugs — memory corruption and race conditions, notably — have more random crash reports, but are still identifiable as such. Random hardware failures generally cause crash reports that look like someone put your bits in a mixer and made a smoothie. (NB: random hardware failures like we're talking about here. CPU design/logic erratas can and do look like code bugs; but that's not what's being discussed here.)
- db48x 2y agoAll computers have a baseline crash rate due to environmental factors such as overheating, cosmic rays, low bit error rates of modern hardware, etc, etc. My 5950X has never crashed, but I bet over a large enough population the crash rate would be measurable and consistent from game to game.
- sysadmin420 2y agoI used to be a hard core Intel guy through my career in IT and systems, but man my AM4 has been a super great system, the 5600x was stable, then upgraded to the 5950x to extend the system and it made me a true AMD fanboy overnight.
- thadt 2y agoIf they were the only ones noticing, I might agree. But there are some groups I trust to know their hardware inside and out, and one of those is RadGames[1]. [1] https://www.radgametools.com/oodleintel.htm https://www.radgametools.com/oodleintel.htm
- ndiddy 2y agoPrevious discussion: https://news.ycombinator.com/item?id=39478551 https://news.ycombinator.com/item?id=39478551
- deleted 2y ago[deleted]
- the8472 2y ago> If it’s as they say, switching to AMD shouldn’t lead to 100x fewer crashes, it should lead to no crashes. No, this does not follow, at all. Few things in life are monocausal and systems (here the aggregate of hardware and software) tend to have more than one bug. If one bug dominates and you remove it then of course you're still left with residuals. It doesn't have to be intel's fault though. It could be mobo vendors defaulting to unsafe settings.
- sliken 2y ago> It doesn't have to be intel's fault though. It could be mobo vendors defaulting to unsafe settings Sure. However workstation and server class motherboards have a much lower risk of not following Intel recommendations on voltages and clock speeds. The error rate on the workstation boards are pretty high, on the order of a failure per week for 50% of workstations they were collecting telemetry from.
- tonyarkles 2y agoYeah, I'm super curious here. It's a very bold claim. It's possible that they've got some kind of data race/memory barrier/other threading issue that's more likely to get tickled by Intel processors than AMD processors. Hell, it could even be a compiler bug that generates incorrect assembly that trips more often on Intel processors.
- eqvinox 2y agoThere is no such thing as "incorrect assembly", modern CPUs don't have undefined instructions like the 6502 did. Either it gets executed deterministically, or you get an #UD exception.
- cpgxiii 2y agoThat's not quite true, though. While there are no singular undefined instructions, there are sequences of instructions that can result in undesired/undefined behavior. Read through the published errata for modern CPUs, there will be a large number of combinations of instructions/actions that should be avoided entirely or modified to avoid misbehavior.
- eqvinox 2y ago> Read through the published errata for modern CPUs I totally agree, and this entire HN post is about the fact that these Intel CPUs probably have some yet-undiscovered erratas. Best theory at this point is that it's power/thermal related though. > there will be a large number of combinations of instructions/actions that should be avoided As a matter of fact, reading those erratas is part of my job, and no, it is very rare for such erratas to have the effect you're implying. Rare enough that this issue that is seen here would already have been matched against an errata. I don't see anything in Intel's current errata — other than their in-progress responses to these reports — that could explain this. Here's the errata sheet for 13th/14th gen: https://edc.intel.com/content/www/us/en/design/products/platforms/details/raptor-lake-s/13th-generation-core-processor-specification-update/summary-tables-of-changes/ https://edc.intel.com/content/www/us/en/design/products/plat... … do you see anything?
- deleted 2y ago[deleted]
- fourfour3 2y ago13th and 14th gen desktop parts are nearly identical - so lumping them together makes sense. Edit: 14th gen top end parts are basically 13th gen parts with different binning and higher clocks.
- coconut08 2y agothis isn't something that is just coming from them. there has been a lot of buzz from a lot of sources discussing this exact topic. https://www.youtube.com/watch?v=QzHcrbT5D_Y https://www.youtube.com/watch?v=QzHcrbT5D_Y
- jchw 2y ago> Lumping both 13th and 14th gen hardware together makes it harder to take the claims as easily at face value. 13th gen have been around long enough that you’d think they’d be able to dig up a kernel mailing list bug report or some other corroboration of their accusation. This is pretty damn reasonable actually. The 14th gen Intel desktop hardware is for all intents and purposes, essentially the same as the 13th gen hardware. It behaves nearly identically, it benchmarks nearly identically.
- blain 2y agoI agree there should be something more specific but in the resources links they provide there is an epic games instructions page [1] to do msinfo for people getting crashes for these specific CPUs, so its no just them saying it. [1] https://www.epicgames.com/help/en-US/c-Category_Fortnite/c-Fortnite_TechnicalSupport/frequent-crashes-in-fortnite-on-i9-13900k-kf-ks-or-i9-14900k-kf-ks-cpus-a000086852?sessionInvalidated=true https://www.epicgames.com/help/en-US/c-Category_Fortnite/c-F...
- rcarmo 2y agoThere's been a few interesting videos on this, some like Wendell's - https://www.youtube.com/watch?v=QzHcrbT5D_Y https://www.youtube.com/watch?v=QzHcrbT5D_Y He's reached out to folk who gather game telemetry and got some interesting data to play with.
- doctorpangloss 2y agoThe problem is real. Based on Fortnite’s suggested fix in their help, it sounds like there there is some flaw in the Intel hardware that makes it crash later in life when using the shipping voltage, which is maybe its most important parameter. When that is forced to be higher the part gets much hotter and performance per watt declines.
- jeffbee 2y agoYou can't necessarily publish a 100% repro for glitchy CPUs. It's stochastic both for whether it happens and whether it's detectable or silent corruption. My glitchy 14900K fails with illegal instruction, or sometimes machine checks. But who can say what else is happening that wasn't detected? I can definitely tell you that at least my instance of the 14900K cannot complete a `bazel test` of Abseil without crashing.
- maxwell_smart 2y agoThe links that are included shed more light on things. The write-up linked is a little light on details, and I also expected an actual "flaw". What the other links show is a series of anecdotal reports that the k-series intel processors have been found by end-users to become unstable, and that setting more conservative power management settings in the BIOS helps these chips regain stability. The argument is that Intel is selling chips that at their out-of-the-box recommended use patterns can wear out extra quickly, and by default are over-volted. So it's not a "flaw" so much as not having a great engineering safety margin for premature wear.
- paulmd 2y agothe same series of videos that launched this, also confirm that chips like 13700T are degrading (or at least crashing). So it’s not nearly so simple as “too much power/voltage” - this is happening in 35W chips. There are at least 2 causes and failures modes, intel has already confirmed eTVB=off was a problem and could cause degradation due to excessive heat. The other suspect right now is the ring bus degrading, perhaps.
- eqvinox 2y ago> This seems like deflection to me. There’s no reproducer, no technical info, no objective facts besides what is basically a meta study that lumps their code quality with intel’s hardware correctness as one metric. It's just laziness. The RAD game tools statement/article[1] has much more detail. I'm not sure to what degree the people behind the current article verified that they're seeing the same issue, but there is definitely a issue (apparently power/scaling/temperature related). [1] https://www.radgametools.com/oodleintel.htm https://www.radgametools.com/oodleintel.htm
- hdhshdhshdjd 2y agoI’ve had extensive problems with the 13th gen, nearly all of the reporting I’m seeing is what I’ve experienced. Including weird NVME issues that I couldn’t figure out until I saw reports of it being related to these issues. I’m done with Intel.
- sliken 2y ago> Lumping both 13th and 14th gen hardware together makes it harder to take the claims as easily at face value. 13th gen have been around long enough that you’d think they’d be able to dig up a kernel mailing list bug report or some other corroboration of their accusation. First of all the 13th and 14th gen are nearly the same CPUs. Slight differences in clock, TDPs, and core counts, but all using the identical raptor lake cores. Intel's been widely ridiculed for calling them 14th gen when there's no difference in the cores. There have been various posts about these stability problems going back months on numerous forums. Often things like BSOD, not enough VRAM errors, decompression errors, failure to load game assets, which looks like a NVMe error, etc. Part of the reason it's flying under the radar is a combination of factors. It's not all chips, it's not a specific instruction, and the behavior changes over time by getting worse. Sadly if a game or OS crashes nobody is terribly surprised, which can make hardware problems less obvious. From what I can tell there Intel clock/voltage/temp curves were pushed, trying to be competitive with AMD despite a disadvantage Intel has with their 10nm fabs. Not only does this cause errors, but it causes damage, so the error rate increases over time. The best source found seems to be the telemetry built into unreal engine that provides an unbiased report on crashes (but not hangs, BSoD aren't reported). Sure some gamers overclock, have poor airflow, poorly applied headsink goo, etc. But I found the following pretty compelling since it's based on workstation hardware, without overclocking, using the W series workstation chipset not the Z series consumer chipset which allows overclocking: In a test population of more than 210 W680-based systems, 47.1% of these systems experience at least one incident of instability over a 168 hour test window. This distribution is the same to within 0.4% between Asus brand W680 and Supermicro W680 based boards. For people running these boards on the server side they actually are charging more because of support issues related to CPU replacements. Certainly a failure of 50% per week per node is crazy high and unacceptable on workstation class hardware.
- jovial_cavalier 2y agoDo you hold INTC by any chance? >If it’s as they say, switching to AMD shouldn’t lead to 100x fewer crashes, it should lead to no crashes. As many others have pointed out, this is baloney. Computers are not perfect machines and even if they were, gathering aggregate user data is also imperfect. > I clicked expecting basically an undocumented errata with specific, 100% reproducible instructions and a deep dive into the internal architecture changes that could have caused this. I don't know why you would expect that. You know what an intermittent bug is, this feels like disingenuous reasoning. Furthermore, there are many other people who are noticing this, documenting it, and attempting to report on it.