20 ms·
Intel Processor Instability Causing Oodle Decompression Failures
- eqvinox 3y agoThe article is a bit unclear on whether this happens with standard/default settings, tough that's probably because they don't know themselves. The workarounds changing things from "Auto" to "disabled" or even increasing voltage settings certainly seems like it also applies with defaults, and isn't some overclocking/tuning side effect. If that is the case… ouch.
- yetihehe 3y agoIt seems like it happens only on some select cpu specimens (apparently works after replacing cpu with another one of the same model), so probably a small binning failure?
- eqvinox 3y agoI'm not sure I would call a binning failure "small" — mostly because I can't remember this ever happening before. Binning is a core aspect of managing yields. And it seems that this is breaking for a sufficient number of people to have a game tooling vendor investigate. How many bug reports would it take to get them into action?
- rygorous 3y agoPerson who actually did the investigation here. It took exactly one bug report. RAD/Epic Games Tools is a small B2B company. Oodle has one person working full-time on it, namely me, and I do coding, build/release engineering, docs, tech support, the works. There's no multiple support tiers or anything like that, all issues go straight into my inbox. Oodle Data in particular is a lossless data compression API and many customers use two entry points total, "compress" and "decompress". I get a single-digit number of support requests in any given month, most of which is actually covered in the docs and takes me all of 5 minutes to resolve to the customer's satisfaction. The 3-4 actual bug reports I get in any given year, I will investigate.
- wmf 3y agoThis is the usual silicon lottery. Every chip will work at stock settings. Some will be stable when overclocked and some won't.
- scrlk 3y agoThis sounds like motherboard manufacturers pushing aggressive OOTB performance settings, likely in excess of the Intel spec.
- eqvinox 3y agoThat's… an assumption. At least 3 motherboard vendors are affected, and going by the Gigabyte/MSI workarounds at the end of the article, it looks like things need to be adjusted away from Intel defaults. …it'll need a statement from Intel for some clarity on this…
- scrlk 3y ago> "Intel's default maximum TDP for the 13900K is 253 watts, though it can easily consume 300 watts or more when given a higher power limit. In our testing, manually setting the power limit to 275–300 watts and the amperage limit to 350A, proved to be perfectly stable for our 13900K. That required going into the advanced CPU settings in the BIOS to change the PL1/PL2 limits — called short and long duration power limits in our particular case. The motherboard's default "Auto" power and current limits meanwhile created instability issues — which correspond to a power limit of 4,096 watts and 4,096 amps." [0] The motherboard manufacturers are setting default/auto power and current limits that are way outside of Intel's specs (253 W, 307 A) [1]. [0] https://www.tomshardware.com/pc-components/cpus/is-your-intel-core-i9-13900k-crashing-in-games-your-motherboard-bios-settings-may-be-to-blame-other-high-end-intel-cpus-also-affected https://www.tomshardware.com/pc-components/cpus/is-your-inte... [1] https://www.intel.com/content/www/us/en/content-details/743844/13th-generation-intel-core-and-intel-core-14th-generation-processors-datasheet-volume-1-of-2.html https://www.intel.com/content/www/us/en/content-details/7438... (see pg. 98 and 184, the 13900K/14900K is 8P + 16E 125 W)
- eqvinox 3y agoYour [0] also says: > It's not exactly clear why the 13900K suffers from these instability problems, and how exactly downclocking, lowering the power/current limits, and undervolting prevent further crashes. Clearly, something is going wrong with some CPUs. Are they "defective" or merely not capable of running the out of spec settings used by many motherboards?
- rygorous 3y ago(I'm Oodle maintainer and did most of this investigation.) For the majority of systems "in the wild", I don't know. We had two people with affected machines contact us and consent to do some testing for us, and in both cases the issue still reproduced after resetting the BIOS settings to defaults.
- yetihehe 3y agoTL;DR: some motherboards by default overclock too much on some intel processors, causing instability.
- worewood 3y agoIt's MCE all over again
- eqvinox 3y agoThat's not my interpretation. Cf the following: > For MSI: > Solution A): In BIOS, select "OC", select "CPU Core Voltage Mode", select "Offset Mode", select "+(By PWM)", adjust the voltage until the system is stable, recommend not to exceed 0.025V for a single increase. This really sounds like the Intel defaults are broken too.
- yetihehe 3y agoYes, that or insufficient quality checks, meaning some units will fail, some will work. Apparently it was only a subset of each model failing.
- haunter 3y ago>default overclock too much Per the article MSI literally suggests to OC the CPU to fix the problem
- xcv123 3y agoArticle recommends disabling overclocking. The MSI recommendation is only to increase voltage.
- londons_explore 3y agoIf Oodle has control of this code, the logical thing for them to do, when they detect a decompression checksum failure, is to re-decompress the same data (perhaps single threaded rather then multithreaded). Sure, the user has a broken CPU, but if you can work around it and still let the user play their games, you should.
- yetihehe 3y agoYes, but then the processor will fail at another task during the game and corrupt some other memory. The only solution for unstable processor is to make it stable or replace.
- mike_hock 3y agoYes. Props to Oodle for not passing on the hot potato but trying to get the root cause fixed. This hack would have been the easy way out for them so their product doesn't get blamed.
- lifthrasiir 3y agoAs noted in the linked page, this issue would affect any heavy use of CPU. Oodle happened to be optimized well to hit this issue earlier than most other applications, but nothing can't be really trusted at that point. There is a reason that they recommend to disable overclocking if possible, because such issue is in general linked to the instability due to excessive overclocking.
- barrkel 3y agoRetry is not an unusual response to unreliable hardware, and all hardware is ultimately unreliable. Software running at scale in the cloud is written to be resilient to errors of this nature; jobs are cattle, if jobs get stuck or fail they are retried, and sometimes duplicate jobs are started concurrently to finish the entire batch earlier.
- lifthrasiir 3y ago
- lifthrasiir 3y agoThis page doesn't seem to be linked from any other public page, so I think it was a response to unwanted complaints from users who tried to track the "oodle" thing in the error log---like SQLite back in 2006 [1]. [1] https://news.ycombinator.com/item?id=36302805 https://news.ycombinator.com/item?id=36302805
- eqvinox 3y agoIt's linked from https://www.radgametools.com/tech.htm https://www.radgametools.com/tech.htm (click "support" at the top, look next to "Oodle" logo → "Note: If you are having trouble with an Intel 13900K or 14900K CPU, please [[read this page]].")
- lifthrasiir 3y agoOoh, thank you! I looked so long at the Oodle section and skimmed other sections as well (even searched for the `oodleintel.htm` link in their source codes), but somehow missed that...
- atesti 3y agoThere are some pages that are not linked, wondered what happened to these products https://www.radgametools.com/granny.html https://www.radgametools.com/granny.html https://www.radgametools.com/iggy.htm https://www.radgametools.com/iggy.htm https://www.radgametools.com/milesperf.htm https://www.radgametools.com/milesperf.htm
- rygorous 3y agoGranny, Iggy and Miles are all discontinued as stand-alone products. We're still providing support to existing customers but not selling any new licenses.
- pixelpoet 3y agoWhile we've got you, any chance you'll attend another demoparty here in Germany? :) Big thanks for your awesome blog, learnt much from it over the years.
- imdsm 3y ago1994 all over again!
- cwillu 3y agoFDIV was a bug in the logical design, this is a over-aggressive clock tuning, I fail to see any resemblance whatsoever beyond intel being inside.
- crest 3y agoIt's no longer black and white like the FDIV bug, but if the default configuration leads to data corruption in heavy SIMD workloads... sure you can reduce clock speed or increase voltage until it works, but unless the mainboards violate the specs this is at least partly an Intel CPU flaw leading to data corruption.
- mhio 3y agoThis sounds familiar... ye olde pentium III 1.13 GHz https://www.tomshardware.com/reviews/intel-admits-problems-pentium-iii-1,235-3.html https://www.tomshardware.com/reviews/intel-admits-problems-p...
- ManuelKiessling 3y agoUnrelated to the actual topic, but kudos to the Tom’s Hardware site that they serve a 24 years old web posting flawlessly.
- franzb 3y agoReminds me of this saga I went through as an early adopter of AMD Threadripper 3970X: https://forum.level1techs.com/t/amd-threadripper-3970x-under-heavy-avx2-load-defective-design-no-but-there-is-an-issue/153883 https://forum.level1techs.com/t/amd-threadripper-3970x-under... HN discussion: https://news.ycombinator.com/item?id=22382946 https://news.ycombinator.com/item?id=22382946 Ended up investigating the issue with AMD for several months, was generously compensated by AMD for all the troubles (sending motherboards and CPUs back and forth, a real PITA), but the outcome is that I've been running since then with a custom BIOS image provided by AMD. I think at the end the fault was on Gigabyte's side.
- rwmj 3y agoReminded me of the Intel Skylake bug found by the OCaml compiler developers: https://tech.ahrefs.com/skylake-bug-a-detective-story-ab1ad2beddcd https://tech.ahrefs.com/skylake-bug-a-detective-story-ab1ad2...
- rkagerer 3y agoHoly cow I had no idea CPU vendors would do this for you.
- devmor 3y agoWhen you’re not only helping them debug their own hardware but are also spending money on their ridiculously overpriced HEDT platform, it probably makes them want to keep you happy.
- zitterbewegung 3y agoThat is true and also lots of people use OCaml
- zare_st 3y agoSupermicro gave us same type of assistance. Then new feature of bifurcation did not work correctly. Without it, enterprise telecommunications peripheral that costs 10x more than 4 socket Xeon motherboard can't run at nominal speed, and it was ran on real lines, not test data. They sent us custom BIOSes until it got stabilized and said they'll put the patch in the following BIOS releases. The thing is neither Intel nor AMD nor Supermicro can test edge cases at max usage in niche environments without paying money, but they would really love to claim with backup they can be integrated for such solutions. If Intel wants to test stuff in space for free they have to cooperate with NASA; the alternative is in-house launch.
- vdaea 3y agoI have a 13900K and I am affected. Out of the box BIOS settings cause my CPU to fail Prime95, and it's always the same CPU cores failing. Lowering the power limit slightly will make it stable. I intended to better refrigerate the CPU and change the power limit back to the default and if the problems continued I would RMA the CPU, but now I'm not so sure that the BIOS is not pushing it beyond the operating limits.
- mips_r4300i 3y agoCan I ask what mobo vendor? Do you know what power limits the BIOS was targeting that caused the error? On my Asus/14900k, it was uncapping PL1/2 and I saw absurd temps and power every time anything even touched the CPU. I programmed PL1/2 to 125/253w per Intel ARK and everything normalized. I did not do Prime95 at the insane default power limits but I suspect similar.
- vdaea 3y agoMOBO is MSI, setting was "cpu cooler tuning" which was set at 4096W and I had to change to 253W (the limit according to Intel)
- Havoc 3y agoSame cpus as the unity engine (or was it unreal?) with issues Not a good look but at least it’s fixable with bios tweaks rather than a silicon flaw that’s permanent
- zvmaz 3y agoDoes it mean that a formally verified piece of software like seL4 can still fail because of a potential "bug" in the hardware?
- flumpcakes 3y agoI would assume that _any_ software, formally verified or not, could fail due to a hardware problem. A cosmic ray could flip a bit in a CPU register. The chances of that happening, and that effecting anything in any meaningful way is probably astronomically low. We probably have thousands of hardware failures every day and don't notice them. This is why I think rust in a kernel is probably a bad idea if it doesn't change from the default 'panic on error'.
- hmottestad 3y agoI would assume that software can always fail in the event of a bug in the hardware. That's why systems that are really redundant, for instance flight control computers, have several computers that have to form a consensus of sorts.
- eqvinox 3y agoIt doesn't even need a bug in the hardware; cosmic rays or alpha particles can also cause the same type of issue. For those, making systems redundant is indeed a good solution. For the situation of an actual (consistent) hardware bug, redundancy wouldn't help… the redundant system would have the same bug. Redundancy only helps for random-style issues. (Which, to be fair, the one we're talking about here seems to be.)
- davrosthedalek 3y agoThat's why some redundant systems use alternative implementations for the parallel paths. Less likely that a hardware bug will manifest the same way in all implementations.
- rygorous 3y agoAbsolutely, yes. It can also misbehave without any hardware bugs due to glitching. Rates of incidence of this must be quite low or that would be considered a HW bug, but it's never zero. Run code for enough hours on enough machines collecting stack traces or core dumps on crashes and you will notice that there's a low base rate of failures that make absolutely no sense. (E.g. a null pointer dereference literally right after a successful non-null pointer check 2 instructions above it in the disassembly.) You will also notice that many machines in a big fleet that log such errors do so exactly once and never again, but some reoccur several times and have a noticeably elevated failure rate even though they're running the exact same code as everyone else. This too is normal. These machines are, due to manufacturing variation on the CPU, RAM, or whatever, much glitchier than the baseline. Once you've identified such a machine, you will want to replace it before it causes any persistent data corruption, not just transient crashes or glitches.
- enraf 3y agoI got one of the faulty 13900k, at least in my case I can confirm that the fault appeared using the default settings for pl1/pl2. I was doing reinforcement learning on that system and it was always crashing, I spent quite a bit of time trying to find the problem, swapped the CPU for a 13700kf I was using in another PC, the problem was solved. So I contact Intel to start the RMA process, Intel said that the MSI motherboard I was using doesn't support Linux, I emailed them the official Intel GitHub repo with the microcode that enables the support, they switched agents at that point but I was clear to me at that moment that Intel was trying their best to avoid the RMA, luckily I live in Europe, so I contacted my local consumer protection agency and did the RMA through them, in the meanwhile I saw a good offer for a 7950x + motherboard in an online retailer, bought it and sold in the second market my old motherboard and the RMA 13900k when I got it. Not buying Intel ever again, I was using Intel because they sponsor some projects in DS but damn.
- hopfenspergerj 3y agoI’ve had instability with my 7700k since I bought it, and 16 months of bios updates haven’t helped. Maybe this latest generation of processors just has more trouble than older, simpler designs.
- acdha 3y agoIntel has been struggling with CPU performance for a decade, and has been trying to regain their position in absolute performance and performance/{price,watt} comparisons. I think that means they’re being less conservative than they used to be on the hardware margins and also that their teams are likely demoralized, too.
- smolder 3y agoPossibly. I would start swapping parts around at that point. Different memory, different CPU, or different motherboard. Just 1 more anecdote, but my r7-7700x has been a dream (won the silicon lottery). It runs at the maximum undervolt & RAM at 6000 with no stability problems.
- 123232qe 3y ago[flagged]
- mattgreenrocks 3y agoI can't say I'm surprised by this at all. I bought my 4790k's ASUS TUF board awhile back because I wanted something basic enough and wasn't interested in overclocking or tweaking. The BIOS had other ideas. I had to manually configure a lot more things just to avoid overclocking, including setting RAM timing and going through each BIOS setting to ensure it wasn't overclocking in some way. The "optimal" setting would turn on aggressive changes like playing with bus speed multipliers, etc.
- layer8 3y agoFew people buy a K processor who aren’t interested in overclocking and tweaking. I wouldn’t be surprised if the BIOS of a gaming mainboard sets the “optimal” defaults on that basis, since the gaming market is all about benchmarks.
- phil21 3y agoI'm pretty much the same as OP. I almost always buy the K version of the processor, but never intend to overclock. I just figure I want the theoretical ability to, and the more volume they have on those SKUs the less likely they are to take it away entirely. That or I'm just rewarding shitty corporate product segmentation behavior. I never can quite decide. I do agree over the recent years getting a "boring" higher-end configuration is getting more and more difficult.
- bonton89 3y agoK chips often came with higher default clocks and definitely have better resale value so they're often worth buying even if you don't overclock.
- smolder 3y agoYes, the overclockable chips are better-binned/faster chips even without enabling overclocking. (Unless you're talking about X3D chips, which have most overclocking features turned off due to thermal limitations of stacked cache.)
- IYasha 3y ago© 1991 - 2024 Epic Games Tools LLC Wow. RAD was bought by Epic? I kinda missed that. Feels old. :(
- mobilio 3y agoYup https://www.epicgames.com/site/en-US/news/epic-acquires-rad-game-tools https://www.epicgames.com/site/en-US/news/epic-acquires-rad-...
- tibbydudeza 3y agoWhy I chose a i9 13900 (non K) variant rather- my PC earns me money as a freelance software dev so I can't stand weird issues like this
- svantana 3y agoAs another software dev, I would pay big money for the "worst possible computer" that exhibits all of the glitches and issues that end users see. It's so annoying to get bug reports that I can't reproduce.
- tibbydudeza 3y agoI had my time during my embedded days - did a site visit 1000 km away and discovered no wonder the serial port and scanner/printer is going wonky. No shielding, earth - using the crappiest/cheapest PC they could get instead of using the recommended kit as the sales droid wanted a bigger commission. Said call me when you replaced the h/w - I walked out and went to the airport. They never called me.
- fabianhjr 3y agoIf that was the case why not go for Ryzen + ECC memory?
- tibbydudeza 3y agoGot ECC memory in my server - I am a value for money - my previous kit was i7-6700 system 48GB so I really sweated it until Jetbrains let me know "She canno go more Captain". DDR4/Intel motherboards are cheaper than AM5/DDR5 - also a Ryzen laptop foobarred on my daughter so to me Intel kit was just more stable - no weird XMP issues or overclocking to the nines.
- codexon 3y agoI got a 13900 non-K on a Linux server and it randomly locked up the system after a month.
- op00to 3y agoThis sounds a lot like the behavior I see when I have over locked my processor too far, and try to run AVX heavy workloads! Cranking down the frequency during AVX seems to stabilize things.
- Arech 3y agoHad the same experience overclocking old AMD Phenom II a while ago. Worked flawlessly in all publicly available test software I tried, until I run some custom heavily vectorized code, which eventually required to shave off almost all overclocking :D
- op00to 3y agoThere's a way (at least on my Intel) to tell the processor to clock down a certain number of steps depending on whether AVX is being executed. So, for the majority of stuff that didn't use AVX I let 'er rip, but when AVX is running it clocks down a couple steps. I could use less voltage, and this CPU is fast enough. I think it's a 13900k.
- Arech 3y agoAh, that's an interesting feature of new CPUs, I didn't know about it! Thanks for telling!
- newsclues 3y agoremember when. intel had intel branded reference boards? I would like a comeback please
- mbrumlow 3y agoI recently built a new system with a i9 149kf and a Ausus Formula motherboard. For a VFIO system so I could run windows and play some games. It was a nightmare to get running stable. None is the default settings the motherboard used worked. Games crashed, kernel and emacs compiles failed. End result I had to cap turbo to 5.4ghz on a 6ghz chip, and enable settings that capped max watts and temperature for throttling to 90c. System seems stable now. Can get sustained 5.4ghz without throttling and enjoying games at 120fps with 4k resolution. Even though it is working I do feel a way about not being able to run the system at any of the advertised numbers I paid for.
- doubled112 3y agoWhat I'm not happy about is the marketing around turbo boost. You know how ISPs used to sell "up to X Mbps"? Same idea. Your chip will turbo boost "up to 6.00 GHz". It's basically automated overclocking, and as you learned, sometimes it can't even do it in a stable fashion. Some of those chips will never clock "up to 6.00 GHz" but they didn't lie. "up to"
- deleted 3y ago[deleted]
- wtallis 3y agoIt's particularly bad when they stop telling you what clock speeds are achievable with more than one core active. At best these days you get a "base clock" spec that's very slow that doesn't correspond to any operating mode that occurs in real life. You used to get a table of x GHz for y active cores, but then the core counts got too large and the limits got fuzzier. And laptops have another layer of bullshit, because the theoretical boost clocks the chip is capable of will in practice be limited by the power delivery and cooling provided by that specific machine, and the OEMs never tell you what those limits are. So they'll happily take an extra $200 for another 100MHz that you'll never see for more than a few milliseconds while a different model with a slower-on-paper CPU with better cooling can easily be more than 20% faster.
- 3y ago
- jeffbee 3y agoSimilar experience with an Asus motherboard. With their automatic tuning, instability leading to compiler crashes. Had to manually set the BIOS for sanity. I believe the problems are compounded by the way their SuperIO controls the cooler, because the crashes were associated with temperature excursions to 100C. It's too slow to ramp up and too quick to ramp down. It is possible to tune this from userspace under Linux. But really the up ramp should be controlled by a leading indicator like the voltage regulator instead of a lagging indicator. Alternately the Linux p-state controller could anticipate the power levels and program a higher fan speed.
- mips_r4300i 3y agoI'm dealing with this right now (Asus ROG Z790, 14900k, Noctua NH-D15). The stock fan curves seem to be ineffective and also annoying as they are hunting constantly. Single P core temps bounce around constantly causing the fans to be spastic. I have read increasing the ramp up time would smooth out the fan behavior but your experience says this can cause processor fails.
- jeffbee 3y agoExact same setup, I'm afraid. Best solution I have been able to come up with is to read the datasheet of the SuperIO on that board and tune the hysteresis parameters from Linux after boot.
- PawBer 3y agoReminds me of this Raymond Chen classic: https://devblogs.microsoft.com/oldnewthing/20050412-47/?p=35923 https://devblogs.microsoft.com/oldnewthing/20050412-47/?p=35...
- lostmsu 3y agoI wonder why didn't they add a system crash analyzer component that would tell user their CPU is misbehaving (xor eax, eax) to save themselves some hard to debug support volume.
- perryizgr8 3y agoIf I ever encounter a CPU bug causing problems in my production code, I will consider my life complete. I will be satisfied that I've practiced my profession to a high degree of completeness.
- dist-epoch 3y agoYou should go work for Facebook. At their scale they are encountering CPU bugs daily: > This has resulted in hundreds of CPUs detected for these errors https://arxiv.org/abs/2102.11245 https://arxiv.org/abs/2102.11245
- Ochi 3y agoSo ideally, we should disable hyper threading to mitigate security issues and now also disable turbo mode to mitigate memory corruption issues. Maybe we should also disable C states to avoid side-channel attacks and disable efficiency cores to avoid scheduler issues... and at some point we are back to a feature set from 20+ years ago. :P
- bee_rider 3y agoIMO it is worth noting that the “turbo mode,” as you call it, seems to be an overlock that some motherboards do by default. Not the stock boost frequencies. The hyperthread and c-state stuff, eh, if you want to run code that might be a virus you will have to limit your system. I dunno. It would be a shame if we lost the ability to ignore that advice. Most desktops are single-user after all.
- blibble 3y agoturbo boost is an advertised feature of the chip these chips that have been specially binned because they are supposedly stable at those frequencies (within an envelope set by intel) if intel can't get it to work they shouldn't be selling these chips at all
- bee_rider 3y agoUnless I misread the blog post, there doesn’t seem to be any issue with the stock turbo behavior.
- alwayslikethis 3y agoProvided enough cooling, a chip that can boost to its turbo frequency for a few seconds should also run stably at that frequency indefinitely. Nowadays these boost clocks are so high that there is often not much gained by pushing any further.
- dist-epoch 3y agoIntel should police their own ecosystem.
- ryukoposting 3y agoI vaguely recall motherboard vendors ignoring Intel's power recommendations a couple years ago, which was causing weird thermal/performance issues (was it Asus?). I get the impression that's what's happening here, again.
- adamc 3y agoWhile I appreciate their point of view, from a consumer pov this would definitely be a failure of their software, since an implicit requirement is that it has to run on the customer machines. People aren't going to throw away their CPU for this, they are going to return the game (if possible), and certainly express the bad user experience they had with it.
- xcv123 3y agoThis is a hardware fault causing other software to fail. Intel and mainboard manufacturers have recommended workarounds. Customers are not stupid and they know who is at fault.
- adamc 3y agoI predict you are wrong and there will be returns. It's not really a question of stupidity, but of what options are available to them.
- xcv123 3y agoNo they just follow the provided instructions, go into the BIOS setup, and fix the settings. These are $1k CPUs. Purchased by enthusiasts, not retards. If their CPU is unstable they will want to fix it.
- deleted 3y ago[deleted]
- johnklos 3y agoFor years I've had this impression that Intel CPUs were, to put it simply, trying too hard. I administer servers for various companies, and some use Intel even though I generally recommend AMD or non-x86. A pattern I've noticed is that some of the AMD systems I administer have never crashed or panicked. Several are almost ten years old and have had years of continuous uptime. Some have had panics that've been related to failing hardware (bad memory, storage, power supply), but none has become unstable without the underlying cause eventually being discovered. Intel systems, on the other hand, have had panics that just have had no explanation, have had no underlying hardware failures, and have had no discernible patterns. Multiple systems, running an OS and software that was bit-for-bit identical to what has been running on AMD systems, have panicked. Whereas some of the AMD systems that had bad memory had consumer motherboards with non-ECC memory, the Intel systems have typically been Supermicro or Dell "server" systems with ECC. In one case two identical Supermicro Xeon D systems with ECC were paired with two identical Steamroller (pre-Ryzen) AMD systems. All systems provided primary and backup NAT, routing, firewalling, DNS, et cetera. The Xeon systems were put in place after the AMD systems because certain people wanted "server grade" hardware, which is understandable, and low power AMD server systems weren't a thing in that time period. Over the course of several years, the Xeon systems had random panics, whereas one of the AMD systems had a failed SSD, but no unplanned or unexplained panic or outage, and the other had never had a panic or unplanned reboot in all the years it was in continuous service. Had I collected information more deliberately from the very beginning of these side-by-side AMD and Intel installations, I'd have something more than anecdotal, but I'm comfortable calling the conclusion real: multiple generations of Intel systems, even with server hardware and ECC, have issues with random crashes and panics, on the order of perhaps one every year or two. I do not see a similar instability on AMD, though. With brand new Intel CPUs taking substantially more power than similarly performing AMD CPUs, we have a more literal example of what I think is the underlying cause: Intel is trying way too hard to get every tiny bit of performance out of their CPUs, often to the detriment of the overall balance of the system. Between the not insignificantly higher number of CPU vulnerabilities on Intel due to shortcuts illustrated by the performance losses from enabling mitigations, and the rather shocking power draw of stock Intel CPUs that have turbo boosting enabled, I can't recommend any Intel system for any use where stability matters.
- 3y ago
- terrelln 3y agoWe also regularly run into hardware issues with Zstd. Often the decompressor is the first thing to interact with data coming over the network. Or like in this case the decompressor is generally very sensitive to bit-flips, with or without checksumming enabled, so notices other hardware problems more than other processes running on the same host. One decision that Zstd made was to include only a checksum of the original data. This is sufficient to ensure data integrity. But, it makes it harder to rule out the decompressor as the source of the corruption, because you can't determine if the compressed data is corrupt.
- mjevans 3y agoCompressed data is like a backup. It's not valid until it's tested.
- FileSorter 3y agoI recently had to RMA my i9-13900KS because it was faulty. I was experiencing some of the weirdest behavior I have ever seen on a PC. For example, 1. Whenever I tried to install Nvidia drivers I would get "7-zip: Data error" 2. A fresh install of Windows would give me SxS error when trying to launch edge 3. I could not open the control panel 4. BSOD loop on boot
- dontupvoteme 3y agoDecompression failure immediately went to a far grimmer failure mode in my mind..
- colombiunpride2 3y agoI wonder if this is partly related to the LGA1700 frame problems that tend to bend the heat spreader. There are two after market contact frames that drop the temperature around 10 Celsius and ensure flat contact with the head spreader. The stock frame causes the center of the head spreader to dip. I wonder if the turbo boost is controlled by a Proportional–integral–derivative controller.(PID) The idea that the parameters are fine tuned to slow down processor speed as it heats up but before it overshoots its maximum threshold. If those PID values are tuned to assume flat heat spreader/heat sink contact, I can see where a bent heat spreader could cause the cpu to overshoot its safe limit and cause errors.
- ezekiel68 3y agoBased on the article contents, this doesn't seem to be CPU errata. We already know that overclocking above a certain point will cause OS crashes. This seems to be system instability just below the threshold of crashing. Aggressive power and clock settings manifest as this instability without causing an actual crash. I don't find this situation much different than needing to dial back BIOS settings when actual crashes are observed.
- sedatk 3y agoExactly. I remember when I overcloked my 486DX4-100 to 120Mhz, everything would work fine but the floppy drive. It just wouldn’t work for whatever reason. Never thought it was a CPU issue, I’d just asked for it.
- SergeAx 3y ago> overly optimistic BIOS settings First time I see this nice euphemism for overclocking.