14 ms·
AMD Threadripper 3970X under heavy AVX2 load: Defective by design?
- Filligree 7y agoHe's not alone; I've had similar problems with my 3960X. It seems to be a power delivery issue, and fortunately fixable if you disable all spread-spectrum and VRM power-saving options, but the Zen series seems a tricky beast. I've had machine crashes triggered by using the "wrong" CPU scheduler under Linux. It's amazing, in a horrible way.
- xxs 7y agoThat would be a motherboard issue not a CPU one, though. Has anyone attached an oscilloscope?
- franzb 7y agoYes: https://forum.level1techs.com/t/3970x-prime95-stability/153206 https://forum.level1techs.com/t/3970x-prime95-stability/1532...
- codyb 7y agoGreat thread thanks for sharing. Really makes you want to order an oscilloscope and some CPUs to putz around with. Everything in computer science is such a rabbit hole, great field to be in.
- gothroach 7y agoIf you're interested in a relatively inexpensive oscilloscope to play around with, check out the Rigol 1054Z (or 1074Z Plus if you want MSO capabilities). It's four channels and has more than enough features for playing around with, especially for the price. Using the Riglol website, you can unlock all software options on it, including increasing the bandwidth to 100mHz and the memory to 24MP. It was the first 'real' scope I bought and I still use it a fair amount despite having upgraded to a Rohde and Schwarz MSO model. I've been amazed how much I use an oscilloscope after getting one, from measuring ripple on power supplies to diagnosing serial communication issues.
- fivesixzero 7y agoThanks for the recommendations! I have a DSLabs DScope (100 MHz, 2-channel FPGA scope) and while it’s handy I’d prefer to have a proper hardware scope someday. Rigol’s scopes look like they nicely fit in between the basic DSO/FPGA stuff and the “proper” 4-5 digit priced test bench gear. Any recommendations for learning resources that could help with understanding DC power supply analysis for non-EE types? While refurbishing laptops and working with microcontrollers I’ve run into some odd things where ruling out transient power supply issues would probably be helpful.
- gothroach 7y agoThe low end Rigols make good entry-level scopes, and have a surprising amount of capability for the price. As for learning resources, I came across a decent article on the subject when I was starting out (1), and most of the oscilloscope manufacturers have whitepapers on SMPS diagnostics, the Tektronics one I read a while back (2) gave a good overview. A lot of the whitepapers have a manufacturer-specific focus, but they still have good information that can be applied to almost any oscilloscope. If you want to get really into the power supply and do high-side measurements you'll need an isolated differential probe, which can cost as much as an inexpensive oscilloscope, but for DC output measurement you shouldn't need anything special. Current probes are a lot more affordable if you're interested in looking at loads or current fluctuations/harmonics, but that's more useful after you've figured out a bit more what specific properties you're trying to measure. 1: https://www.testandmeasurementtips.com/test-switching-power-supplies/ https://www.testandmeasurementtips.com/test-switching-power-... 2: https://download.tek.com/document/3GW_23612_7.pdf https://download.tek.com/document/3GW_23612_7.pdf Edit: I forgot to mention that the EEVblog forums are a good resource also, but they sometimes aren't as friendly as they could be towards people just starting out.
- tzs 7y ago> Using the Riglol website, you can unlock all software options on it, including increasing the bandwidth to 100mHz and the memory to 24MP. It's interesting how Rigol did the locking of those extra features. The unlock key for a feature set for your scope has to be signed by a Rigol private key using an elliptic curve signature system. But they are only using a 56 bit private key. That was quickly brute forced, and key generators proliferated. They used a good library for the cryptography stuff, and except for the short key seem to have used it well and knew what they were doing. This suggests that the choice of a weak key was deliberate. Each family of scopes has its own private key. As few families came out with new private keys, Rigol continued to use 56 bits. When major firmware upgrades came out in existing families, where they could have easily changed to a longer private key, they kept the same 56 bit that was now widely circulated on the net. It seems pretty clear that they are not interested in stopping people from free unlocking.
- ComputerGuru 7y agoI'm really confused the OP there didn't try a different sTRX4 motherboard.
- franzb 7y agoOP here. I wish I had! Unfortunately that's my work machine and I don't have the time to somehow get my hands on a new fairly rare, very expensive motherboard, disassemble my workstation and rebuild it just to see if my motherboard is the culprit. From the experience of other owners of the 3970X, this problem happens with other TRX40 motherboards.
- ComputerGuru 7y agoMakes sense. I’ve done Amazon Prime and returned it because I consider that a defective mobo and fair game to return in the past.
- xxs 7y agoAt least EU Amazon is extremely good about RMAs, e.g. international expresses delivery (next day) even without returning the defective component 1st. Return shipment is also covered, i.e. free for the customer. Dunno, if they'd do that for high priced stuff, though. p.s. lovely/lively oscilloscope shots!
- leeter 7y agoI remember when Buildzoid of AHOC did the mobo breakdowns of the TR4 boards and thinking that while some of the super high end boards were probably good for this sort of beating, the mid and low range might struggle hard with the 64/128 part if it ever came into being (3990X was just a rumor at that time). But it looks like they need more caps to handle the transient response time, and probably also some firmware fixes to slow ramp because I don't think all the SMD caps in the world are going to handle that sort of ramp. It's just not possible to get them close enough to the actual CPU without literally putting them under the IHS.
- franzb 7y agoHere's BuildZoid review of the motherboard I'm using (GIGABYTE TRX40 Aorus Xtreme): https://www.youtube.com/watch?v=HMUWzDSAS9c https://www.youtube.com/watch?v=HMUWzDSAS9c A very interesting watch if you're interested in electronics in general and in power delivery in particular (the whole YouTube channel is awesome to be honest).
- leeter 7y agoThat was one of the few that I thought could handle the 64/128 part. However look at the output filtering: It's roughly the same if not a cap or two larger as what you'd find on x299... which is a higher voltage and thus has lower amperage requirements. I have yet to find a back of board shot but unless it has a ton of SMD AL-poly caps back there that board would still struggle with the ramp described in the thread. Even then I'm not sure the socket resistance wouldn't cause enough V-droop to cause a crash anyway.
- close04 7y agosTRX4 can take one of three CPUs ranging from 24 to 64 cores. The TRX40 Aorus Xtreme that's mentioned in the thread should be the absolute top of the line even if it was launched before the 64 core monster was available. So I'd expect it to work just fine with the 3970X which is a 32 core part. But I wonder if this is a Gigabyte issue who have a history of playing around with the power delivery and using "fake phases" (to the point where they now have to advertise their boards as having 16 "real phases"). As far as I can tell many (most?) reviewers benchmarked the board with 24 core CPUs and most likely skipped on the power intensive tests.
- wtallis 7y ago"A tricky beast" definitely seems fair. These things have to scale per-core power consumption from ~13W down to ~3W on the fly to stay within their 280W limit. To my knowledge, AMD hasn't implemented any proactive measures that are as severe as Intel's AVX512 strategy, where the instructions get split and handled by the narrower vector units for a surprisingly long time while powering up the full-width vector units (and dropping CPU clocks). This AMD instability is only using AVX2: 256-bit SIMD rather than 512-bit. But spread across so many cores, that's still a lot of FPUs to be lighting up.
- tedunangst 7y agoThat would suggest the instability might be resolved by adding threads one at a time, but does prime95 have an option like that?
- franzb 7y agoThat's an interesting thought!
- Hello71 7y agoyou can probably bodge it by running a bunch of copies with limited workers, then run your main copy, then kill the starters. or, you could fiddle with the CPU affinity.
- Filligree 7y agoThat's something you can do outside of prime95; the schedutil scheduler does that by default, which I believe is why I had hard crashes with ondemand and not schedutil. (Ok, it actually ramps up frequencies slower rather than lock out cores, but the effect is the same.)
- nottorp 7y agoSo basically Intel underclocks itself to do the AVX stuff while AMD underclocks itself ... less?
- 7y ago
- wyldfire 7y ago> fortunately fixable if you disable all spread-spectrum If these are on by default, are they required for FCC certification conformance? Presumably the device has not been EMI/EMC tested in the mode where spread-spectrum was disabled.
- Filligree 7y agoThey are, but I'm not physically in the USA. At any rate, isn't a computer case basically a Faraday cage? The frequencies are up in the microwave band, so I can't imagine there's much range.
- wyldfire 7y ago> computer case basically a Faraday cage? If it was designed well ;) But that's only effective for addressing radiated emissions, not conducted ones.
- nottorp 7y agoHmm are we having another "AMD motherboards are crap" moment? Or is it simply that delivering 200+W at load through a CPU socket can't be reliably done at consumer prices? Anyone has had this problem with less high end CPUs? Something at 95-65 W?
- close04 7y agoAnything with sTRX4 socket should be able to handle at least between 250W and 280W. But it's not out of the question that many motherboard designs draw a lot from AM4 boards which have much lower requirements, to keep costs down. This may be the result. The "CPU defective by design" in the title might be a bit misguided since the suggested workarounds do not address a CPU issue but a motherboard one.
- franzb 7y agoI agree (my fault, sorry) but the fact is that several people are encountering the exact same issue with a variety of motherboards. So either all those motherboards are "defective" or the problem is more complex than it looks. My motherboard is the highest end consumer motherboard GIGABYTE has ever built (as far as I know), and I wouldn't say that they tried to keep costs down (it's listed at $849 on PartPicker, and it was introduced at $999) and the power delivery stage is insane. Here's BuildZoid's _in-depth_ review and analysis of its VRM: https://www.youtube.com/watch?v=HMUWzDSAS9c https://www.youtube.com/watch?v=HMUWzDSAS9c
- ComputerGuru 7y agoI learned a long time ago to never get the most expensive motherboard (ASUS Republic of Gamers many years back): they're not as heavily tested as the cheaper ones. They are only ostensibly better; in practice, the mid-level "workstation-class" motherboards that feature similar capabilities but missing fancier options are put through the works and far more thoroughly tested and debugged. Just look at the BIOS changelogs and even the hardware revisions for high-but-not-flagship-high motherboards as compared to the flagship models. When you buy the top-of-the-line, you're on your own. I now literally go out of my way not to get the flagship motherboard models, even if it means holding off on a purchase until a lower-spec'd model comes out, and have never regretted it since. (I also will never again buy Gigabyte motherboards, either.)
- ntauthority 7y agoMy 3970X on the ASRock Taichi (with default settings, generally) does not seem to reproduce this issue at this time - the system remains operational despite the FMA3 path being used (I'm assuming this is behind the AVX2 flag? disabling FMA3 leads to a plain AVX path) while running an all-core test with 16K FFTs in Prime95. Either a slight background workload (Windows seems to be trying to use half a core for an OS update) resolves this, or this board does not have a broken power design?
- franzb 7y agoThanks for sharing your experience. What version of Prime95 are you using? Make sure to use the latest one (v29.8 build 6). The only CPU options I see in the torture test settings are: (1) Disable AVX-512 (grayed out since unsupported on this CPU), (2) Disable AVX2, (3) Disable AVX. There's nothing about FMA3.
- ntauthority 7y agoYeah - I explicitly updated to the latest one; the FMA3 setting is one that existed in prior builds in local.txt so I toggled it off there just to be sure I was hitting an AVX2 code path (in case it didn't mean the UI saying FMA3 in each worker window), but it seems to interpret AVX2 as being synonymous to FMA3 I guess.
- hrgiger 7y agosame here on 1950x, I built from source, when toggled avx2 it shows "using type-1" I dont know what it is: https://pastebin.com/tPuYzYC0 https://pastebin.com/tPuYzYC0
- c2h5oh 7y agoConsider trying a couple different versions/builds of prime95 If you look at just the list of issues fixed in v29.x series there is more than one that is close to the code path that is problematic here: https://www.mersenneforum.org/showpost.php?p=508842&postcount=2 https://www.mersenneforum.org/showpost.php?p=508842&postcoun...
- ComputerGuru 7y agoCan anyone suggest a different CPU load-testing tool other than prime95, that might catch things prime95 wouldn't? I have a machine running a 1950X and I get random ffmpeg segfaults anywhere from six to eight hours in to an encoding session with all 16 cores fully loaded, but the machine is prime95 stable for a week+, so I suspect it's an AVX/AVX2 issue.
- leetcrew 7y agoyou can test avx with newer versions of prime95. you probably shouldn't run that for a week with small ffts though.
- ComputerGuru 7y agoI never realized prime95 was still updated! I seem to remember there was a time when the newest prime95 releases were several years old and simply assumed that was still the case. I just ran whatever copy I had in my downloads archive at the time; I just checked and it was version 29.3 build 1.
- DoofusOfDeath 7y agoDo you really mean segfault (i.e., SIGSEGV), or do you just mean that the program crashes in general? Either way, you might consider bringing this to the attention of the ffmpeg developers [0]. If they don't have a fix already, that may appreciate your help in root-causing the bug. [0] https://lists.ffmpeg.org/mailman/listinfo/ffmpeg-devel/ https://lists.ffmpeg.org/mailman/listinfo/ffmpeg-devel/
- ComputerGuru 7y agoYes, a literal SIGSEGV but the dump doesn't indicate anything amiss. It's not an ffmpeg issue because other long-running software also crashes when run alongside an encode job (e.g. I've had rav1e crash) some hours in, too. I prefer not to send software developers goose hunting unless I have some sort of valid repro case, which I don't, not really.
- t0mas88 7y agoMost performance motherboards with Intel unlocked K models will downclock the maximum boost when using AVX instructions. The reason is very high power draw and temperatures. For example my i5 9600k runs at 5ghz turbo boost on all cores but 4.7 when using AVX. If I disable that option it crashed with prolonged usage like benchmarks. Edit: To be clear, the i5 9600k is sold as 3.7ghz with boost up to 4.6 on a single core. So there is a difference with the AMD case in that this doesn't happen on the setting Intel sell it at.
- paulmd 7y agoAVX offset is a configurable parameter with unlocked (K- or X- series) processors. You can run with 0 offset at all but the highest overclocks. At some point it does stop being worth it though, because the power/voltage implications of 5 GHz AVX are so severe/potentially damaging to the chip. It is a lot of current and current kills chips. SiliconLottery does all their validation with a 200 MHz offset.
- leetcrew 7y agothere's no real reason to try and hit 0 offset anyways. if you get it stable, it usually implies you could just increase the multiplier and offset by some n>0 for greater overall performance.
- paulmd 7y agoThe problem is that real-world code (including games) often includes at least some AVX instructions, so the AVX number is often more representative of "real" performance. The way AMD does it where it's a smooth transition based on current/thermals is definitely better than Intel's "whoops, AVX instruction, pump the brakes!". But yes, if you can get higher clocks at least some of the time then you might as well. The other other downside is that flipping between power states can cause problems/crashes too. It shouldn't but it can.
- leetcrew 7y ago
- frou_dh 7y agoThe original Ryzen (1000 series) shipped plenty units with hardware defects that could be exposed by running parallel compiles. The so-called segfault bug: https://www.phoronix.com/scan.php?page=article&item=new-ryzen-fixed&num=1 https://www.phoronix.com/scan.php?page=article&item=new-ryze...
- ComputerGuru 7y agoFrom your link: > AMD has confirmed this issue doesn't affect EPYC or Threadripper processor The CPU Michael was using was a Ryzen 7, not a Threadripper. But I don't know if TRs were later found to have the same problem, which wouldn't surprise me.
- paulmd 7y agoIMO the likely cause was some kind of cache bug due to manufacturing. Threadripper processors were binned for tighter cache timings (they had the same timings as the 2000 series desktop chips) and thus didn't suffer from it. EPYC was on a different stepping entirely and didn't share dies with the consumer or HEDT processors. I suspect they likely had the tighter cache timings although I'm not sure on that. Regardless, it's also possible that they were just binned out, or nobody ever encountered it. EPYC was not a very high-volume product, most people did test installations and said "yeah, we'll wait for Rome". Quad-numa on a package and lower-than-Intel IPC on a pretty poor node was not a winning formula. Once the problem was realized, desktop Ryzen processors started getting binned for it as well. A few slipped through here and there, so it wasn't a manufacturing change, I think binning is the most likely explanation.
- bdd 7y ago“Unable to perform AVX2 instructions correctly under heavy load” is also a common “WTF Intel!?”–inducing phenomenon. I’m certain SREs who work at companies with more than 1 million servers have a bunch of hair pulling stories. Most (all?) Intel server CPUs in fact decrease clock speed when executing AVX2 (and some other) instructions to keep things a bit more sane. Vlad from Cloudflare wrote about this, more specific to AVX-512 back in 2017: https://blog.cloudflare.com/on-the-dangers-of-intels-frequency-scaling/ https://blog.cloudflare.com/on-the-dangers-of-intels-frequen... Then there is PROCHOT signal. Which is supposed to protect the CPU from getting too hot but keeps getting raised in lopsided AVX2 loads not because CPU is too hot but voltage regulation gets whacked. You may wonder: what is an example of AVX2 heavy load. RSA multiplication is a good candidate. AES constructions or modes (CBC with SHA, GCM) are implemented in AVX2-BMI2 as well.
- fivesixzero 7y agoI’m curious if this behavior defined by something in hardware, microcode, boot-time BIOS flags, or higher level kernel/hypervisor/application code.
- gameswithgo 7y agoon unlocked intel cpus you can change the avx multiplier in the bios.
- olliej 7y agoOh neat (I haven’t messed with over clocking in years) - is it just avx that you can tailor? (Beyond the old school bus multipliers)
- gameswithgo 7y agoThe newest AMD CPUs decrease clock speed based on parameters like heat, which AVX2 under heavy load will cause. So, AMD also decreases clock speed when executing AVX2, indirectly. Though in a very different fashion, continuously, rather than a distinct mode.
- 7y ago
- grokas 7y agoGreat to see L1T posted here.
- johnklos 7y agoDo some more investigating. Drop the memory clocks to stock, drop the CPU clocks to, perhaps, 3 GHz, and see if the same issues happen. If they do, there's a systemic issue that needs to be addressed. If the issue disappears, try raising the clock incrementally until the issue reappears. Get a Kill-a-watt and look at power usage for each frequency and graph the results.
- citilife 7y agoIt appears this could just be a prime95 bug from reading the comments.
- xxs 7y agoThere is a beautiful oscilloscope shot, that shows it's a power delivery issue - the voltage sharpy drops and no software bug should cause it. It appears to be a VRM issue, either a design flaw or incorrect setup by the BIOS.
- cma 7y ago> Finally, a note on CPU temperatures: At idle the CPU hovers around 39-50 °C and tops around 72-78 °C under full load. I’m using the best air cooling setup I could think of and get my hands on, but it’s still air cooling, and my system is installed in a closed case (but with extreme attention to airflow). I know air coolers can be competitive, but it says right on the outside of the 2950X box that you should use liquid cooling.
- retrovm 7y ago78C is not a really objectionable temperature for a CPU, if that's really the on-die temperature. Usually thermal protections are set to 90C.
- ehutch79 7y agoSo, does this mean threadripper is unusable and we shouldn't buy them?
- ehutch79 7y agoI'm not sure why I'm being downvoted. This is a legitimate question. Does this bug make the cpu unusable? Does the work around for it slow down the per core performance to the point of unusability?
- fivesixzero 7y agoWhether the issue is with the CPU hardware, the mainboard design (VRM, etc), mainboard BIOS, kernel, or the Prime95 app itself still appears to be an open question. Based on oscilloscope analysis of the VRM output in a linked thread elsewhere in the comments it looks like the board’s VRM design, or its configuration by the board’s BIOS, may be the most likely suspect. But there are less-researched reports of similar issues on other boards as well, which makes things a bit more murky. Given the uncertainties there it may put some people off from buying into the TR/sTRX40 platform in general. But to offer a blanket recommendation to avoid is a bit premature.
- topspin 7y agoIt's a little early to start condemning things. It might be motherboard power delivery, a prime95 bug or something else. If it is the CPU there could be a firmware correctable flaw. This is the bleeding edge; these devices have only been available for 13 weeks and someone is indeed bleeding, it seems. It could even be an uncorrectable flaw in the device design, in which case AMD will likely do as they did in 2017: replace the parts.
- boris 7y ago> This is the bleeding edge; these devices have only been available for 13 weeks [...] Shouldn't these devices have been tested by AMD and motherboard manufacturers for months before they were released? Including on workloads like Prime95 which are well known to uncover system instabilities. Keep in mind also that this is not the first time in the Zen/Zen2 history that AMD has shipped buggy products: there was the segfault bug, the random number generator instruction bug, and the unbootable Threadripper on recent (at the time of release) Linux kernel. The worst part, to me personally, is that they were all "surface bugs" that could have been detected by AMD with a little bit of testing. I've been wondering why Supermicro isn't releasing any Threadripper motherboards (while they do for Xeon-W). Maybe this explains it: the consumer CPU business at AMD is a circus and Supermicro wants none of it?
- m0zg 7y agoProbably the motherboard, or VRM brownout to be more exact. That said, I'm glad I did not pick up a 3970X like I was planning to, yet. AMD is pretty great with its warranty, motherboard manufacturers can be a chore.
- daneel_w 7y agoI can reproduce this problem on my non-Threadripper Ryzen 5 3600 Zen 2 CPU. I don't think it's specific to TR. With AVX2 enabled, Prime95's torture test is only stable when I use 3 workers or less. With 4 workers one of them will abort due to an error within 20 seconds. The more workers, the sooner a crash; with 5 workers it happens within 10 seconds, and with 6 workers it happens within 2-3 seconds. If I play with the tests on and off for a while, seemingly increasing the quiescent temperature of the CPU, the whole experience and testing actually becomes a bit more stable. My motherboard uses the B450M chipset.
- Scramblejams 7y agoWhat motherboard make/model?
- daneel_w 7y agoGigabyte B450M DS3H. Currently running with latest BIOS (F50).
- robocat 7y agoNote that @DerAlbi thinks the CPU is fine, instead he suspects the VRM on his Gigabyte mobo for this issue (based on his oscilloscope readings): https://forum.level1techs.com/t/3970x-prime95-stability/153206/21 https://forum.level1techs.com/t/3970x-prime95-stability/1532...
- daneel_w 7y agoThanks for the info!
- d1zzy 7y agoDo you have PBO or overlocking enabled? PBO is not considered stock by AMD, ensure that it's disabled in BIOS (some motherboards incorrectly enable it by default)
- 7y ago
- deleted 7y ago[deleted]
- maljx 7y agoI ran the same test on my ryzen 3900X with no issues, MSI X570 ACE motherboard, seasonic PSU.