3 ms·
Didn't know that. Since they know how to fix it on TR/Epyc, why would they release a brand new line with the same issue?
by mrchicity 9y ago
Didn't know that. Since they know how to fix it on TR/Epyc, why would they release a brand new line with the same issue?
- paulmd 9y agoWell, we don't actually know if they did fix it on TR/Epyc. Epyc is on a newer stepping (that has not been released to the consumer market yet), but the issue was only acknowledged a couple months ago. The B2 stepping may already have been taped out at that point. We don't really know because you pretty much can't get ahold of Epyc systems at the moment (at the moment, the only way is by ordering prebuilt servers from SuperMicro). There hasn't been the degree of testing that you have on Ryzen. But, if the B2 stepping does contain a fix, that would be the reason that Epyc is immune. Threadripper is on the same stepping, but it's cherrypicked silicon (AMD claims they're the top 5% of silicon off the line). There is a wide range of severity reported with Ryzen processors, some people segfault in seconds, others it takes 24h+ of testing before it manifests. That strongly suggests that it's a litho problem, and the fact that Threadripper is on cherrypicked silicon may insulate it somewhat from this issue. But in general I don't see any guarantee there either. There isn't anywhere near as much testing as there is on Ryzen here either (although certainly more than Epyc). From what I've read the finger points pretty strongly at a cache problem (likely the uop cache) and it's also possible that TR/Epyc's NUMA configuration somehow avoids the trigger conditions for this bug. I don't think this case is particularly likely but it's one possible explanation. Frankly though I don't understand why AMD hasn't released the B2 stepping to the consumer market. They are using it for Epyc but they are still producing Ryzen and TR on the original B1 stepping and appear to be releasing Raven Ridge on the B1 stepping as well. I don't know whether the B2 stepping contains a fix for this issue, but I'm sure it has some general fixes that would be good to have.
- mrchicity 9y agoMerely anecdotal, but I've run heavy multi-threaded compile tests on my 1950x for days to check whether the chip was stable when overclocked, and have never seen it segfault. That was my main purpose for purchasing the CPU. I use it for that all day every day. If I were spending millions on Epyc servers, I'd certainly test this on a reasonably large sample before committing.
- paulmd 9y agoYeah, given the binning that's in play I'm not particularly worried about TR. Out of an abundance of paranoia I would recommend a 24h run of kill-ryzen.sh on literally any Ryzen/TR processor (including RMA replacements) but TR is not likely to be a problem. For that matter, the chips that are taking 24h+ to segfault are not likely to be a problem either. It's the ones that are segfaulting in a matter of seconds or minutes that are going to be problematic for normal users. However, that's my thoughts as to why we're not seeing it on Threadripper despite it being on the same stepping. If it were a simple microcode fix (even if you needed to disable some codepath and cut a few percent off) then there is no reason not to pass that down to consumer Ryzen. So far, however, a microcode fix has not been forthcoming. So it's either TR is better binned and thus less subject to the fault, or that something about TR's NUMA layout is breaking one of the necessary conditions for the cache fault to occur. I do think AMD stepped up their QA after the fault was discovered. I've actually heard of post-week-30 retail samples passing kill-ryzen runs, whereas essentially 100% of the pre-week-25 chips display the fault. However, there are definitely post-week-30 chips that do still display the fault. I assume what's going on there is AMD is doing a quick run of kill-ryzen to weed out the shittiest chips, the ones that die in seconds/minutes. But since there isn't a way to deterministically reproduce this fault, and AMD can't realistically spend 24h testing every single chip for one fault, some of the only-somewhat-shitty chips are still leaking through their QA testing. I would take a random week 33 chip over a random week 08 chip for sure, the quality is definitely up, but week 30+ is no guarantee that your chip won't segfault either.