3 ms·
While the full root cause has not yet been found or resolved the limited issues have been pretty reliably resolved by disabling ASLR. Whatever the root cause it
by examancer 9y ago
While the full root cause has not yet been found or resolved the limited issues have been pretty reliably resolved by disabling ASLR. Whatever the root cause it is likely the issue can be fixed through BIOS/microcode updates.
The number of people affected are low. My Ryzen machine has only ever run linux and compiles a lot and has never exhibited this behavior. Also, most new platforms have issues, even new server platforms. These will be worked through during substantial validation server OEMs will go through.
Lastly, look up the errata list for any Xeon CPUs. Intel releases microcode updates for them several times a year to fix bugs. Modern CPUs are complex and will pretty much always have bugs. Luckily some combination of BIOS or microcode updates will almost always resolve them.
- dis-sys 9y agoI am a Ryzen user, actually built my Ryzen system on its release day. For your claims on Xeon's bugs, just wondering when was the last time Intel had to patch the microcode to fix bugs that can continuously crashing day to day workloads such as compiling some linux packages? In case you don't fully understand the situation - when Ryzen was released, it doesn't work with many memory modules on the consumer market, as of today, for pretty high probability, you still don't get the top speed of RAM you paid for, it crashes on day to day compilation jobs, the AMD GPIO linux module maintained/contributed by AMD is too buggy to run on Gigabytes motherboards, oh, let's don't forget the FMA3 bug. Downplaying the issues causing troubles for Ryzen users do not get Ryzen better.
- sliken 9y agoHow about the intel fdiv bug? My understanding is that BIOS updates have improved the speed, stability, and clock speeds available. This is far from unusual, read newegg for any new intel socket/chip and the reviews are full of dimms XYZ didn't work with motherboard ABC. Every motherboard manufacturer posts a list of compatible DIMMS they test with, unfortunately they are rarely the dimms available to consumers to buy in quantity 2-4. Thus the market opportunity for crucial that does their own testing and has a generous replacement policy.
- deleted 9y ago[deleted]
- jmgao 9y ago> For your claims on Xeon's bugs, just wondering when was the last time Intel had to patch the microcode to fix bugs that can continuously crashing day to day workloads such as compiling some linux packages? Intel had an even worse one in 2014: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=762195 https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=762195 Glibc merged a patch to use Intel's shiny new TSX transactional memory extensions when available to do hardware lock elision, except TSX was completely busted, and any time a process used pthread_mutex_lock, there was a chance that it would immediately crash, or corrupt memory silently. (In practice, this would happen all the time.) Their solution was to release microcode that just turned off TSX entirely.
- smilekzs 9y agoTSX has been an experimental new feature --- disable it and all is good again, since there is no production-quality software targeting it. On the other hand, to me it seems that Ryzen segfault bug cannot be simply eliminated. Disabling ucode cache might help. Disabling ASLR "pretty reliably" (per grandparent) resolves it. But I would imagine "pretty reliable", sans official investigation report from AMD, not reliable enough for those operating server farms, which jeopardizes the main value proposition of EPYC. By no means am I bashing AMD's new tech though --- it's just CPUs are so critical that you want strong guarantees on "from this ucode patch onwards, you won't get hit by this particular problem, at least 99.9999% of the time".