4 ms·
I hoped this just affected Ryzen CPUs, but this Reddit post indicates that it affects Epyc also: https://www.reddit.com/r/Amd/comments/6rmq6q/epyc_7551_mining_p
by powercf 9y ago
I hoped this just affected Ryzen CPUs, but this Reddit post indicates that it affects Epyc also:
https://www.reddit.com/r/Amd/comments/6rmq6q/epyc_7551_mining_performance/ https://www.reddit.com/r/Amd/comments/6rmq6q/epyc_7551_minin...
The first post on on AMD's community forum (https://community.amd.com/thread/215773?start=0&tstart=0 https://community.amd.com/thread/215773?start=0&tstart=0) is almost three months old, so AMD have known about this for a long time. If it's not something that can be fixed in a UEFI update, then it's bad news for everyone: a weakened AMD means more stagnation in amd64
- old-gregg 9y agoSo he's got a segfault every couple of minutes, wow... I've been running the same test for over 4 hours now on my Ryzen 1700 (and I've had several uneventful 30-40 minute runs before). To date, I only got one "internal compiler error: Illegal instruction" but no segfaults. Whatever it is, it doesn't affect every chip the same way.
- userbinator 9y agoWhatever it is, it doesn't affect every chip the same way If it is marginal timing in some part of the chip, that combined with statistical process and environment variations, and the increasingly tiny geometries (which serve to amplify the variation) mean the problem could really occur quite randomly. Modern CPUs are pushing the limits in more ways than one, and IMHO this is what happens when they go too far.
- dis-sys 9y ago> Whatever it is, it doesn't affect every chip the same way. Such uncertainty is bad. segfault is not the only issue, the risk of having some interally corrupted data is the biggest risk. I've stopped using the Ryzen junk I bought on its release day.
- snarfy 9y agoThat's kind of how interrupts work though. It's a random disruption of normal control flow. Something like this could fail a different way every time you run it. If IRETQ is failing to RET under load then this is a huge issue. How do you even fix that? Load balance interrupt processing across cores? Throttle any process in a tight loop? It's all ugly hacks. AMD needs to fix this.
- dogma1138 9y agoThere were some comments that indicated this scales up with the number of CCXs on a given CPU. https://www.reddit.com/r/Amd/comments/6ltdqd/comment/djx6g9r https://www.reddit.com/r/Amd/comments/6ltdqd/comment/djx6g9r There is also a POC for invoking this on Windows on GitHub I'm really wondering what this issue will lead too.