5 ms·
The real “big caveat” here is to make sure that this actually improves your real workload in real life with real inputs. Zen and Zen+ have a lot of edge cases,
by seriesf 7y ago
The real “big caveat” here is to make sure that this actually improves your real workload in real life with real inputs. Zen and Zen+ have a lot of edge cases, for example their pdep/pext support is just crazy slow and how slow it is depends on the input values (it gets slower if more bits are set, which is bananas). But this also goes for genuine intel hardware. MKL is not always optimal. Even -march=native does not always produce the best code. Your program may be faster with the latest instructions disabled.
- partingshots 7y agoIf you go to the MKL library, it specifically says it speeds up performance on non-Intel processors as well. That they would state this and still sabotage the performance for AMD processors by directly disabling performance enhancements for them shows to me that this was a calculated decision that was willfully made.
- seriesf 7y agoOr it shows that the mkl developers are aware of some obscene edge case that kills the performance of a major customer’s application. This article tested ONE thing.
- lhl 7y agoIt's certainly not an "obscene edge case" - MKL cripples AMD on every single matrix operation tested here: https://github.com/flame/blis/blob/master/docs/graphs/large/l3_perf_epyc_nt1.pdf https://github.com/flame/blis/blob/master/docs/graphs/large/...