7 ms·
Right the arm stuff is probably in the "it runs" camp. Largely because its SVE, which is barely available, and the code written to utilize it has largely probab
by StillBored 3y ago
Right the arm stuff is probably in the "it runs" camp. Largely because its SVE, which is barely available, and the code written to utilize it has largely probably been tuned for the a64fx, or maybe the gravaton v1's.
Both of which have considerably different memory and vector size/issue characteristics. So three different SVE variations now, and the previous two show significant uplift when given custom tuning (ex: see gcc -mtune=neoverse-512tvb, vs the custom a64fx compiler benchmarks). Arm put a bunch of effort into creating an instruction set that is microarch agnostic, but then its not exactly worked the first couple tries. Maybe that will be fixed with V2 and all SVE cores going forward.
- rbanffy 3y agoIndeed. Right now there is about 0% HPC code tuned to Grace and Grace Hopper. I'd love if Nvidia made reasonably priced Grace and Grace Hopper ATX boards (or a Nvidia Studio stylish desktop, priced like a Mac mini) developers could buy so that we can do our best to optimize code for Grace for free in our spare time. Same goes for AMD and their MI300 family, in case AMD is listening. There is less to be gained, as the x86 side is pretty well cared for ATM, but, still, I'd love to see such a beast.