4 ms·
"The Intel processor has nifty 256-bit SIMD instructions. The Apple chip has nothing of the sort as part of its main CPU. So I could easily come up with example
by bfgoodrich 6y ago
"The Intel processor has nifty 256-bit SIMD instructions. The Apple chip has nothing of the sort as part of its main CPU. So I could easily come up with examples that make the M1 look bad."
The M1 has 128-bit NEON SIMD, and given its decode pipeline and cache efficiency it seems more likely to actually benefit entirely from it. AVX on Intel devices is often of limited value because it's either memory starved or gets throttled (many SKUs throttle once you start using AVX).
I've had cases where vectorizing ongoing sequential processing in the most optimal fashion available barely gave a single digital percentage increase. 9 times out of 10 it's snake oil.
But regardless, if there's an example that would make the M1 look bad, do it.
- thewebcount 6y agoThis is one of the things I missed about AltiVec from the PPC days. They really designed it nicely. It was a different set of registers that worked normally. The permute instructions were very useful! When trying to port AltiVec code to Intel SSE at the time, it would often come out worse because of all the constraints. You couldn't intermix floating point and SSE code because they used the same registers. There were stalls if you did as it switched back and forth. A friend actually hired Intel engineers to port his AltiVec code to Intel at the time and even they couldn't make it work as fast as it was on his PPC Mac. So I'm hoping the Neon instructions bring back some of the elegance and sanity we had in the PPC/AltiVec days.
- kitsunesoba 6y agoMan that takes me back, it used to be pretty common to see Mac apps advertise their leverage of AltiVec, even for “consumer” stuff. It may be entirely imaginary but I always felt like the PowerPC versions of OS X were more responsive than their Intel counterparts, despite running on CPUs that were at a disadvantage in terms of clock speed. I wonder how much of that was due to the PowerPC arch itself.
- watersb 6y agoYeah I really wanted to dig into AlitVec a bit more, but by the time I took a look at it, we had moved on to Intel. So I just tried to use the OS X vector libraries where possible, and hope for the best. There was a paper that really got my attention, from US Air Force research I think, that reported a micro benchmark on a Mac Mini G4 where the AltiVec scores blew away their best Silicon Graphics beast. The M1 reports feel a lot like that. Or better, because of the reports coming in on huge performance gains in real world apps.