5 ms·
I have been working on a number of projects involving ARM64 assembly lately. My take is that we haven't even begun to max out the potential of Apple Silicon. F
by garbagecoder 5y ago
I have been working on a number of projects involving ARM64 assembly lately. My take is that we haven't even begun to max out the potential of Apple Silicon.
First, many apps still run on Rosetta2. When they are finally converted, libraries and all, to native ARM code, that is real low-hanging fruit yet to be picked.
Second, my sense--this isn't something I have data for--is that the code that is output by compilers can be further optimized much much more than for x64. One example is library routines not using all of the available hardware algorithms, e.g., in cryptogrpahy.
Third, optimizing multithreaded programs for performance vs. efficiency cores. If there's an easy way of doing this, I couldn't find it. I implemented it myself. Making sure intensive tasks are on the performance cores and not randomly assigned to the efficiency cores makes a difference. OS handles this most of the time, I think.
Fourth, Apple could open up some of the internal compute stuff that you need their frameworks for.
Some combination of these things could see our existing M1s feel faster for a long time. If the M2 is a big improvement (hopefully in cores and RAM), I look forward to trying one.
- nebula8804 5y agoThis is music to my ears. Thank you for your efforts. When I first started using my M1 Mac, it was like a computer from another planet. It was so freaking fast it took me back to the 90s/early 2000s when every upgrade was a night and day improvement. I could never believe a computer could be so snappy in everyday usage. It just brings joy every time I use it. Meanwhile I have this 2019 Core i9 Macbook pro at work and it is a slug compared to the M1.
- garbagecoder 5y agoI had the same experience. When I got the DTK I couldn't believe it was running almost everything in emulation and was going to sell for about $500. And that wasn't even an M1. Exploring all of its unique nooks and crannies has definitely taken me back too!
- inkyoto 5y ago> Second, my sense--this isn't something I have data for--is that the code that is output by compilers can be further optimized much much more than for x64. One example is library routines not using all of the available hardware algorithms, e.g., in cryptogrpahy. Your senses are not failing you. Let's pick apart a specific example being the stock openssl binary (/usr/bin/openssl) compiled with unknown C compiler flags vs the manually built one (I have used the one that homebrew ships). The stock binary reports: $ /usr/bin/openssl speed sha256 Doing sha256 for 3s on 16 size blocks: 24823543 sha256's in 2.99s Doing sha256 for 3s on 64 size blocks: 17875138 sha256's in 2.99s Doing sha256 for 3s on 256 size blocks: 13158887 sha256's in 2.99s Doing sha256 for 3s on 1024 size blocks: 5565350 sha256's in 2.99s Doing sha256 for 3s on 8192 size blocks: 874073 sha256's in 2.99s LibreSSL 2.8.3 built on: date not available options:bn(64,64) rc4(ptr,int) des(idx,cisc,16,int) aes(partial) blowfish(idx) compiler: information not available The 'numbers' are in 1000s of bytes per second processed. type 16 bytes 64 bytes 256 bytes 1024 bytes 8192 bytes sha256 132632.74k 382026.73k 1125063.44k 1902869.91k 2391966.26k Let's tinker with the homebrew's CFLAGS and add ARM64 v8.4 specific optimisation flags + unroll loops + free up the frame pointer register for the general use (-Ofast -funroll-loops -fomit-frame-pointer -pipe -w -pipe -march=armv8.4-a+simd+crypto+i8mm+bf16+fp16): $ /opt/homebrew/opt/openssl@1.1/bin/openssl speed sha256 Doing sha256 for 3s on 16 size blocks: 76652798 sha256's in 2.99s Doing sha256 for 3s on 64 size blocks: 52778508 sha256's in 2.99s Doing sha256 for 3s on 256 size blocks: 22320676 sha256's in 3.00s Doing sha256 for 3s on 1024 size blocks: 6744512 sha256's in 2.99s Doing sha256 for 3s on 8192 size blocks: 899002 sha256's in 2.99s Doing sha256 for 3s on 16384 size blocks: 451670 sha256's in 3.00s OpenSSL 1.1.1m 14 Dec 2021 built on: Tue Dec 14 15:45:01 2021 UTC options:bn(64,64) rc4(int) des(int) aes(partial) idea(int) blowfish(ptr) compiler: /usr/bin/clang -fPIC -arch arm64 -Ofast -funroll-loops -fomit-frame-pointer -pipe -w -pipe -march=armv8.4-a+simd+crypto+i8mm+bf16+fp16 -mmacosx-version-min=12 -isysroot/Library/Developer/CommandLineTools/SDKs/MacOSX12.sdk -DL_ENDIAN -DOPENSSL_PIC -DOPENSSL_CPUID_OBJ -DOPENSSL_BN_ASM_MONT -DSHA1_ASM -DSHA256_ASM -DSHA512_ASM -DKECCAK1600_ASM -DVPAES_ASM -DECP_NISTZ256_ASM -DPOLY1305_ASM -D_REENTRANT -DNDEBUG -isystem/opt/homebrew/include -F/opt/homebrew/Frameworks -isysroot/Library/Developer/CommandLineTools/SDKs/MacOSX12.sdk The 'numbers' are in 1000s of bytes per second processed. type 16 bytes 64 bytes 256 bytes 1024 bytes 8192 bytes 16384 bytes sha256 410182.20k 1129707.19k 1904697.69k 2309826.18k 2463085.08k 2466720.43k By simply specifying ARM64 v8.4 specific code generation flags, the speedup for the 16 and 64 byte blocks is nearly 4x, is nearly 2x for the 256 byte test and a 21% speedup for the 1024 byte test. What I have begrudgingly noticed, though, is that way too many open source packages go a very long way to forcefully override CFLAGS / CXXFLAGS environment variables (I set mines in the shell rc file) to set them to something else. What are the most usual / typical choices, though? Not very imaginative: «-O2 -g», «-Ofast/-O3», «-O -g» or simply «-O». My personal plea to open source maintainers: can you please stop overriding CFLAGS and CXXFLAGS variables in your Makefile's / CMakefiles unless you know that a specific combination of the optimisation flags breaks your package. «make» / «cmake» et al allow for the conditional or additive setting of variables (i.e. ?= or += vs :=), so please stick with those and inherent the values when they are present in the environment.
- Tagbert 5y agoThe percentage of apps that are recompiled for AS is now around 68% https://isapplesiliconready.com/statistics https://isapplesiliconready.com/statistics that’s pretty good for a system that has only been available for about a year.