3 ms·
> Citation needed. I've seen people claiming this for years, and I've yet to see a single case where handwritten assembler actually did better than spending the
by cipherboy 6y ago
> Citation needed. I've seen people claiming this for years, and I've yet to see a single case where handwritten assembler actually did better than spending the same amount of effort on speeding up the compiled program (e.g. taking 5 minutes to actually set the right target architecture).
Take some time looking at established open source projects where IBM has heavily invested in optimizing for the POWER architecture. I'm thinking things like glibc, gmp, Golang, openssl, ... The results they get from their hand rolled assembly far exceeds what gcc/llvm spit out. (At least, it did 2-3 years ago.)
Back when the Linux Technology Center (LTC) was a thing, I was fortunate enough to meet many of the individuals working on these projects. They are all wizards. Brilliant.
One of them retired recently and went on to start pveclib to make some of these optimizations in number crunching more accessible to others: https://github.com/open-power-sdk/pveclib https://github.com/open-power-sdk/pveclib
I was once in his office asking about assembly instructions for SHA-3 (a project that sadly didn't go very far) and he could quote useful scalar/vector instructions to me and their timing/latency stats faster than I could find them in the manual.