4 ms·
As mentioned by at least one other comment (buried in a thread), in reality, don't do this. If you care about the performance of the loop /that/ much, just hand
by owlbite 5y ago
As mentioned by at least one other comment (buried in a thread), in reality, don't do this. If you care about the performance of the loop /that/ much, just hand code the vectorization/assembly yourself. Use some sort of profiler to make sure it hits peak throughput and declare victory.
- josefx 5y ago> just hand code the vectorization/assembly yourself. Yes, shoot portability in the foot so you can keep your code free of well defined language constructs.
- kaba0 5y agoAs opposed to writing worse performing and less readable code? Then just write it in the simplest way and hope autovectorization will help you.
- josefx 5y ago> As opposed to writing worse performing and less readable code? I can count the number of coworkers I have that have experience with inline assembly on one hand (it is less than 1). Also the first reaction I usually get to vector intrinsics are questions about the wtfness of shuffle instructions, no you can't make that readable without sacrificing performance. > Then just write it in the simplest way and hope autovectorization will help you. Spoiler: It wont't. Compilers often don't have the context and some of the biggest hot spots I had to deal with simply used the single value versions of vector instructions. Or rather: If that had worked I wouldn't be there hand optimizing the code.