3 ms·
This article highly underestimates the value of keeping 128b vector performance high. Most code doesn't get recompiled or compiled with the appropriate flags. T
by universal_sinc 3y ago
This article highly underestimates the value of keeping 128b vector performance high. Most code doesn't get recompiled or compiled with the appropriate flags. There is significant overhead involved in supporting 1x512b operations, 2x256b operations, or 4x128b operations per cycle with the same datapaths, forwarding network, and register files. Until 128b vector performance gets deprecated this tension incentives narrow implementations.
- RaisingSpear 3y ago> There is significant overhead involved Current CPUs already do this, and have been doing so for quite some time. And AVX10/128 doesn't alleviate it either.
- Symmetry 3y agoUnless Intel is planning on deprecating SSE then this is already baked in and AVX10/128 existing or not existing won't be changing anything.