4 ms·
AVX512 has a compressed store which might be a bit easier than a normal masked store?
by celrod 5y ago
AVX512 has a compressed store which might be a bit easier than a normal masked store?
- dzaima 5y agocompress has less throughput, and probably a decent bit more latency. It would save offsetting the result pointer by the leading zero count, but I don't think that'd be enough to compensate for the slower compress. I don't know where regular masked store is done in the pipeline though, so maybe it has its own latency comparable to that of compress, but I doubt it.
- celrod 5y agoYeah, 2p5 (two uops on port 5) for a compressed store vs 1p05 (1 uop on either port 0 or port 5) for a masked move. For throughput's sake, shifting the pointer is better.