3 ms·
I think the better question is why use a single bit or byte for a single bool when sum types exist? You're pulling in an entire cache line (64 bytes) with any
by Veliladon 2y ago
I think the better question is why use a single bit or byte for a single bool when sum types exist?
You're pulling in an entire cache line (64 bytes) with any load. Why would you not just turn the bool into sum type that carries the payload with them? That way you can actually use the rest of the cache line instead of loading 64 bytes to work with a single bit, throwing the other 511 bits in the trash, and then doing another load on top for the data.
It's even worse when you do multithreading with packed bools because threads can keep trashing each other's cache lines forcing them to wait for the load from L3, or worse, DRAM.
- sph 2y agoIn fact in some archs it might be even faster to make bool word-sized (i.e 8 bytes on 64 bit) if unaligned loads and stores are slower or disallowed.
- Veliladon 2y agoIt'd have to be a really old one. Almost every CPU with an L1 cache loads in cache lines and loads don't have to be cache line aligned because it's supposed to be transparent to the programmer. So there's little to no penalty on modern architectures for non-aligned words. I know there hasn't been one on ARM or Intel for at least a decade. If you try to load and use consecutive words in a hot loop that are, say, 72 bytes apart you'll see a huge performance drop but that's more because you're only using 1/8th of the cache line and not because they're unaligned.