3 ms·
Which is exactly why I said you block the computation to stay at a low cache level. With SIMD loads and stores I don't think this matters quite as much as you s
by mlochbaum 3y ago
Which is exactly why I said you block the computation to stay at a low cache level. With SIMD loads and stores I don't think this matters quite as much as you suggest, even without blocking. It's pretty much only arithmetic that can saturate L1. I timed the BQN compiler on various files (some old version of itself, repeated). For 18K it runs at 21.4MB/s; for 1.7M, 16.5MB/s; for 17M, 12.0MB/s. So even when the source won't fit in L3 (mine's 8MB) the degradation is under a factor of 2 (and of course the compiler makes no consideration of cache, who writes a megabyte of BQN?).