4 ms·
Okay that was a bad guess. I see in the previous chapter that the processor is only clocked at 2 GHz, the results still feel a bit low, but certainly way more r
by NohatCoder 5y ago
Okay that was a bad guess. I see in the previous chapter that the processor is only clocked at 2 GHz, the results still feel a bit low, but certainly way more reasonable in that light. I still wonder exactly where the extra clock cycles go.
- dahart 5y agoIf the article was based on a 2GHz processor, then I think I might have been making some bad guesses too. ;) Doesn’t the prefetch solution providing almost 2x mean that on average cache misses in the baseline naive method are costing almost half the cycles? I forgot to say above, but my data point was using a 1GB array, so well above where the author’s cliff dropping off due to being out of cache. I assume the whole reason he’s getting higher perf with arrays < 16MB is because he initialized the array right before computing the prefix sum.