4 ms·
I have a Dell XPS 13 with a Tiger Lake CPU. Out of curiosity, running the script: > ./two_or_three N = 1000000000, 953.7 MB starting experiments. two
by celrod 6y ago
I have a Dell XPS 13 with a Tiger Lake CPU.
Out of curiosity, running the script:
> ./two_or_three
N = 1000000000, 953.7 MB
starting experiments.
two : 30.6 ns
two+ : 39.6 ns
three: 45.1 ns
bogus 1422321000
This is much slower than the times Lemire reported for the M1. `two+` is 62% of the way between `two` and and `three`, vs 88% for the M1.
EDIT: Adding `-march=native` didn't really change the results, which makes sense given that it's a memory benchmark.
- gspr 6y agoStrange on my XPS 15 7590 (i7-9750H), I get N = 1000000000, 953.7 MB starting experiments. two : 12.8 ns two+ : 13.7 ns three: 19.5 ns bogus 1422321000
- celrod 6y agoYeah, that is strange. Why was it so slow? Trying on a desktop with a 7900X, I get N = 1000000000, 953.7 MB starting experiments. two : 17.7 ns two+ : 19.1 ns three: 26.4 ns bogus 1422321000 This is again close to 50% slower than your time, but nearly twice as fast. I'll try again on the laptop and make sure I don't have other processes running.
- gspr 6y agoStrange. I suppose the compiler shouldn't matter much here, right? At any rate, I'm using GCC 10.2.1.
- celrod 6y agoI'm using gcc 10.2.0. I tried clang 11 and got more or less the same thing, so it doesn't seem to make much of a difference. Neither did messing with flags, like (I tried -fno-semantic-interposition -march=native and a few others).
- celrod 6y agoI just ran it again, and got more or less the same results: N = 1000000000, 953.7 MB starting experiments. two : 29.7 ns two+ : 36.5 ns three: 43.8 ns This surprises me. Normally, it does very well in most benchmarks I run. Looking a little closer at the script, it loads numbers from "random", a vector of 3 million `Int` (this is hard coded, separate from `N`). This vector is about 11.4 MiB. The Tiger Lake CPU has 12 MiB of L3 cache (same as your i7-9750H), so it barely fits. Meanwhile, the L1 cache is 48 KiB and the L2 cache is 1.5 MiB -- huge compared to most recent CPUS, and a lot of benefit in most benchmarks, but at the cost of higher latency. https://www.anandtech.com/show/16084/intel-tiger-lake-review-deep-dive-core-11th-gen/4 https://www.anandtech.com/show/16084/intel-tiger-lake-review... Skylake's L3 latency was 26-37 cycles, and in Willow Cove's (Tiger Lake), it is 39-45 cycles. That difference by itself isn't big enough to account for the difference we're seeing, so something else must be going on.
- nkurz 6y agoThe caching of the random[] array (almost) shouldn't matter, as the access is sequential. I'm wondering if the difference is the the number of active memory channels. How many channels does your respective computers support? Do you have enough RAM installed so all channels are in use? Are you able to do a RAM bandwidth test by some other means to verify? Another possibility is that for some reason the base latency is just different between your machines. A commenter added a pointer-chasing variation of Daniel's test on his blog. Maybe run this to find the full latency and see how the times differ? Finally, there was one more commenter on the blog who reported anomalously fast times on a Windows laptop. It's possible there is a bug with Daniel's time measurements on windows.
- gspr 6y agoI'll run the pointer-chasing version later. Thanks Re timing issues: I am on Linux.