7 ms·
Not sure if that’s relevant, but when I do micro-benchmarks like that measuring time intervals way smaller than 1 second, I use __rdtsc() compiler intrinsic ins
by Const-me 1y ago
Not sure if that’s relevant, but when I do micro-benchmarks like that measuring time intervals way smaller than 1 second, I use __rdtsc() compiler intrinsic instead of standard library functions.
On all modern processors, that instruction measures wallclock time with a counter which increments at the base frequency of the CPU unaffected by dynamic frequency scaling.
Apart from the great resolution, that time measuring method has an upside of being very cheap, couple orders of magnitude faster than an OS kernel call.
- sa46 1y agoIsn't gettimeofday implemented with vDSO to avoid kernel context switching (and therefore, most of the overhead)? My understanding is that using tsc directly is tricky. The rate might not be constant, and the rate differs across cores. [1] [1]: https://www.pingcap.com/blog/how-we-trace-a-kv-database-with-less-than-5-percent-performance-impact/ https://www.pingcap.com/blog/how-we-trace-a-kv-database-with...
- Dylan16807 1y agoIf you have something newer than a pentium 4 the rate will be constant. I'm not sure of the details for when cores end up with different numbers.
- quotemstr 1y agoWizardly workarounds for broken APIs persist long after those APIs are fixed. People still avoid things like flock(2) because at one time NFS didn't handle file locking well. CLOCK_MONOTONIC_RAW is fine these days with the vDSO.
- denotational 1y agoSadly GPFS still doesn’t support flock(2), so I still avoid it.
- quotemstr 1y agoDoesn't it? https://sambaxp.org/archive-data-samba/sxp09/SambaXP2009-DATA/Henning_Henkel.pdf https://sambaxp.org/archive-data-samba/sxp09/SambaXP2009-DAT... It would be weird, even for AIX, to support POSIX byte range locks and not the much simpler flock.
- denotational 1y agoIt doesn't, at least on the version I have access to, as it is configured on that cluster. I’m using Linux rather than AIX. fcntl(2) locks are supported (as long as they aren't OFD), but flock(2) locks don't work across nodes.
- toast0 1y agoI think most current systems have invariant tsc, I skimmed your article and was surprised to see an offset (but not totally shocked), but the rate looked the same. You could cpu pin the thread that's reading the tsc, except you can't pin threads in OpenBSD :p
- wahern 1y agoBut just to be clear (for others), you don't need to do that because using RDTSC/RDTSCP is exactly how gettimeofday and clock_gettime work these days, even on OpenBSD. Where using the TSC is practical and reliable, the optimization is already there. OpenBSD actually only implemented this optimization relatively recently. Though most TSCs will be invariant, they still need to be trained across cores, and there are other minutiae (sleeping states?) that made it a PITA to implement in a reliable way, and OpenBSD doesn't have as much manpower as Linux. Some of those non-obvious issues would be relevant to someone trying to do this manually, unless they could rely on their specific hardware behavior.
- RossBencina 1y agoOut of interest, does training across cores result in any residual offset? If so, is the offset nondeterministic?
- wahern 1y agoI was curious myself, poked around, and found some references. But I'm still woefully incapable of answering that with any confidence and don't want to risk saying anything misleading, so here's the code and some other breadcrumbs: 1. Apparently OpenBSD gave up on trying to fix desync'd TSCs. See https://github.com/openbsd/src/commit/78156938567f79506a923cf635bd525907207a76 https://github.com/openbsd/src/commit/78156938567f79506a923c... 2. Relevant OpenBSD kernel code: https://github.com/openbsd/src/blob/master/sys/arch/amd64/amd64/tsc.c#L346 https://github.com/openbsd/src/blob/master/sys/arch/amd64/am... 3. Relevant Linux kernel code: https://github.com/torvalds/linux/blob/master/arch/x86/kernel/tsc_sync.c https://github.com/torvalds/linux/blob/master/arch/x86/kerne..., https://github.com/torvalds/linux/blob/master/arch/x86/kernel/tsc.c https://github.com/torvalds/linux/blob/master/arch/x86/kerne... 4. Linux kernel doc (out-of-date?): https://www.kernel.org/doc/Documentation/virtual/kvm/timekeeping.txt https://www.kernel.org/doc/Documentation/virtual/kvm/timekee... 5. Detailed SUSE blog post with many links: https://www.suse.com/c/cpu-isolation-nohz_full-troubleshooting-tsc-clocksource-by-suse-labs-part-6/ https://www.suse.com/c/cpu-isolation-nohz_full-troubleshooti... 6. Linux patch (uncommitted?) to attempt to directly sync TSCs: https://lkml.rescloud.iu.edu/2208.1/00313.html https://lkml.rescloud.iu.edu/2208.1/00313.html
- triknomeister 1y agoTSC is about cycles consumed by a core. Not about actual time. And so for microbenchmarking, it actually makes sense, because you are often much more interested in CPU benchmarks than network benchmarks in microbenchmarking.
- ainiriand 1y agoYou have to benchmark tsc against a fixed CPU speed, say 1000Mhz, then you have a reliable comparison.
- deleted 1y ago[deleted]
- tonyarkles 1y agoIt was a while ago (2009-10ish) but I ran into an exceptionally interesting performance issue that was partly identified with RDTSC. For a course project in grad school I was measuring the effects of the Python GIL when running multi-threaded Python code on multi-core processors. I expected the overhead/lock contention to get worse as I added threads/cores but the performance fell off a cliff in a way that I hadn't expected. Great outcome for a course project, it made the presentation way more interesting. The issue ended up being that my multi-threaded code when running on a single core pinned that core at 100% CPU usage, as expected, but when running it across 4 cores it was running 4 cores at 25% usage each. This resulted in the clock governor turning down the frequency on the cores from ~2GHz to 900MHz and causing the execution speed to drop even worse than just the expected lock contention. It was a fun mystery to dig into for a while.
- mananaysiempre 1y agoThis does not account for frequency scaling on laptops, context switches, core migrations, time spent in syscalls (if you don’t want to count it), etc. On Linux, you can get the kernel to expose the real (non-“reference”) cycle counter for you to access with __rdpmc() (no syscall needed) and put the corrective offset in an memory-mapped page. See the example code under cap_user_rdpmc on the manpage for perf_event_open() [1] and NOTE WELL the -1 in rdpmc(idx-1) there (I definitely did not waste an hour on that). If you want that on Windows, well, it’s possible, but you’re going to have to do it asynchronously from a different thread and also compute the offsets your own damn self[2]. Alternatively, on AMD processors only, starting with Zen 2, you can get the real cycle count with __aperf() or __rdpru(__RDPRU_APERF) or manual inline assembly depending on your compiler. (The official AMD docs will admonish you not to assign meaning to anything but the fraction APERF / MPERF in one place, but the conjunction of what they tell you in other places implies that MPERF must be the reference cycle count and APERF must be the real cycle count.) This is definitely less of a hassle, but in my experience the cap_user_rdpmc method on Linux is much less noisy. [1] https://man7.org/linux/man-pages/man2/perf_event_open.2.html https://man7.org/linux/man-pages/man2/perf_event_open.2.html [2] https://www.computerenhance.com/p/halloween-spooktacular-day-8-mmozeikos https://www.computerenhance.com/p/halloween-spooktacular-day...
- Const-me 1y ago> does not account for frequency scaling on laptops Are you sure about that? > time spent in syscalls (if you don’t want to count it) The time spent in syscalls was the main objective the OP was measuring. > cycle counter While technically interesting, most of the time I do my micro-benchmark I only care about wallclock time. Contradictory to what you see in search engines and ChatGPT, RDTSC instruction is not a cycle counter, it’s a high resolution wallclock timer. That instruction was counting CPU cycles like 20 years ago, doesn’t do that anymore.
- mananaysiempre 1y ago>> does not account for frequency scaling on laptops > Are you sure about that? > [...] RDTSC instruction is not a cycle counter, it’s a high resolution wallclock timer [...] So we are in agreement here: with RDTSC you’re not counting cycles, you’re counting seconds. (That’s what I meant by “does not account for frequency scaling”.) I guess there are legitimate reasons to do that, but I’ve found organizing an experimental setup for wall-clock measurements to be excruciatingly difficult: getting 10–20% differences depending on whether your window is open or AC is on, or on how long the rebuild of the benchmark executable took, is not a good time. In a microbenchmark, I’d argue that makes RDTSC the wrong tool even if it’s technically usable with enough work. In other situations, it might be the only tool you have, and then sure, go ahead and use it. > The time spent in syscalls was the main objective the OP was measuring. I mean, of course I’m not covering TFA’s use case when I’m only speaking about Linux and Windows, but if you do want to include time in syscalls on Linux that’s also only a flag away. (With a caveat for shared resources—you’re still not counting time in kswapd or interrupt handlers, of course.)
- junon 1y agordtsc isn't available on all platforms, for what it's worth. It's often disabled as there's a CPU flag to allow its use in user space, and it's well know to not be so accurate.
- loeg 1y agoWhat platforms disable rdtsc for userspace? What accuracy issues do you think it has?
- junon 1y agordtsc instruction access is gated by a permission bit. Sometimes it's allowed from userspace, sometimes it's not. There were issues with it in the past, I forget which off the top of my head. It's also not as accurate as a the High Precision timer (HPET). I'm not sure which platforms gate/expose which these days but it's a grab bag.
- loeg 1y agoPersonally I'm not aware of any platform blocking rdtsc, so I was curious to learn which ones do.
- bonzini 1y ago> It's also not as accurate as a the High Precision timer (HPET) This hasn't been true for about 10 years.
- junon 1y agoYou're right, I was thinking about the interrupt precision over the default APIC timer. My point about it being disabled on some platforms has historically been true, however.
- bonzini 1y agoI think you're confusing this and the kernel's blacklisting of the TSC for timekeeping if it is not synchronized across CPUs; but while there's a knob to block userspace's access to the TSC, I am not sure that has been used anywhere except for debugging reasons (e.g. record/replay).
- mrlongroots 1y agoThey could've just used `clock_gettime(CLOCK_MONOTONIC)`
- oguz-ismail 1y ago> I use __rdtsc() compiler intrinsic What do you do on ARM?
- Narishma 1y agoUse a Raspberry Pi or something.
- loeg 1y agohttps://github.com/facebook/folly/blob/main/folly/chrono/Hardware.h#L48 https://github.com/facebook/folly/blob/main/folly/chrono/Har...
- oasisaimlessly 1y agoRead `cntvct_el0` with the `mrs` instruction. [1] [1]: https://developer.arm.com/documentation/102379/0104/The-processor-timers/Count-and-frequency?lang=en https://developer.arm.com/documentation/102379/0104/The-proc...
- JdeBP 1y agoSee manual page and changelog. * https://jdebp.uk/Softwares/djbwares/guide/commands/clockspeed.xml https://jdebp.uk/Softwares/djbwares/guide/commands/clockspee... * https://github.com/jdebp/djbwares/commit/8d2c20930c8700b1786db917a2586bf32acd94f1 https://github.com/jdebp/djbwares/commit/8d2c20930c8700b1786... Yes, 27 years later it now compiles on a non-Intel architecture. (-:
- madog 1y agoJust use gettimeofday/clock_gettime via vDSO. struct timespec ts; clock_gettime(CLOCK_MONOTONIC, &ts); On arm64 it directly uses the cntvct_el0 register under the hood but with a standard/easy to use API instead of messing about with inline assembly. Also avoids a context switch because it's vDSO.
- ethan_smith 1y agoWhile __rdtsc() is fast, be cautious with multi-core benchmarks as TSC synchronization between cores isn't guaranteed on all hardware, especially older systems. Modern Intel/AMD CPUs have "invariant TSC" which helps, but it's worth checking CPU flags first.
- signa11 1y agoI don't think it is (guaranteed to be) synchronized across cores. I might be wrong about that though.
- redleader55 1y agoSuccesive rdtsc calls, even on the same CPU, are not guaranteed to be executed in the expected order by the CPU - [1]. 1 - https://lore.kernel.org/all/da9e8bee-71c2-4a59-a865-3dd6c5c9f092@paulmck-laptop/ https://lore.kernel.org/all/da9e8bee-71c2-4a59-a865-3dd6c5c9...