5 ms·
What's the precision of these ns level measurements?
by s800 6y ago
What's the precision of these ns level measurements?
- mhh__ 6y agoThe answer to that is usually very context dependant, and on what you're measuring. As long as you use a histogram first and don't blindly calculate (say) the mean it should he obvious. Two examples( that are slightly bigger than this but the same principles apply): If you benchmark a std::vector at insertion, you'll see a flat graph with n tall spikes at ratios of it's reallocation amount apart, and it scales very very well. The measurements are clean. If, however, you do the same for a linked list you get a linearly increasing graph but it's absolutely all over the place because it doesn't play nice with the memory hierarchy. The std_dev of a given value of n might be a hundred times worse than the vector.
- CyberRabbi 6y agoClock_gettime(CLOCK_REALTIME) on macos provides nanosecond-level precision.
- geocar 6y agoI seem to recall OSX didn't used to have clock_gettime, so it's news to me that it even exists -- I might have been away from OSX too long. Is there any performance difference between that and mach_absolute_time() ?
- saagarjha 6y agoIt's new in macOS Sierra. I believe mach_absolute_time is slightly faster but not by much–both just read the commpage these days to save on a syscall.
- lilyball 6y agoIt was added some years ago, and I believe mach_absolute_time is actually now implemented in terms of (the implementation of) clock_gettime. The documentation on mach_absolute_time now even says you should use clock_gettime_nsec_np(CLOCK_UPTIME_RAW) instead. macOS also has clock constants for a monotonic clock that increases while sleeping (unlike CLOCK_UPTIME_RAW and mach_absolute_time).
- saagarjha 6y agoNot yet, at least :) _mach_absolute_time: 00000000000012ec pushq %rbp 00000000000012ed movq %rsp, %rbp 00000000000012f0 movabsq $0x7fffffe00050, %rsi ## imm = 0x7FFFFFE00050 00000000000012fa movl 0x18(%rsi), %r8d 00000000000012fe testl %r8d, %r8d 0000000000001301 je 0x12fa 0000000000001303 lfence 0000000000001306 rdtsc 0000000000001308 lfence 000000000000130b shlq $0x20, %rdx 000000000000130f orq %rdx, %rax 0000000000001312 movl 0xc(%rsi), %ecx 0000000000001315 andl $0x1f, %ecx 0000000000001318 subq (%rsi), %rax 000000000000131b shlq %cl, %rax 000000000000131e movl 0x8(%rsi), %ecx 0000000000001321 mulq %rcx 0000000000001324 shrdq $0x20, %rdx, %rax 0000000000001329 addq 0x10(%rsi), %rax 000000000000132d cmpl 0x18(%rsi), %r8d 0000000000001331 jne 0x12fa 0000000000001333 popq %rbp 0000000000001334 retq
- Skunkleton 6y agoThat may be the result of inlining clock_gettime, though that would imply a pretty different implementation from the one I am familiar with. AFAIR on x86 a locked rdtsc is ~20 cycles. So to answer the gp question, it has around a precision in the few nanoseconds range. Accuracy is a different question, IE compare numbers from the same die, but be a little more suspicious across dies. No clue how this is implemented on the M1, or if the M1 has the same modern tsc guarantees that x86 has grown over the last few generations of chips.
- saagarjha 6y agoYeah, clock_gettime is somewhat more complicated than this. If anything, it might have an inlined mach_absolute_time in it…
- lilyball 6y agoI was actually misremembering a bit. Sufficiently old versions of mach_absolute_time used a function called clock_get_time() on i386 (if the COMM_PAGE_VERSION was not 1). This changed in macOS 10.5 to a tiny bit of assembly that just read from _COMM_PAGE_NANOTIME on i386/x86_64/ppc (the arm implementation(!!) triggers a software interrupt). The i386/x86_64/ppc definitions were also copied into xnu. For the next few years it kind of bounced back and forth between libc and xnu, and the routine was complicated by adding timebase conversion as needed. And at some point arm support was added back (it vanished when it first went to xnu), but this time using the commpage if possible. As for M1, I assume it's using the arm64 routine in xnu, which can be found at https://opensource.apple.com/source/xnu/xnu-7195.50.7.100.1/libsyscall/wrappers/mach_absolute_time.s https://opensource.apple.com/source/xnu/xnu-7195.50.7.100.1/.... As for clock_gettime_nsec_np(), at least as of Big Sur, it's in libc¹ instead of xnu and defers to mach_continuous_time()/mach_continuous_approximate_time()/mach_absolute_time()/mach_approximate_time() for the CLOCK__RAW[_] clocks. And clock_gettime() for those clocks is implemented in terms of clock_gettime_nsec_np(). ¹https://opensource.apple.com/source/Libc/Libc-1439.40.11/gen/clock_gettime.c.auto.html https://opensource.apple.com/source/Libc/Libc-1439.40.11/gen...
- varjag 6y agoStill single digit level nanosecond precision sounds marginal. 1ns = 2 clock cycles at 2GHz.