10 ms·
Two frequently used system calls are ~77% slower on AWS EC2
- gtirloni 10y agoPrevious related discussion: https://news.ycombinator.com/item?id=13697555 https://news.ycombinator.com/item?id=13697555
- deleted 10y ago[deleted]
- masklinn 10y agoSo… it's not that the syscalls are slower, it's that the Linux-specific mechanism the Linux kernel uses to bypass having to actually perform these calls does not currently work on Xen (and thus EC2).
- yellowapple 10y agoDepends on if you're looking at this from userspace or kernelspace. From the latter, you're spot on. From the former, the headline's spot on.
- masklinn 10y ago> From the former, the headline's spot on. Only if you're using Linux guests and assuming vDSO so not really. The headline made me first go to issues with the host/virtual hardware and some syscalls being much slower than normal across the board.
- drewg123 10y agoAnother option is to reduce usage of gettimeofday() when possible. It is not always free. Roughly 10 years ago, when I was the driver author for one of the first full-speed 10GbE NICs, we'd get complaints from customers that were sure our NIC could not do 10Gbs, as iperf showed it was limited to 3Gb/s or less. I would ask them to re-try with netperf, and they'd see full bandwidth. I eventually figured out that the complaints were coming from customers running distros without the vdso stuff, and/or running other OSes which (at the time) didn't support that (Mac OS, FreeBSD). It turns out that the difference was that iperf would call gettimeofday() around every socket write to measure bandwidth. But netperf would just issue gettimeofday calls at the start and the end of the benchmark, so iperf was effectively gettimeofday bound. Ugh.
- LanceH 10y agoMaybe you could set up some caching.
- drewg123 10y agoCaching is basically what the vdso things do. In my recollection, they basically grab a good time from the kernel occasionally, and then use userspace accessible things like rdtsc() to offset from that authoritative timestamp. So it turns millions of syscalls into one.
- sly010 10y agoFor time?
- LanceH 10y agoSorry, I couldn't resist.
- monktastic1 10y agoYeah, with a TTL, and on each round you just check the time to see if it's expired.
- throwaway2048 10y agoyeah just call gettimeofday() to see if it expired yet.
- OJFord 10y agoObviously you don't use clock-TTL.
- coldtea 10y agoYes, for time too. For one, if you don't need over second precision, they why have some of your servers e.g. ask for the current time thousands of times per second? There are ways to get a soft expiration that don't involve asking for the time.
- jdamato 10y agoAuthor here, greetings. Anyone who finds this interesting may also enjoy our writeup describing every Linux system call method in detail [1]. [1]: https://blog.packagecloud.io/eng/2016/04/05/the-definitive-guide-to-linux-system-calls/ https://blog.packagecloud.io/eng/2016/04/05/the-definitive-g...
- Rapzid 10y agoI will def check that out. Anyone who find that interesting may also enjoy "The Linux Programming Interface" :D
- a_t48 10y agoNitpick - `77 percent faster` is not the inverse of `77 percent slower`. The line that says `The results of this microbenchmark show that the vDSO method is about 77% faster` should read `446% faster`.
- woolly 10y agoShould that not be 346% faster? If A takes 1 second and B takes two seconds, then B is 100% faster than A. So the calculation would be (B/A - 1) * 100. Applying this here gives around 346%. EDIT: B would, of course, take 100% longer than A, rather than be 100% faster.
- mulmen 10y agoHow can something that takes twice as long be faster?
- woolly 10y agoYou're right, of course: hadn't had the morning coffee. It should have been 'takes 100% longer' in the 1 second/2 seconds example. The point I was trying to make is that you have to factor in the initial 100% which doesn't contribute to the final value.
- a_t48 10y ago
- brendangregg 10y agoYes, this is why we (Netflix) default to tsc over the xen clocksource. I found the xen clocksource had become a problem a few years ago, quantified using flame graphs, and investigated using my own microbenchmark. Summarized details here: https://www.slideshare.net/brendangregg/performance-tuning-ec2-instances/42 https://www.slideshare.net/brendangregg/performance-tuning-e...
- aluminussoma 10y agoCan you share if you needed to do anything to deal with time drift issues when using tsc? For my own systems, incorrect timestamps would cause a lot of issues.
- brendangregg 10y agoWell, it's been a few years and we haven't switched it back. :) We have had a number of clock issues, and one of the first things I try is taking and instance and switching it back to xen for a few days, but those issues have not turned out to be the clocksource. Usually NTP. AWS can comment more about the state (safety/risk) of these clocksources (given they have access to all the SW/HW internals).
- bdarfler 10y agoLooks like AWS recommends TSC as well. https://www.slideshare.net/AmazonWebServices/cmp402-amazon-ec2-instances-deep-dive/24 https://www.slideshare.net/AmazonWebServices/cmp402-amazon-e...
- brendangregg 10y agoThis reminds me: I should give an updated version of that talk for 2017...
- amr_01 10y agoPlease do, I would be very interested in this!
- binarycrusader 10y agoI prefer the way Solaris solved this problem: 1) first, by eliminating the need for a context switch for libc calls such as gettimeofday(), gethrtime(), etc. (there is no public/supported interface on Solaris for syscalls, so libc would be used) 2) by providing additional, specific interfaces with certain guarantees: https://docs.oracle.com/cd/E53394_01/html/E54766/get-sec-fromepoch-3c.html https://docs.oracle.com/cd/E53394_01/html/E54766/get-sec-fro... This was accomplished by creating a shared page in which the time is updated in the kernel in a page that is created during system startup. At process exec time that page is mapped into every process address space. Solaris' libc was of course updated to simply read directly from this memory page. Of course, this is more practical on Solaris because libc and the kernel are tightly integrated, and because system calls are not public interfaces, but this seems greatly preferable to the VDSO mechanism.
- jdamato 10y agoThis is precisely what the vDSO does. The clocksources mentioned explicitly list themselves as not supporting this action, hence the fallback to a regular system call.
- binarycrusader 10y agoNot quite; vdso is a general syscall-wrapper mechanism. The Solaris solution is specifically just for the gettimeofday(), gethrtime() interfaces, etc. The difference is that on Solaris, since there is no public system call interface, there's also no need for a fallback. Every program is just faster, no matter how Solaris is virtualized, since every program is using libc. There's also no need for an administrative interface to control clocksource; the best one is always used.
- jdamato 10y agoNot quite. The vDSO provides a general syscall-wrapper mechanism for certain types of system call interfaces. It also provides implementations of gettimeofday clock_gettime and 2 other system calls completely in userland and acts precisely as you've described. Please see this[1] for a detailed explanation. For a shorter explanation, please see the vDSO man page[2]. Thanks for reading my blog post! [1]: https://blog.packagecloud.io/eng/2016/04/05/the-definitive-guide-to-linux-system-calls/#virtual-system-calls https://blog.packagecloud.io/eng/2016/04/05/the-definitive-g... [2]: http://man7.org/linux/man-pages/man7/vdso.7.html http://man7.org/linux/man-pages/man7/vdso.7.html
- xenophonf 10y agoIs this just an EC2 problem, or does it affect any Xen/KVM guest? I ran the test program on a Hyper-V VM running CentOS 7 and got the same result: 100 calls to the gettimeofday syscall. Conversely, I tested a vSphere guest (also running CentOS 7), which didn't call gettimeofday at all.
- officelineback 10y ago>Is this just an EC2 problem, or does it affect any Xen/KVM guest? Looks like it's how the Xen hypervisor works.
- ahoka 10y agoIt is slower because it misses an optimization where you can get the current time without having to enter the kernel. The trick is using the RDTSC instruction, which is not a privileged instruction, so you can call it from userspace. The Time Stamp Counter is a 64 bit register (MSR actually), which gets incremented monotonically. You can get the current time by calibrating it against a known duration on boot or get the frequency from a system table first, then with a simple division and adding an offset. There are sone caveats though, like you have to check if the CPU has an invariant TSC using CPUID and every core has a separate register. I think the problem with XEN is that the VM could be moved across hypervisors or CPUs which would suddenly change the value of the counter. The latter could be mitigated by syncing the TSCs across cores (did I mention that they are writable?) and XEN supports emulating the RDTSC instruction too. I'm not sure how it's configured on AWS, so it may be perfectly safe or mostly safe.
- Twirrim 10y agoIt depends on a number of things. My DigitalOcean instance has this problem. The virtual machine I spun up on Oracle's Bare Metal Cloud platform doesn't (disclaimer, I work for the team).
- andygrunwald 10y agoThis was also presented at the last AWS re:Invent in December. See AWS EC2 Deep Dive: https://de.slideshare.net/mobile/AmazonWebServices/aws-reinvent-2016-deep-dive-on-amazon-ec2-instances-featuring-performance-optimization-best-practices-cmp301 https://de.slideshare.net/mobile/AmazonWebServices/aws-reinv...
- apetresc 10y agoDoes anyone have any intuition around how this affects a variety of typical workflows? I imagine that these two syscalls are disproportionally likely to affect benchmarks more than real-world usage. How many times is this syscall happening on a system doing things like serving HTTP, or running batch jobs, or hosting a database, etc?
- octo_t 10y ago> hosting a database this will very likely be calling time related system calls, especially clock_gettime with CLOCK_MONOTONIC.
- TheDong 10y agoYou can use strace and see! Go to your staging environment, use `strace -f -c -p $PID -e trace=clock_gettime` (or don't use -p and just launch the binary directly), replay a bit of production traffic against it, and then interrupt it and check the summary. HTTP servers typically return a date header, often internally dates are used to figure out expiration and caching, and logging almost always includes dates. It's incredibly easy to check the numbers of syscalls with strace, so you really should be able to get an intuition fairly easily by just playing around in staging.
- deleted 10y ago[deleted]
- JoshTriplett 10y agoFor anyone looking at the mentions of KVM "under some circumstances" having the same issue and wondering how to avoid it with KVM: KVM appears to support fast vDSO-based time calls as long as: - You have a stable hardware TSC (you can check this in /proc/cpuinfo on the host, but all reasonably recent hardware should support this). - The host has the host-side bits of the KVM pvclock enabled. As long as you meet those two conditions, KVM should support fast vDSO-based time calls.
- anonymous_iam 10y agoI wonder if they tried this: https://blog.packagecloud.io/eng/2017/02/21/set-environment-variable-save-thousands-of-system-calls/ https://blog.packagecloud.io/eng/2017/02/21/set-environment-...
- teddyuk 10y agoHow common are get time calls so that they would actually be an issue? I've worked on quite a few systems and can't think of a time where an api for getting the time would have been called so much that it would affect performance?
- jankedeen 10y agoAny application code that includes logging of any sort is going to grab time. All of my code (quite a private set and scientific in nature) calls time() at critical points with identifiers so I can easily investigate issues.
- tyingq 10y agoTimestamped logs, transaction timeouts, http keepalive timeouts, cache expiration/eviction, etc. Apache and nginx for example, both call gettimeofday() a lot. Edit: Quick google searches indicate software like redis and memcached also call it quite often.
- Anderkent 10y agoSo does cassandra.
- chillydawg 10y agoInteresting way to find out the version of the hypervisor kernel. If the gtod call returns faster than the direct syscall for it, then you know the kernel version is prior to that of the patch fixing the issue in xen. I expect there are many such patches that you could use to narrow down the version range of the host kernel. Once you've that information, you may be in a better position to exploit it, knowing which bugs are and are not patched.
- pgaddict 10y agoI wonder why the blog post claims setting clock source to 'tsc' is considered dangerous.
- bandrami 10y agoBecause if the clock rate changes, tsc can become out of sync. https://lwn.net/Articles/209101/ https://lwn.net/Articles/209101/
- pgaddict 10y agoNot really. Recent CPUs (at least those from Intel, which is what EC2 runs on) implement constant_tsc, so the frequency does not affect the tsc. A worse issue is that the counters may not be synchronized between cpus, which may be an issue when the process moves between sockets. But I wouldn't call that "dangerous", it's simply a feature of the clock source. If that's an issue for your program, you should use CLOCK_MONOTONIC anyway and not rely on gettimeofday() doing the right thing.
- blibble 10y agohow does constant_tsc interact with VMs being silently migrated from one physical machine to another?
- kondro 10y agoEC2 doesn't do any kind of live machine migration. The only times a machine may start on a different host is if it is stopped and then started. Even reboots don't allow them to move. You see this a lot when AWS lets you know about maintenance on a physical host and gives you the option to avoid the automated move by doing these steps manually at a time of your choosing before the maintenance window.
- pgaddict 10y agoNot sure, but it can't be better than moving processes between CPUs I guess. Also, does EC2 silently move VMs like this?
- nneonneo 10y agoThe title is misleading. 77% slower sounds like the system calls take 1.77x the time on EC2. In fact, the results indicate that the normal calls are 77% faster - in other words, EC2 gettimeofday and clock_gettime calls take nearly 4.5x longer to run on EC2 than they do on ordinary systems. This is a big speed hit. Some programs can use gettimeofday extremely frequently - for example, many programs call timing functions when logging, performing sleeps, or even constantly during computations (e.g. to implement a poor-man's computation timeout). The article suggests changing the time source to tsc as a workaround, but also warns that it could cause unwanted backwards time warps - making it dangerous to use in production. I'd be curious to hear from those who are using it in production how they avoided the "time warp" issue.
- klodolph 10y ago77% faster is not correct either. "Speed" would probably by ops/s. 4.5x longer = 350% slower.
- stouset 10y agoEven this is confusing as hell. Just say the native calls take 22% of the time they do on EC2. Or that the EC2 calls take 450% of the time of their native counterparts. "Faster" and "slower" when going with percentages are ripe with confusion. Please don't use them.
- derefr 10y ago> Some programs can use gettimeofday extremely frequently This is what's usually considered the "root cause" of this problem, though. It's easy enough, if it's your own program, to wrap the OS time APIs to cache the evaluated timestamp for one event-loop (or for a given length of realtime by checking with the TSC.) Most modern interpreters/VM runtimes also do this.
- westbywest 10y agoOpenJDK has an open issue about this in their JVM: https://bugs.openjdk.java.net/browse/JDK-8165437 https://bugs.openjdk.java.net/browse/JDK-8165437
- revmoo 10y agoInteresting, I ran into this issue in PHP a while back. I found a really basic loop that ran thousands of times per pageload that was running REALLY slowly. Caching the time before running the loop solved the problem entirely.
- known 10y agoJust curious to know the status on Azure;
- amluto 10y agovDSO maintainer here. There are patches floating around to support vDSO timing on Xen. But isn't AWS moving away from Xen or are they just moving away from Xen PV?
- nodesocket 10y agoIf anybody is interested, Google Compute Engine VM's result. blog ~ touch test.c blog ~ nano test.c blog ~ gcc -o test test.c blog ~ strace -ce gettimeofday ./test % time seconds usecs/call calls errors syscall ------ ----------- ----------- --------- --------- ---------------- 0.00 0.000000 0 100 gettimeofday ------ ----------- ----------- --------- --------- ---------------- 100.00 0.000000
- MayeulC 10y agoWasn't a workaround posted for this some time ago, that requires setting the TZ environment variable? https://news.ycombinator.com/item?id=13697555 https://news.ycombinator.com/item?id=13697555 It seems very closely related, unless I am mistaken.
- daenney 10y agoYou are not mistaken in that the topics are (somewhat) related, they all have to do with time. But setting the TZ environment variable doesn't mean your programs don't execute the syscalls discussed in this article. This is about the speed of execution of the mentioned syscalls, which will be called regardless of the TZ environment variable, and how vDSO changes that. However, by setting the TZ environment variable you can avoid an additional call to stat to as it tries to determine if /etc/localtime exists.
- peterwwillis 10y ago> All programmers deploying software to production environments should regularly strace their applications in development mode and question all output they find. Or, instead, you could just not do that. Then you could go back to being productive, instead of wasting time tracking down unstable small tweaks for edge cases that you can barely notice after looping the same syscall 5 million times in a row. When will people learn not to micro-optimize?
- jankedeen 10y agoCrapulent and without merit.
- damagednoob 10y agow