8 ms·
Running a database on EC2? Your clock could be slowing you down
- aarongolliver 8y agoan older discussion of this: https://news.ycombinator.com/item?id=13813079 https://news.ycombinator.com/item?id=13813079
- kalmar 8y agoYup, that's the blog post I mentioned in the intro. Sometimes obsessively reading HN pays off. Even if it's a year later... :-)
- misiti3780 8y agohere is a link that doesnt have the ssl problem: http://archive.is/7zVmu#selection-1533.0-1536.0 http://archive.is/7zVmu#selection-1533.0-1536.0
- wallstprog 8y agoNice article! If you're interested in clocks on Linux, you might also find this article useful (shameless plug): http://btorpey.github.io/blog/2014/02/18/clock-sources-in-linux/ http://btorpey.github.io/blog/2014/02/18/clock-sources-in-li...
- amluto 8y ago> Note that the 100ns mentioned above is largely due to the fact that my Linux box doesn’t support the RDTSCP instruction, so to get reasonably accurate timings it’s also necessary to issue a CPUID instruction prior to RDTSC to serialize its execution. Huh? That’s definitely not true now, and I don’t think it ever was. Linux uses LFENCE or MFENCE, depending on CPU.
- _msw_ 8y agoUsing CPUID as a serializing instruction before RDTSC{,P} is a bad bad thing to do inside a virtual machine on Intel processors. CPUID will cause a VMEXIT, and the CPUID instruction will be emulated. The Intel Software Development Manual Instruction Set Reference gives good guidance on using MFENCE and LFENCE as required. https://software.intel.com/sites/default/files/managed/39/c5/325462-sdm-vol-1-2abcd-3abcd.pdf#page=1667 https://software.intel.com/sites/default/files/managed/39/c5...
- amluto 8y agoLinux has mostly stopped using CPUID to serialize at all. When full serialization is needed, we use IRET now. In the future, we could optimize a bit by writing to CR2, except on Xen.
- abraham_lincoln 8y agoWe set our server 4 hours ahead!
- scarface74 8y agoWhy would you run your own Postgres instance on EC2 within AWS? That kind of defeats the purpose of paying for AWS. Why not use Postgres RDS or Aurora? It makes some sense with Sql Server and Oracle in a few cases because of licensing but hosting your own Postgres instance on AWS is the worse of both worlds -- you're paying more than with a cheaper VPS and you have to do all of the maintenance yourself and not taking advantage of all of things that AWS provides -- point in time restores, easy cross region read replicas, faster disk io (Aurora), etc.
- hamandcheese 8y agoWell, if your business requires running databases really well it might be preferable to run your own. If they were using RDS they would not have been able to get this level of insight into their infra. Maybe RDS is already configured correctly. Maybe not. But at a certain scale not having your hands tied becomes more valuable than having everything taken care of for you.
- arciini 8y agoGiven that they're an analytics company, they probably have 2 problems with using Aurora/RDS. Note that these are educated guesses from their statement: "Heap’s data is stored in a Postgres cluster running on i3 instances in EC2. These are machines with large amounts of NVMe storage—they’re rated at up to 3 million IOPS, perfect for high transaction volume database use cases" 1. Aurora charges per-request ($0.20/million). Given that analytics comes with tons of events and that they wanted servers that have up to 3 million IOPS, it can get pricey fast. 2. RDS has database instances that have SSDs that provide "up to 40,000 IOPS" per instance in their provisioned case, which is probably not enough.
- nil_pointer 8y agoIf you compare the hourly cost for any instance on EC2 vs RDS, it's significantly more expensive for the managed solution (75%+ more), which is to be expected. I know people who roll their own for cost savings.
- plopz 8y ago
- tofflos 8y agoIs AWS NVMe still ephemeral and how do you deal with it? What happens if a machine, or several, reboots?
- cornellwright 8y agoReboots are actually fine as ephemeral data will persist through a reboot on an EC2 instance. Your question is still valid though in case of halting and how you deal with it is specific to your application, but you have to be able to handle all the data on that ephemeral disk disappearing without warning. One way in the case of a database could be a second EC2 instance configured as a read replica in a different AZ.
- kalmar 8y agoBack when we ran the Citus cluster on EBS, we lost some EBS volumes as well. This manifested as disk not responding, followed several days later by an email from AWS with the subject Your Amazon EBS Volume vol-123456789abcdef telling you the disk was lost irrecoverably. But yeah, you need to be ready for your disks to go away no matter where they are: ephemeral, EBS, physical, whatever.
- kalmar 8y agoPost author here. It's ephemeral, yes. It survives reboots, so that's not a problem. It doesn't survive instance-stop, so if a machine is being decommissioned by AWS we do indeed lose its data. As for how we protect against it, the main thing is replication: the data is stored on more than one machine. If we lose a machine for whatever reason, the shards from that machine are copied from a replica to another DB instance.
- _msw_ 8y agoAs local NVMe storage does not have any interaction with the "classic" block device mapping APIs (the storage shows up as a PCI device, the same way that a GPU or FPGA does, and it doesn't matter in any way how the block device mapping is set up), there is no reason to use "ephemeral" to describe it. Said more directly: no, it is not ephemeral. It is local storage that is tied to the life cycle of the instance.
- deleted 8y ago[deleted]
- Johnny555 8y agoAnd, EC2 does not live migrate VMs across physical hosts. I couldn’t find anything explicit from AWS on this, but it’s something that Google is happy to point out. Is it a good idea for a production database to depend on a feature not being used when the vendor hasn't said that they don't or won't use it? They may very well live-migrate when convenient, but just don't expose that functionality to customers since they don't want customers demanding it.
- KayEss 8y agoWith the instance types they're using the migration isn't really an option because the point of the i3 instances is the locally attached NVMe SSD disks where the database files are.
- AbacusAvenger 8y agoThis is exactly why I wrote the "clockperf" tool during the time I was working at AWS: https://github.com/tycho/clockperf https://github.com/tycho/clockperf At the time, we were trying to benchmark disk I/O for new platforms, but we found that things were underperforming compared to the specifications for the hardware. We figured out that fio was reading the clock before/after each I/O (which isn't really necessary unless you really care about latency measurement) and just by reading the clock we were rate limiting our I/O throughput. By switching to "clocksource=tsc" in our fio config, we managed to get the performance behavior we expected.
- logicallee 8y ago>we managed to get the performance behavior we expected. can you put this into roughly quantitative terms? How much of a performance hit did you remove this way?
- AbacusAvenger 8y agoI don't remember the exact numbers (this was 2011), but the overhead of using e.g. CLOCK_MONOTONIC were substantial. Under Xen, the cost of reading CLOCK_MONOTONIC was a few orders of magnitude higher than reading the TSC. I think on Xen PV it was like 500ns per read, while on HVM it was at about 2000ns-3000ns or something like that. I remember with 8 disks that should have been able to do 60K 4K IOPS each (early SSD models), we were capping out at 90K IOPS with all disks in parallel at a queue depth of 32 while reading from CLOCK_MONOTONIC. When we switched to TSC I think we ended up getting around 320K IOPS. Still not perfect, but we were also capped by the particular HBA we chose (which didn't have multiqueue support).