21 ms·
The database servers powering Let's Encrypt
- StreamBright 6y ago>> We currently use MariaDB, with the InnoDB database engine. It is kind of funny how long InnoDB was the most reliable storage engine. I am not sure if MyISAM is still trying to catch up, it used to be much worse than InnoDB. With the emergence of RocksDB there are multiple options today.
- kirugan 6y agoMyISAM days are gone, no one will seriously consider it as suitable engine in MySQL.
- jeffbee 6y agoIt is quite likely unless you've meticulously avoided it that MySQL is using ISAM on-disk temp tables in the service of your queries.
- kirugan 6y agoI didn't know that, thanks!
- fipar 6y agoActually, that shouldn't be the case since 5.7: https://dev.mysql.com/doc/refman/5.7/en/server-system-variables.html#sysvar_internal_tmp_disk_storage_engine https://dev.mysql.com/doc/refman/5.7/en/server-system-variab... (and related: https://dev.mysql.com/doc/refman/5.7/en/server-system-variables.html#sysvar_default_tmp_storage_engine https://dev.mysql.com/doc/refman/5.7/en/server-system-variab...) And on 8 MyISAM is mostly gone, not even the 'mysql' schema uses it. Edit: Originally linked only to default_tmp_storage_engine).
- VWWHFSfQ 6y agoThe only thing I ever used MyISAM tables for was for storing blobs and full-text search on documents. If your data is mostly read only then it's a decent option out of the box. But if you do even mildly frequent updates then you'll quickly run into problems with the table level locking instead of the row level locking offered by InnoDB
- sgt 6y agoI'm guessing someone out there's thinking: Why aren't they hosting in the cloud? The cloud being either Amazon or Azure. Surely nothing else exists. Is it really possible to host your own PHYSICAL machine? Does that count as the cloud?!
- barkingcat 6y agoFor a service like letsencrypt, the independence factor is also a major reason for self hosting. I can forsee letsencrypt in the future going to building their own cloud (on their own physical infrastructure), but speaking as a letsencrypt user of their free certificate program, I would lose respect and interest in their service if they went with an AWS or GCP or Azure approach. The independence from other major players (and the ability of their team to change and move everything about their service, as needed) is one of the reasons I use letsencrypt.
- dingaling 6y agoFunny you mention AWS as they're one of the corporate sponsors of LE. So long as they don't have a viable independent revenue stream they're arguably less independent than commercial CAs.
- deleted 6y ago[deleted]
- nicoburns 6y ago"one of" being the key point here. Let's encrypt has a huge number of sponsors (AWS being only 1 of 9 even if you only count the "platinum level" sponsors), which should allow them to maintain their independence. https://letsencrypt.org/sponsors/ https://letsencrypt.org/sponsors/
- castillar76 6y agoFirst, this made me giggle because I run into that attitude all the time. "You're hosting things on a SERVER? Why would anyone do THAT? Heck, you should be putting everything in serverless and avoiding even the vague possibility that you would have to touch anything so degrading and low-class as an operating system. Systems administration? Who does that?" In all seriousness, however, the decision (likely) has very little to do with that. They're most likely not hosting in the cloud because the current CA/Browser Forum rules around the operation of public CAs effectively don't permit cloud hosting. That's a work in progress, but for the time being, the actual CA infrastructure can't be hosted in the cloud due to security and auditability requirements.
- multifascia 6y agoAs someone unfamiliar with db management, is it really less operational overhead to have to physically scale your hardware than using a distributed option with more elastic scalability capabilities?
- stu2010 6y agoRelational databases enable some very flexible data access patterns. Once you shard, you lose a lot of that flexibility. If you move away from a relational model, you lose even more flexibility and start having to do much more work in your application layer, and usually start having to use more resources and developer time every step of the way. The productivity enabled by having one master RDBMS is a big deal, and if they can buy commodity servers that satisfy their requirement, this seems like a fine way to operate.
- hinkley 6y agoIf I had a billion dollars, I'd put a research group together to study the prospects of index sharding. That is, full table replication, but individual servers maintaining differing sets of indexes. OLAP and single request transactions could be routed to specialized replicas based on query planning, sending requests to machines that have appropriate indexes, and preferably ones where those indexes are hot.
- ddorian43 6y agoThe problem is the network. You need billion dollars to fix the network so it's as fast as local ram/nvme.
- zinekeller 6y ago... and even if you somehow solved that, the law of physics hits you hard. Latency can be a real performance killer generally and is doubly true in database-type computing.
- 6y ago
- ryanworl 6y agoWhat are they storing on this server that requires 150Tb of storage and millions of IOPS?
- dewey 6y ago> What exactly are we doing with these servers? Our CA software, Boulder, uses MySQL-style schemas and queries to manage subscriber accounts and the entire certificate issuance process.
- jeffbee 6y agoThere's nothing in that sentence that implies they'd need even 100 IOPS, much less 20 million.
- gbrown_ 6y agoThe post doesn't specify requirements or application level targets for performance. They show a couple of good latency improvements but don't describe the business or technical impact. The closest we get is this. > If this database isn’t performing well enough, it can cause API errors and timeouts for our subscribers. What are the SLO's? How was this being met (or not) before vs after the hardware upgrade? There's a lot of additional context that could have been added in this post. It's not a bad post but instead it simply reduces down to this new hardware is faster than our old hardware.
- stefan_ 6y agoWhat exactly needs to be stored once the certificate is created and published in the hash tree? It seems like the kind of data that possibly needn't be stored at all or onto something like Glacier for archival.
- Conclusionist 6y agoGoing to guess it's for OCSP responses.
- 6y ago
- cbg0 6y agoUnless I misunderstood something, it seems they have a single primary that handles read+write and multiple read replicas for it. It shouldn't be too difficult given the current use of MariaDB to start using something like Galera to create a multi-master cluster and improve redundancy of the service, unless there are some non-obvious reasons why they wouldn't be doing this. I think I also see redundant PSUs, would be neat to know if they're connected to different PDUs and if the networking is also redundant.
- lykr0n 6y agoThat's still a very common pattern if you need maximum performance, and can tolerate small periods of downtime. When designing systems, you have to accept some drawbacks. You can forgo a clustered database if you have a strong on call schedule, and redundancy built in to other parts of your infrastructure. Galera is great, but you lose some functionality with transactions and locking that could be a deal breaker. And up until MySQL 8, there were some fairly significant barriers to automation and clustering that could be a turn off for some people. Everything has it's pros and cons.
- jabberwcky 6y agoMulti-master hardly comes for free in terms of complexity or performance, you're at the mercy of latency. Either host the second master in the same building, in which case the redundancy is an illusion, or host it somewhere else in which case watch your write rate tank Asynchronous streaming to a truly redundant second site often makes more sense
- birdman3131 6y agoHow well would same city with fiber between work?
- jarym 6y agoJust goes to show how much a single SQL server can scale before having to worry about sharing and horizontal scaling
- ed25519FUUU 6y agoThis is especially true with your own hardware. Trying this kind of thing in the cloud is usually prohibitively expensive.
- jbverschoor 6y agoIt’s doesn’t have to. Unless you’re conditioned to believe that aws is cheap
- deleted 6y ago[deleted]
- ed25519FUUU 6y agoI'm curious what the cost would be to run this type of hardware at any cloud vendor? Does it even exist?
- jeffbee 6y agoThat's the wrong way to think about the cloud. A better way to think about it would be "how much database traffic (and storage) can I serve from Cloud Whatever for $xxx". Then you need to think about what your realistic effective utilization would be. This server has 153600 GB of raw storage. That kind of storage would cost you $46000 (retail) in Cloud Spanner every month, but I doubt that's the right comparison. The right math would probably be that they have 250 million customers and perhaps 1KB of real information per customer. Now the question becomes why you would ever buy 24×6400 GB of flash memory to store this scale of data.
- whitepoplar 6y agoNot completely comparable, but Hetzner offers the following dedicated server that costs 637 euro/month (maxed out): - 32-core AMD EPYC - 512GB ECC memory - 8x 3.84TB NVMe datacenter drives - Unmetered 1gbps bandwidth
- ed25519FUUU 6y agoWhat a great read. I think the authors here made great hardware and software decisions. OpenZFS is the way to go, and is so much easier to manage than the legacy RAID controllers imho. Ah, I miss actual hardware.
- ganoushoreilly 6y agoI enjoyed it as well, i'm also appreciative that they shared their configuration notes here. I've been running multiple data stores on ZFS for years now and it's taken a while to get out of the hardware mindset (albeit you still need a nice beefy controller anyway). https://github.com/letsencrypt/openzfs-nvme-databases https://github.com/letsencrypt/openzfs-nvme-databases
- gautamcgoel 6y agoCan you explain the advantages of OpenZFS over other filesystems? I know FreeBSD uses ZFS, but I never really understood how it stacks up relative to other technologies...
- tyingq 6y agoI was, long ago, an old-school Unix sysadmin. While I was technically aware of how powerful smallish servers have become, this article really crystallized that for me. 64 cores and 24 NVME drives in a 2U spot on a rack is just insane compared to what we used to have to do to get a beefy database server. And it's not some exotic thing, just a popular mainstream Dell SKU. If you price it out on Dell's site, you get a retail price north of $200k. That is really what made it clear for me. That you could fit $200k+ worth of DIMMS, Drives, CPUS into a 2U spot :)
- vmception 6y agoI have a motherboard from 2012 and I just put 2x 8TB NVMe SSDs on it, on a PCIe 2.0 x16 slot Works great. The PCIe card itself has 2 more slots for SSDs The GPU is on the 2.0 x8 slot because they don't really transfer that much data over the lanes. I honestly didn't realize PCIe was up to 4.0 now, and I am pushing up against the limits of PCIe 2.0 but it still works! And I’m “only” at the limits, and its only a limit when I want faster than 3,000 megabytes per second, which is amazing. Granted, this would have been considered a good enthusiast motherboard in 2012. Buying new but cheap is the mistake.
- foota 6y agoWhat drives did you get? I think you need PCI 4 to stress most SSDs these days?
- deleted 6y ago[deleted]
- vmception 6y agoI have a hunch that the pcie card itself is most important as it is doing bifurcation. So each drive acts like it has its own slower (but fast enough) pcie slot, and then the raid0 combines the bits back to double the performance. Could be wrong but I get 2,900 megabytes per second transfers from RAM to disk and back. And this is PCIe 2.0 x16 so maybe if you want , 3,000, 4,000 or 6,500 megabytes per second then I have nothing to brag about. I’m pretty amazed though and will be content for all my use cases.
- baby 6y agoSo scaling up instead of scaling out. I’m not sure if it’s a viable strategy long term, at the same time we probably don’t want a single CA to handle too many certificates?
- theandrewbailey 6y agoI'm not aware of any other CA giving out free certificates to anyone. I know that some other providers/hosts will do free certificates, but only to their users (last time I checked).
- thamer 6y agoScaling up means each query is faster (3x in this particular case). Scaling out means they can support more clients/domains (more DB shards, more web servers, more concurrency, etc). These are two distinct axes that are not incompatible with each other.
- baby 6y agoNot always true though, scaling out can make your queries faster by alleviating the load
- fabian2k 6y agoI'm curious why they didn't go with the larger 64 core Epyc. I mean it's double the cost, but I suspect that the huge amount of NVMe SSDs is by far the largest part of the cost anyway. And it seems like CPU was the previous bottleneck as it was at 90%.
- jaas 6y agoWe didn't go with the 64-core chips because they have significantly lower clock speeds. Dual 32-core chips give us plenty of cores while keeping clocks higher for single-threaded performance. You are correct that the price of the CPUs is almost irrelevant to the overall cost of a system with this much memory and storage. We were picking the ideal CPU, not selecting on CPU price.
- fabian2k 6y agoThanks for the answer. I would have guessed that the higher core count outweighs the lower frequence for database usage, but obviously I don't know the details. I think the 90% CPU usage graph just made me nervous enough to want the biggest possible CPU in there.
- _joel 6y agoooi how is the NUMA on that setup? Does it still use QPI or is there a newer technology now (I've been out of this space for a few years now)
- sradman 6y ago> We can clearly see how our old CPUs were reaching their limit. In the week before we upgraded our primary database server, its CPU usage (from /proc/stat) averaged over 90% This strikes me as odd. In my experience, traditional OLTP row stores are I/O bound due to contention (locking and latching). Does anyone have an explanation for this? > Once you have a server full of NVMe drives, you have to decide how to manage them. Our previous generation of database servers used hardware RAID in a RAID-10 configuration, but there is no effective hardware RAID for NVMe, so we needed another solution... we got several recommendations for OpenZFS and decided to give it a shot. Again, traditional OLTP row stores have included a mechanism for recovering from media failure: place the WAL log on separate device from the DB. Early MySQL used a proprietary backup add-on as a revenue model so maybe this technique is now obfuscated and/or missing. You may still need/want a mechanism to federate the DB devices and incremental volume snapshots are far superior to full DB backup but placing the WAL log on a separate device is a fantastic technique for both performance and availability. The Let's Encrypt post does not describe how they implement off-machine and off-site backup-and-recovery. I'd like to know if and how they do this.
- snuxoll 6y agoMySQL backup is more or less a “solved” issue with xtrabackup, and I assume that’s exactly what they are using with a database of this size.
- e12e 6y ago> The Let's Encrypt post does not describe how they implement off-machine and off-site backup-and-recovery. I'd like to know if and how they do this. The section: > There wasn’t a lot of information out there about how best to set up and optimize OpenZFS for a pool of NVMe drives and a database workload, so we want to share what we learned. You can find detailed information about our setup in this GitHub repository. points to: https://github.com/letsencrypt/openzfs-nvme-databases https://github.com/letsencrypt/openzfs-nvme-databases Which states: > Our primary database server rapidly replicates to two others, including two locations, and is backed up daily. The most business- and compliance-critical data is also logged separately, outside of our database stack. As long as we can maintain durability for long enough to evacuate the primary (write) role to a healthier database server, that is enough. Which sounds like traditional master/slave setup, with fail over?
- gbrown_ 6y agoThis is a very high level overview and ideally I would have liked to have seen more application level profiling, e.g. where time is being spent (be it on CPU or IO) within the DB rather than high level system stats. For example the following. > CPU usage (from /proc/stat) averaged over 90% Leaves me wondering exactly which metric from /proc/stat they are refering to. I mean it's presumably its user time, but I just dislike attempts to distill systems performance into a few graph comparisons. In reality the realized performance of a system is often better described given a narrative describing what bottlenecks the system.
- rantwasp 6y agogets huge server. does not properly resize the jpeg on the page (it’s 5megs in size and i see it loading). we don’t al have 5TB of ram you know
- avipars 6y agothose EPYC CPUS are really epic
- polskibus 6y agoI wonder how much do they save by not going public cloud?
- tclancy 6y ago+10% for the proper use of decimated
- MayeulC 6y agoYou mean, that one? https://en.wikipedia.org/wiki/Decimation_(Roman_army) https://en.wikipedia.org/wiki/Decimation_(Roman_army) Then it would be 90 ms -> 81 ms, not 90 ms -> 9 ms. The way I see it, at least. With proper decimation, 90% of what was there remains. ("removal of a tenth", as wikipedia puts it).
- tclancy 6y agoOh man, you’re totally right. I ... uh, blame somebody else.
- uncledave 6y agoI’d like to understand more about the workload here. Queries per second, average result size, database size, query complexity etc.
- Proven 6y agoYes, there's no requirement analysis whatsoever. Let's buy the biggest servers we can. Because PCI lanes.
- 1MachineElf 6y agoI'm thankful for their OpenZFS tuning doc which they developed as part of this server migration: https://github.com/letsencrypt/openzfs-nvme-databases https://github.com/letsencrypt/openzfs-nvme-databases The one thing that I get hung up on when it comes to RAID and SSDs is the wear pattern vs. HDDs. Take for example this quote from the README.md: We use RAID-1+0, in order to achieve the best possible performance without being vulnerable to a single-drive failure. Failure on SSDs is predictable and usually expressed with Terabytes Written (TBW). Failure on spinning disk HDDs is comparatively random. In my mind, it makes sense to mirror SSD-based vdevs only for performance reasons and not for data integrity. The reason is that the mirrors are expected to fail after the same amount of TBW, and thus the availability/redundancy guarantee of mirroring is relatively unreliable. Maybe someone with more experience in this area can change my mind, but if it were up to me, I would have configured the mirror drives as spares, and relied on a local HDD-based zpool for quick backup/restore capability. I imagine that would be a better solution, although it probably wouldn't have fit into tryingq's ideal 2U space.
- hedora 6y agoSSD’s still fail, just not often. State of the art systems keep ~1.2 copies (e.g. 10+2 raid 6) on SSD, and an offsite backup or two. The bandwidth required for timely rebuilds is usually the bottleneck. These systems can be ridiculously dense; a few petabytes easily fits in 10U. With that many NAND packages, drive failures are common.
- walrus01 6y ago> Failure on spinning disk HDDs is comparatively random. comparatively, yes, but when averaged out over a large number of hard drives it definitely tends to follow a typical bathtub curve failure model seen in any mechanical product with moving parts. https://www.itl.nist.gov/div898/handbook/apr/section1/gifs/bathtub2.gif https://www.itl.nist.gov/div898/handbook/apr/section1/gifs/b... early failures will be HDDs that die within a few months of being put into service in the middle of the curve, there will be a constant steady rate of random failures towards the end of the lifespan of the hard drives, as they've been spinning and seeking for many years, failures will increase.
- 6y ago
- louwrentius 6y ago@jaas I wonder how many servers you have and in particular how a failure of a system would play out in your situation. The latter is always interesting to me, building things is easy. Building rock solid reliable things is seriously hard work and often difficult. Would make a great blog post....
- lisper 6y agoI don't understand why they are trying so hard to avoid sharding. It seems to me that this is a perfect example of an "embarassingly parallel" problem for which sharding would be borderline trivial. What am I missing?
- amelius 6y agoWell, their current solution is still a lot simpler, I suppose. With multiple servers there's a lot of extra administrative stuff you need to deal with.
- njharman 6y agoRack space ain’t free
- driverdan 6y agoThe cost of 4U vs 2U is trivial compared to the cost of this hardware.
- deleted 6y ago[deleted]
- top_sigrid 6y agoIt still works like this and they have plenty of headroom with the new solution. Chances are good that when they need to upgrade this solution, technology will have advanced far enough as well. Guessing what their performance demands are (far, far more reads than writes) this seems to work fine for them, so why make it more complicated? They do talk about read-only replicas, so they have some distribution.
- rawoke083600 6y agoSharding is like amputating a limb ! It's works really well if the limb is infected with say a flesh-eating bacteria ! But you keep sharding as a last last resort ! It's hard going back
- JulianMorrison 6y agoIt's also an example of a "mustn't fail or you break the internet" problem and a "lots of people with nation state resources have a reason to fuck with us" problem. They are prioritising simplicity as a means to security, and that makes sense to me.
- ksec 6y ago>Intel® SSD DC P4610 Interesting they decide to put in PCIE 3.0 NVMe SSD instead of PCI-E 4? Imagine having 24x Intel Optane [1]. PCI-E 5 is actually just around the corner. I would imagine next time Let's Encrypt could upgrade again and continue to use a Single DB Machine to Serve. [1] https://www.servethehome.com/new-intel-optane-p5800x-100-dwpd-ssd-dominates-pcie-gen4/ https://www.servethehome.com/new-intel-optane-p5800x-100-dwp...
- zinekeller 6y agoDirty secret when it comes to current PCIe 4 drives: they have optimised sequential RW speed too much that they have forgotten random RW speed (which is the main driver in databases).
- rkagerer 6y agoI'm more interested in how they used ZFS to provide redundancy. I always thought ZFS was optimized for spinning platters with SSD's used for persistent caching. In this scenario they used it to set up all their SSD's in mirrored pairs then stripe across that. No ZIL. They've tweaked a few other settings as well [1]. I'd be curious to see more benchmarks and latency data (especially as they're utilizing compression, and of course checksums are computed over all data not just metadata like some other filesystems). [1] https://github.com/letsencrypt/openzfs-nvme-databases https://github.com/letsencrypt/openzfs-nvme-databases
- dan-robertson 6y agoZFS was started around 2001 when SSDs weren’t really a thing. It’s goals were, amongst other things, to manage multiple volumes (providing redundancy), to be reliable (most filesystems aim for this), and to support cheap snapshotting. The last two we’re supposed to come from being copy-on-write and this model was an advantage when SSDs became popular as it worked a bit better with their semantics.
- WrtCdEvrydy 6y agoZFS is optimized for storing data, you can ZIL into an SSD but you can also ZIL into something faster.
- mark242 6y agoStopping by to say, 9ms API response time is just ridiculously quick. You're starting to run into the laws of physics and client proximity to the datacenter where those machines live. That's a pretty amazing feat. I would assume the next step for scaling is getting those read replicas deployed across the world in order to cut down on RTT.
- shitloadofbooks 6y agoWhy would they need to bother when it's mostly machines talking to machines? Certbot doesn't care that it took 90ms vs 9ms.
- walrus01 6y agoThe 9ms is one indicator that the new hardware platform has ample extra capacity for future growth in load and traffic, it probably won't need to be replaced or upgraded for some years.
- jart 6y ago9ms 50%ile node latency is good, but that number is normally 1ms for big internet services. See https://dl.acm.org/cms/attachment/3658918e-7081-4676-beec-aac65159ccf9/t1.jpg https://dl.acm.org/cms/attachment/3658918e-7081-4676-beec-aa... Mission critical stuff goes even faster like BigTable which has numbers 4x better than the figure.
- rkagerer 6y agoWhat form factor are those NVMe drives and how are the connected? I see cables, so I'm assuming they're not all plugged straight into their own PCIe slot. Are there a bunch of M.2 headers on the motherboard?
- Rafuino 6y agoAlmost definitely U.2 drives that slot into the front of the server. Here's a cartoony view of what that looks like (on a P5800X Optane SSD): https://www.servethehome.com/new-intel-optane-p5800x-100-dwpd-ssd-dominates-pcie-gen4/ https://www.servethehome.com/new-intel-optane-p5800x-100-dwp...
- zerd 6y agoU.2. The drives have PCIe lanes connected via the blue cables to the motherboard, which is then wired to the CPU. Here's a review that goes into details on the hardware: https://www.servethehome.com/dell-emc-poweredge-r7525-review-flagship-dell-dual-socket-server-amd-epyc/ https://www.servethehome.com/dell-emc-poweredge-r7525-review...
- fomine3 6y ago"servers" in title but it seems that it's single. Is it correct? (Just curious, I'm non native for English)
- madsushi 6y agoOne primary, which they replicate to other servers in other locations. https://github.com/letsencrypt/openzfs-nvme-databases https://github.com/letsencrypt/openzfs-nvme-databases
- Proven 6y ago> We have a number of replicas of the database active at any given time, and we direct some read operations to replica database servers to reduce load on the primary. No shared storage, no storage efficiency, and RF2 replication with mySQL on top of that... Ouch. And it's completely unclear why NVMe was necessary in the first place. Are they using more than 5% of its performance? Instead they talk about PCI lanes and whatnot.
- farseer 6y agoLets for the sake of argument assume that Lets Encrypt is a malicious actor. Can they easily compromise the security of the websites using their certificates?
- muldvarp 6y agoYes. Not just websites using their own certificates though. As a certificate authority they can create certificates for arbitrary domains. There are however a few countermeasures against illegitimate certificates such as certificate pinning and certificate transparency.
- godojo 6y agoYes, but they better not get caught. It's a trust based model. It's actually Internet stack developers/packagers (anything from protocols in OS or libraries to browsers and devices) that are trusting Let's Encrypt among other certificate authorities.
- Havoc 6y agoSurprised something that key to the internet runs off a single server.
- OliverJones 6y agoIt would be interesting to see your database schema, Let's Encrypt!
- the_biot 6y agoSo they've completely ignored the lesson Google taught us all 20+ years ago -- lots of cheap servers scales better than an ever more expensive big iron SPOF setup -- and which we're now all using to scale the internet. It's unbelievable to see a design like this in 2021. What a stellar bit of incompetence.