7 ms·
Scaling SQLite to 4M QPS on a Single Server (EC2 vs. Bare Metal)
- justinclift 6y agoTitle should probably say "(2018)". Also, it's already been submitted a bunch of times: https://news.ycombinator.com/item?id=23070888 https://news.ycombinator.com/item?id=23070888 https://news.ycombinator.com/item?id=22856746 https://news.ycombinator.com/item?id=22856746 https://news.ycombinator.com/item?id=20778477 https://news.ycombinator.com/item?id=20778477 https://news.ycombinator.com/item?id=16118776 https://news.ycombinator.com/item?id=16118776
- chrisweekly 6y agoImpressive! Thanks for sharing: it's well-written and detailed, with meaningful data and clear dataviz for comparison.
- xwdv 6y agoI know people are itching to jump to the conclusion that bare metal is ultimately way cheaper than cloud, but really, it isn’t. At the end of the day, we want operating expenses, not capital expenses. And if you purchase a ton of bare metal and suddenly your business needs change it will not be easy or quick to move that much metal or install new metal in place. Stick to the cloud.
- adrianN 6y agoWhy not both? Most businesses have a fairly stable baseload need for computing. Replacing EC2 instances with bare metal to satisfy that seems to make sense to me.
- christophilus 6y agoEh. Most businesses don’t really need big scale. I build exclusively in the cloud, but the app I maintain could easily run on a single $2k server, with the possible exception of image resizing.
- BubRoss 6y agoYou would have to be resizing a tremendous amount of images to overwhelm a $2k server.
- paulryanrogers 6y agoOr a very inefficient implementation.
- ies7 6y agoFor comparation, 2 years ago we build thumbor-like apps in python using vibora and opencv. A single aws m8 large($80/month?) can serve hundred millions images resizing in a month. At that time we build it with only 2 people, copy pasting code from internet in a few days and vibora’s github readme also state that its in alpha release. With $2K server, you need to cheat to make those inefficient implementation :) Side note: Several weeks later I found out about thumbor and replace our ‘emergency work’ with dockerized thumbor I also tinkered with thumbot in aws lambda. But the cost raised to $500/month. I think we did something wrong with aws rekognition.
- maallooc 6y agoThere's a high chance that running 100 times of your app's load won't even use 10% of the $2k server.
- xwdv 6y agoYou don’t need a server period for image resizing, that should really be a lambda function in AWS. Super cheap.
- eloff 6y agoThe cost per query difference of cloud versus bare metal, in the article is over ten times. That kind of margin gives a lot of room for error. You could literally over provision by an order of magnitude and come out on top. There's plenty of hidden costs though of doing this yourself that need a certain level of scale to pay off. Like people to manage the hardware, that's a big one. Still renting bare metal servers from e.g. Packet could make much more sense even for small companies than using AWS.
- zkirill 6y agoAre there any services that help you buy the bare metal outright, set it up at a colo, and maintain it for you?
- watsocd 6y agorenting small servers...much more sense even for small companies No. At one time over 10 years ago, I had a dedicated server with one of the hosting companies. They charged me $225/month. I moved to AWS and my cost dropped to ~$30/month. Even now, 10 years later, my total AWS bill for primarily EC2 and RDS services that runs an application that pays my bills is ~$75 per month. If my company suddenly went viral and had a bunch of new customers, I could scale up instantly.
- eloff 6y agoI'm not talking about hobby projects, which also aren't a great fit for AWS. Use digital ocean, vultr, or linode. I'm talking about startups with over $10k/year in spend.
- panarky 6y ago> we want operating expenses, not capital expenses Ultimately it's all capex. The choice is whether the capital asset is on your balance sheet or someone else's. If someone else is tying up their capital in a depreciating and inflexible asset, they'll demand to be paid well for giving you the flexibility to avoid all of that. So you can't really say opex is always better than capex, since it depends on the price you're willing to pay for that flexibility.
- ashtonkem 6y agoUltimately whether or not a company should be using bare metal vs. the cloud is actually a financial decision, not a technical one. There are some technical decisions that can be made to reduce the risk of changes and/or a mixed use case, but ultimately the difference for most businesses is going to be up to the CFO.
- yayr 6y agoThe article makes a good point, that it is not only a financial decision (b/s vs pnl, risk premium etc) but as well an architectural and functional one. No CFO usually has a good enough understanding of the intricacies of those risks.
- ashtonkem 6y agoThat’s fair, although wise architectural choices can help isolate the technical and financial decisions from each other. For example I think this is the strongest argument for k8s; it gives you the ability to isolate your compute from the underlying infrastructure.
- jjeaff 6y agoYou are creating a false dichotomy. If you aren't a mature company like expensify that can pretty accurately predict your needs, then you can rent bare metal for nearly as cheap as buying it. Probably cheaper if you don't have any special needs and commodity hardware will do.
- watsocd 6y agoIt's not just the server cost that matters when you compare the cloud to a bare-metal server. What about: Networking equipment Backup power - UPS and generators Internet backbone - Speed as well as reliability Cooling and other infrastructure costs Then there is capital lock in risk. Once you buy the above hardware, you are probably stuck with it for a few years. If you have stable load now for 25 big bare-metal servers, it may be cheaper to go bare metal. Otherwise, there is a lot more involved in the decision than going to dell.com and pricing a server.
- toast0 6y agoBare metal is a continuum too. You don't have to pick between a full cloud deploy and building your own datacenter from scratch. You can get bare metal servers from a host on monthly or hourly terms. You can get a 1/3rd rack cabinet at a colo facility and bring your own router, and use the colo's IPs You can get a cage or a room at a colo and arrange for connecticity to your various upstreams and peers with your own ASN and IPs. Only after all this is exhausted would I move towards and owned and operated datacenter where all the pieces are under company control and responsibility. (Although, if you rely on it, you're responsible for it, regardless of if you actually control it)
- watsocd 6y agoThat is all true but then you need to pay for the services and equipment as well as have the expertise on staff to configure and manage the equipment. As far as reliability, you can't compare almost any COLO location with an AWS data center.
- ericd 6y ago>As far as reliability, you can't compare almost any COLO location with an AWS data center. Why not? We had much less downtime than AWS' NoVa DC while coloing, and I don't think our DC was exceptional. AWS is probably much more competent, but they're also trying to build something much more complex than your average small-scale colo setup.
- 6y ago
- brandmeyer 6y agoIn the list of features which are somewhat custom to their use case, they listed "Disable POSIX advisory locks". Turns out this is available stock right now as the "unix-none" VFS implementation. We're using it in an almost-but-not-quite-POSIX RTOS that happens to be missing advisory locks. See also https://sqlite.org/vfs.html https://sqlite.org/vfs.html
- gigatexal 6y agoMaking random() deterministic defeats the purpose of random no?
- gopalv 6y agoThe randomization they are working around is most likely > WHERE indexedColumn > RANDOM() LIMIT 10 That can be bound on RANDOM() performance or we can rely on a seeded "non-random" PRNG to produce a random selection vector without running random() for every row. If that has a syscall into /dev/urandom for every row, then that sucks for sure & is not representative of the actual intention. In Apache Hive, we have a similar argument about unix_timestamp(), which was made deterministic, but stateful (MapReduce failure tolerance demands that if a task fails while running a query like > unix_timestamp() + n, the next attempt will produce an identical output, even-though some time has passed between the attempts). So for that, we do the unix_timestamp() replacement per-query, instead of per-row. I assume that's what was done here.
- rakoo 6y agoA deterministic random() is good for reproducing tests. What you want is a way to use it to generate a pseudorandom output based on "more random" input, such as /dev/urandom.
- nijave 6y agoThey're comparing bare metal to a virtual machine here, aren't they? Of course I'd expect some pretty noticeable performance differences there. Not sure they were offered at the time of the post but AWS also has bare metal that seems like it'd but a much more fair comparison Cloud VM vs on-prem bare metal doesn't seem like a very fair comparison
- jjeaff 6y agoWhy not? As long as you are comparing your total cost all in on both services. But did they even say on-prem? I was assuming colocation.