3 ms·
HN: "A single server with SQLite is perfect for typical production workloads, it's so robust" Also HN: "Database X can fail in Jepsen tests during partitions,
by weddpros 6y ago
HN: "A single server with SQLite is perfect for typical production workloads, it's so robust"
Also HN: "Database X can fail in Jepsen tests during partitions, it's unusable"
I see HA as a mandatory part of what makes a system production ready...
- iMerNibor 6y agoFor many websites it's perfectly acceptable to go down for a few hours (or however long it takes you to notice, spin up a new vm/server and restore a backup) should stuff go wrong, which is fairly unlikely if you're just running a few servers. It's likely the additional overhead worrying about HA is not worth it, not to mention the real possibility of HA just not working properly in actual failure scenarios
- weddpros 6y agoI agree not every website needs high availability... but I wouldn't bet my $600k/yr business on a single VM's availability and reliability. One day, his server is going to die and be replaced by his host, or his disks will need replacing or a RAID array will fry... It seems like his business could at least afford an HA setup on a cloud provider. Moving to any hosted db with backups, updates and redundancy could be worth it. As for HA failing, it's still less likely than not-HA failing. Also an HA setup allows for maintenance and upgrades without downtime, which is way better than the very common "we don't upgrade if it's working, because it could break".
- WJW 6y agoRunning on a single server does not mean no backups and uptime monitoring is in place. If the hardware fails, you get a ping on the channel of your choice, manually provision a new VM then download the latest sqlite backup from your backup provider. Easy to make a checklist or a script for this, too. Also there is a third option between HA (with its increased cost and complexity) and "we don't upgrade if it's working, because it could break", which is "take the site down for a few minutes, do the upgrade, bring it back up". It's not for every site, but there's a range of sites for which that is fine.
- mytherin 6y agoI'm not sure how those are comparable. Most people don't need extreme HA for their websites. It's totally fine if a standard webshop goes down from 2AM-3AM, or if a personal blog is down for a minute because an S3 VM crashed and needs to be rebooted. Most websites are not Google or Amazon, and gain very little from more 9's past the initial 2~3, which is totally manageable on a single server. Why spend thousands setting up and managing distributed servers if going from 99.0% to 99.999% gains you hundreds? If you need more nines and stronger guarantees than that, absolutely, use distributed solutions, nobody is telling you otherwise. Jepsen tests are valuable once you have decided to use a distributed solution. They showcase problems that would be extremely hard to track down in production settings but can cause serious problems there. If encountered in practice these problems can cause data corruption or other very strange bugs that would take days if not weeks to track down. I wouldn't say that a system with failures in Jepsen tests is unusable, every system has bugs, but if a system has these problems on purpose (e.g. to make it look better on benchmarks) and there is no movement to fix these problems then I would definitely steer clear of that system.
- znpy 6y agoYes, people should evaluate ha needs in terms of business needs: what happens if my website goes down? Will I lose money? How much money? "How much money?" Is probably the most import question, as it actually helps in evaluating if there's a business case for investing into the financial commitment, engineering effort and maintenance burden of an ha solution. On a different note, I've seen mysql-based multi-master setups and it's really a joy to be able to treat a mysql crashing as a minor annoyance instead of a call to fire-fighting.
- BrianOnHN 6y agoGoogle crawls 24/7, so as long as you aren't relying on SEO it's fine. Google doesn't want to send traffic to servers with frequent 500 errors. Edit: there is also more to be said about this. I'd research crawl-budget if you're interested.