4 ms·
Always avoidable if that's a priority - schema changes can be done online in MySQL. Patches can be done on subsets of servers. Erlang even supports hot code rel
by greenleafjacob 10y ago
Always avoidable if that's a priority - schema changes can be done online in MySQL. Patches can be done on subsets of servers. Erlang even supports hot code reloading so that even if you had a single point of failure you can upgrade without losing file descriptors or in memory state. It is a lot simpler if you have the choice though, since you don't have to have multiple versions online at the same time. "Divisions of Ericsson that do [hot code reloading] spend as much time testing them as they do testing their applications themselves." [1]
[1]: http://learnyousomeerlang.com/relups http://learnyousomeerlang.com/relups
- abritinthebay 10y agoAlso immutable deploys.
- icebraining 10y agoThere's no such thing if you have a database.
- abritinthebay 10y agoTrue but that's not the only reason a site would take itself down
- endymi0n 10y agoThere's more and more nuanced reasons actually: 1. Companies don't know how to do the engineering for maximum uptime, like you describe. It's way more complicated than the usual CRUD operations 2. Companies know how to but they decide not to invest this time (we often traded one hour of downtime against 2-3 man-days for preparing online schema changes with nasty and inconsistent backfilling in the early days). And 3. Don't forget disaster recovery. I've seen some of the smartest companies go down for hours due to a DB misconfiguration, or a Rack PSU faulting with only one side of the servers connected, even with a reasonably highly available setup. Stuff like this happens - and then you better have a proper 503 Maintenance page up and running to prevent Google from delisting your site. In this case though, "maintenance" is rather an euphemism :)
- gaius 10y agoCompanies know how to but they decide not to invest this time Cost increases exponentially for diminishing returns once you get into serious availability. For most businesses, the investment in moving from 99.9% to 99.99% or 99.99% to 99.999% uptime just isn't worth it - most customers are quite willing to "try again later" in practice, especially if you give them advance notice or have a regular maintenance slot.
- TheAceOfHearts 10y agoFrom a purely abstract point of view it's probably avoidable, but I'd conjecture many teams don't have the collective knowledge to effectively pull it off. Even if you plan things out carefully, something usually goes wrong :(. It only takes a small oversight to have it come crashing down. I think it's better to let your customers know ahead of time that you'll be performing maintenance, assume that something will go down, but still try to avoid it anyway. MySQL only added support for online schema migrations with 5.6, prior to that you had to use a tool like pt-online-schema-change. I've heard claims (which I haven't verified, so it's entirely possible they are incorrect) of performance issues when performing migrations, which effectively bring the database down. Doesn't RDS sometimes require downtime for maintenance and upgrades? Is there any safe way around that?