12 ms·
How to sleep at night having a cloud service: common architecture do's
- bcrosby95 7y ago> A 4 9’s means you can only have 6 minutes down a year. 4 9's is 52 minutes of downtime a year. Keep in mind that single region EC2 SLA is only 99.99%. And if you rely on a host of services with an SLA of 99.99, yours is actually worse than 99.99. So if you want to actually get to 99.99, your components have to be better than this, meaning you will have to go multi-region. So achieving this is actually way harder than this simple step.
- ses1984 7y agoIt's still meaningful to discuss 99.99% on top of things that are around 99.99%. For example, let's say you have a service on AWS and all your clients are on AWS. If AWS is down, you are down but so are your clients. But your clients want you to be up 99.99% of the time that AWS is up. As long as both sides are aware of the implications, this is fine. As long as you're within the same order of magnitude, it can make sense. If a customer wanted me to be up 99.99% of the time on top of a service that is only up 99.9% of the time, I would push back.
- gav 7y agoThis is a very salient point. If your service relies on N other services, each with a SLA of 99.99%, the chance of a single request having at least one failure is: 1 - .9999^N Which means if you make 10 requests, you go from 99.99% to 99.9% or from 52 minutes to 8.77 hours of downtime a year. In most cases you're likely to be making a lot more than 10 service calls.
- james_s_tayler 7y agoDepends on if those 9s are in series or in parallel. In series it multiplies to produce lower availability but in parallel they give you higher availability.
- drieddust 7y ago> AWS will use commercially reasonable efforts to make the Included Services each available for each AWS region with a Monthly Uptime Percentage of at least 99.99%, in each case during any monthly billing cycle.... So to achieve 99.99% within a region, every component should have at least 3 nodes and to better it deployment should go multi-region which will escalate the costs quickly. Most application in reality don't even need four 9s so this works b beautifully for everyone. I work in outsourcing industry and in bad old days we had huge penalties and many rounds of explanations even for applications with no redundancy requirements ;). But it's just Amazon credit nowadays and no one blinks and eye so it's win win the all.
- james_s_tayler 7y ago3 nodes of a component in parallel would give you 99.9999% for that component.
- drieddust 7y agoYes but not in AWS land. Committed SLA for availability of entire region is still 4 nines irrespective.
- james_s_tayler 7y agoHmmm. That's good to know. So in that case you have to replicate across three regions to get 6 nines. So one component needs 9 copies running around the world to have 6 nines for the component.
- drieddust 7y agoPreety much. As I said above it works because most internal apps within the Enterprise don't even need 2 nines.
- dshacker 7y agoUpdated, Thanks!
- haolez 7y agoAs a solo founder, I have almost everything mentioned in this article set up, except CI/CD. I can certainly see its value, but being able to easily take down parts of my production system and replace them with instrumented variants is very useful to me when things go wrong. I find that CI usually gets in the way of this. Maybe it's just a bad habit that I need to ditch :)
- dbt00 7y agoDepends on the kind of instrumentation you're talking about. I think metrics should always be collected, and I love having software where reasonable logs are on by default, and unreasonable levels of logging can be enabled/disabled at runtime without toggling.
- dshacker 7y agoI think the first step is having CI but not reacting to it. For a while I shared your view, but after just "enabling it in the background" nowadays it's really useful. Even if it's just an "it compiles" check.
- veeralpatel979 7y ago> being able to easily take down parts of my production system and replace them with instrumented variants is very useful to me when things go wrong Sorry what did you mean by this?
- haolez 7y agoWhen some service is misbehaving, I have a script to take it down and replace it instantly by an instrumented version with more logs and whatnot.
- eropple 7y agoService reconfiguration is a thing and, in systems that handle dependency management well, tends to not be too painful. (I'm in the process of writing a TypeScript library for doing exactly this, designed for NestJS but usable outside of it.)
- swader999 7y agoGood article. Fire drills are worthy of mention. Simulate parts going down, practice recovery.
- fcvarela 7y agoNice article, may I ask what tools you used to produce the illustrations?
- dshacker 7y agoOneNote! :) I'm a SDE in OneNote
- encoderer 7y agoI like guides like this that can help beginners bridge the gap between hobby and professional quality development. I’ll add one more tip, the one I think has saved me more sleep and prevented more headache than any other as I’ve developed a SaaS app over the last 5 years. It’s simple: Handle failure cases in your code, and write software that has some ability to heal itself. Here are a few things I’ve developed that have saved my butt over the years: 1) An application that is deployed alongside the primary application, tails error logs and replays failed requests. (Idempotent requests make this possible) 2) many built-in health checks like checking back pressure on queues and auto-throttling event emitters when queues get backed up 3) Local event buffering to deal with latency spikes in things like SQS. I hope to eventually write more about these systems on our blog but I never seem to find the time
- matwood 7y ago> write software that has some ability to heal itself. IMO, this is the biggest change that helps me sleep at night and starts with treating all servers as cattle. For me it means every server can be rebuilt and deployed at the press of a button which leads to having failed health checks automatically redeploy servers.
- intrepidhero 7y agoThat's good advice but it's hardly simple.
- deleted 7y ago[deleted]
- theshrike79 7y ago4) Sometimes it's better to fail fast and try again than spend time writing extensive error handling and retry logic.
- ambicapter 7y agoBut unless you inform the audience of those times, they will continue being ignorant of when to use your advice, as if you'd never given this advice at all.
- savrajsingh 7y agoJust about everything mentioned here is well-handled by Google App Engine. I still think it’s the way to go for most projects, but I don’t think they’ve marketed themselves well lately. I’m sure there are other good providers too; I don’t see the downside to using PAAS.
- stickfigure 7y agoI came here to post this same thing. You get all of this for "free" from GAE. I built a $100MM company with three engineers on GAE, and we not only slept fine, we'd all go camping together offgrid.
- hckr1292 7y agoGAE is incredible and poorly marketed. Its the only serverless product I know of that allows me to use whatever server framework I want (flask, rails, spring) but be blissfully ignorant of the underlying VMs. I spent a week looking all the other major alternatives out there, and I don't think GAE has any real competitors. Its just a different kind of serverless...in a really good way. Having said that, it has some serious shortcomings: baked in monitoring (at least for Python) is much worse than, say, Datadog + Sentry. Additionally, Google doesn't have any great relational serverless databases (which is what I personally want for a regular webapp) -- they do have some solid non-relational databases. Also, no secret store...its very tricky to securely store secrets inside GAE. To me, the perfect platform for a webapp is GAE + Aurora + some undiscovered secrets store.
- koevet 7y agoWhat's Aurora?
- james_s_tayler 7y agoAurora Serverless database from AWS.
- judge2020 7y agoAre there any particular downsides you have with storing secrets as environment variables? It's working in my app, albeit configuration is done via the web UI [of elastic beanstalk] to keep secrets out of SCM.
- synack 7y agoThis is all good advice for the app tier, but in my experience the most painful outages relate to the data store. Understand your read/write volume, have a plan for scaling up/out, implement caching wherever practical, and have backups.
- dshacker 7y agoI wanted to lay down some of the common things in the app tier, I think data stores get complex really fast really quickly, it's not easy replicating and sharding quickly unless you've got some experience with it under your belt or you use a tool.
- staticassertion 7y agoOne thing missing here is to avoid synchronous communication. Sync comms tie client state to server state; if the server fails, the client will be responsible for handling it. If you use queue-based services your clients can 'fire and forget', and then your error handling logic can be encapsulated by the queue/ consumers. This means that if you deploy broken code rather than a cascading failure across all of your systems you just have a queue backup. Queue backups are also really easy to monitor, and make a great smoke-signal alert. The other way to go, for sync comms, would be circuit breakers. My current project uses queue-based communications exclusively and it's great. I have retry-queues, which use over-provisioned compute, and a dead-letter for manually investigating messages that caused persistent failures. Isolation of state is probably the #1 suggestion I have for building scalable, resilient, self-healing services. 100% agree with and would echo the content in the article, otherwise. edit: Also, idempotency. It's worth taking the time to write idempotent services.
- sbov 7y agoIt's certainly fairly simple to use queues for straightforward, independent actions, such as sending off an email when someone says they forgot their password. It's less obvious to me how your proposal lines up with things that are less so. Such as a user placing an order. So I'm having trouble envisioning how your system actually works. At least in the stuff I work on, realistically, very few things are "fire and forget". Most things are initiated by a user and they expect to see something as a result of their actions, regardless of how the back end is implemented.
- staticassertion 7y ago> It's less obvious to me how your proposal lines up with things that are less so. Usually you have a sync wrapper around async work, maybe poll based. As an example, I believe that Amazon's "place in cart" is completely async with queues in the background. But, of course, you may want to synchronously wait on the client side for that event to propagate around. You get all of the benefits in your backend services - retry logic is encapsulated, failures won't cascade, scaling is trivial, etc. The client is tied to service state, but so be it. You'll want to ensure idempotency, certainly. Actually, yeah, that belongs in the article too. Idempotent services are so much easier to reason about. So, assuming an idempotent API, the client would "send", then poll "check", and call "send" again upon a timeout. Or, more likely, a simple backend service handles that for you, providing a sync API. Going from Sync to Async usually means splitting up your states explicitly. For example: Given two communication types: Sync (<->) Async (->) We might have a sync diagram like this: A <-> B A calls B, and B 'calls back' into A (via a response). The async diagram would look like one of these two diagrams: A -> B -> A or: A -> B -> C Whatever code in A happens after what the sync call would have been gets split out into its own handler. If your system is extremely simple this may not be worth it. But you could say that about anything in the article, really.
- jto1218 7y agoI'd recommend using an APM product off the shelf to get a lot of the mentioned functionality in the article (Monitoring, Tracing, Anomaly Detection). I would definitely _not_ recommend trying to roll all that yourself, unless you have a ton of time and resources. There's a few good ones out there, we use Instana and it's working really well.
- z3t4 7y agoHaving only one mirror is scary. If one goes down, its like murphy's law kicks in. So you want at least 3 things to go wrong in order to take down your system, 2 is not enough. Also have redundancy everywhere if your checker agent stops working for example. You want 2 of everything and at least 3 of those that should never fail.
- rubyn00bie 7y agoYou know the one thing that has helped me out the most, an error reporting service AND then addressing _every_ error. That is to say, my service should emit zero 500 errors. Then my reporting is easy to interpret and consistently meaningful. I don't have to worry about bullshit noise "oh that's just X it does that sometimes." Sleeping at night is a lot easier when you have less keeping you awake.
- enobrev 7y agoI like to keep a slack channel (and saved kibana search) for 500s for this exact reason. System-wide we should have no 500s, and when they happen I like to tackle them immediately. I also have daily reports for other various errors, like caught exceptions, invalid auths, etc just so I can see where things aren't going quite right in case it's indicative of something weird going on.
- jadams3 7y agoThis. I have a really hard time measuring it, but ever since we really worked on error reporting our week-end sleep factor has greatly improved. For a complex system though, don't under estimate how hard this is to do though ... - Every cloud service needs to be routed to a common service - All of your software, every language, even that cool Go experiment - All of the third party software - logs all have to agree on a format, JSON is not always an option. Finally ... justification of time spent fixing things with no observable side effect(s). Most cloud stuff is reliable against first orders of failure and so are tolerant to a lot of stuff, it's designed that way. But once the wheels come off, and they will come off, ... buckle up if you haven't been fixing those errors. If you aren't clean on second order failures, you're in for a rough ride.
- mcintyre1994 7y agoWe use AWS, and one benefit of their hosted ElasticSearch is that they can build you a lambda that syncs Cloudwatch logs to ES, handling a variety of different formats. So we have our beanstalk web requests + some lambda infra + our main web backend etc. all synced to ES with very little effort. You do have the downside that they don’t have eg nicely synced structure, but that also has the upside that the structure is closer to what the dev is used to so nobody ever needs to go back to CloudWatch or any other logs to get more details or a less processed message. The other downside is you have to write a different monitor for each index, though this has the upside that you can also have different triggers per index. In our small team we just message different slack channels which makes for a nice lightweight opt in/out for each error type. It’d definitely be tricky to get everything aligned in eg the same JSON format, but this sort of middle ground isn’t too hard and still has benefits - you just need to be already syncing in any format to CloudWatch - which if you’re in AWS you probably are.
- jturpin 7y agoGood article. I would add one thing to this - pick a database that scales horizontally and is distributed. CockroachDB, Elasticsearch, Mongo, Cassandra/Scylla are all good choices. If you lose one node, you don't have to be afraid of your cluster going down, meaning you can do maintenance and reconfiguration without downtime. If your load is low or bursty you can even get away with running these on some small servers such as t3 (probably minimally t3.larges). Running a cloud managed database is also a good option.
- ummonk 7y agoYes, and together with that, I recommend putting all state in the distributed database (or distributed file storage for large blobs). This allows you to gracefully handle crashes, stop and restart servers, etc. because you don’t lose any state in the process.
- vishaalk 7y agoGreat article Sada :). Hope OneNote is treating you well!
- dshacker 7y agoHey thanks Vishaal! Having fun everyday :) We miss you over here.
- peterwwillis 7y agoI applaud the author for sharing their notes. But also, this is why HN (and general upvote-anything-that-looks-interesting forums) sucks. If you are actually defining architecture, you should not be reading these kind of blog posts. I get that they are interesting to the layman, but so is The Anarchist's Cookbook. Don't make whatever you read in The Anarchist's Cookbook. And I'm crabbing about this because I am easily susceptible to Anarchists Cookbooks. I have had to implement X tech before, and googled for "How do I X", and some blog post came up saying "For X, Use Y". I'm too lazy to read 5 books on the general concept, so I just dive in and immediately download Y and run through the quick-start guide. After spending a while getting it going and getting past the "quick start", I wonder, "Ok, where's the long-start? What's next?" And that doesn't exist. And later, after a lot of digging, it turns out Y actually really sucks. But the blog post didn't go into that. I wasted my time (my own fault) because I read a short blog post. A lot of people live by Infrastructure as Code, and so they will reach for literally anything which has that phrase in its description. But you don't need it to throw together an MVP, and a lot of the IaC "solutions" out there are annoying pieces of crap. I guarantee you that if you pick any of them up, you are in for months of occasionally painful edge cases where the answer to your problem is "You just weren't using it the right way." In reality, if you want to be DevOps (yes, I'm using DevOps as an adjective, ugh) you should probably develop your entire development and deployment workflows by hand, and only when you've accomplished all of the basic requirements of a production service by hand (bootstrapping, configuration, provisioning, testing, deployment, security, metrics, logging, alerts, backup/restore, networking, scalability, load testing, continuous integration, immutable infrastructure & deployments, version-controlled configuration, documentation, etc), then you can start automating it all. If you've done all of these things before, automating it all from the start may be a breeze. If you haven't, you may spend a ton of time on automation, only later to learn that the above need to be changed, requiring rework of the automation.
- dshacker 7y agoYeah, in reality I was weary about adding the links, these solutions are often created after a problem arose in the system. But I wanted to provide a few places to see what's the initial paths for a beginner to find out more. I can't count the times when I wasn't as knowledgeable approached a conversation about Chef and Puppet that didn't make sense to me, even after I read what Chef and Puppet did. It could be, in the other side, that you really don't grasp the need for these technologies until you've had to manually implement one. Things like the UUIDs are basic to hook up to anything. I can often relate this feeling with nutritional advice. "If you want to be more healthy, eat more advocados", if you only eat advocados, but don't understand what's behind it, or the premise behind it, you'll probably get fat. But if someone tells you "Advocados are a good way to supplement your fats without blah blah blah" and you understand that there isn't a unique solution, then you'll probably be healthier.
- jupp0r 7y agoPretty funny that HN traffic seems to have killed the site.
- corentin88 7y agoHaven’t seen anything related to third-parties service that your cloud service relies on. I’m talking mostly about APIs that you can use that might crash at some time. Any recommendations on that part?
- luord 7y agoThis is a great list. I feel a little happy with myself that I knew about most of these. Except for identifying each request, I had never heard about that. It's so simple yet so brilliant, gotta start doing it.