4 ms·
Could you please list some resources that could help a complete n00b like me start from somewhere wrt resilient and self healing infra?
by krtkush 6y ago
Could you please list some resources that could help a complete n00b like me start from somewhere wrt resilient and self healing infra?
- fxtentacle 6y agoAvoid Java and clouds, use raid and monit. Buy much more memory and storage that you think you need so that you'll have a safety buffer.
- omneity 6y agoThe specific tools we use might not apply to you (the backend is a cluster), but happy to share a few ideas: 1- Use a scheduler that autorestarts: systemd, pm2, nomad, ... (we use nomad) 2- Setup healthchecks to detect when your app is not behaving correctly even if it's still running (for example some exception crippled the program). An HTTP healthcheck is an endpoint (for example /health) that returns a 200 status code when everything is fine. If the endpoint is down or returns something else, the service is not considered healthy and the service is restarted (you can limit the number of restarts when errors cannot be solved with a restart) * Systemd supports socket based healthchecks * pm2 doesn't have built-in support for healthchecks at all but there are some npm modules for that * Nomad does HTTP healthchecks (through consul, not alone) * GCP and AWS (and others) support healthchecks at the level of your server and can restart the entire server when the healthcheck goes wrong 3- Monitoring & alerts: I'll cut to the chase and tell you that honestly the best monitoring solution that worked for us is the built in one from our cloud provider (you still need to setup the agent in your server). 3rd party managed solutions are expensive, and I don't want to self deploy something so critical and add to the complexity of our infra. The main idea in monitoring is not just to be alerted when your servers are down, but to detect issues before they become critical. Common issues like disk or CPU at 70%... 4- High availability: Here be dragons put a load balancer in front of 3 (or more 2n+1) servers, all running the same copy of your app. Make sure your app is stateless! There are risks of race conditions, stale data ... so try to explore the other options first I hope these pointers will help you sleep better at night! You can read more about these topics and look for the tools that match your stack :)