7 ms·
Don't forget "pets vs cattle", thinking of servers as ephemeral and working towards quickly being able to scale up/down based on demand. So often I see people "
by kator 4y ago
Don't forget "pets vs cattle", thinking of servers as ephemeral and working towards quickly being able to scale up/down based on demand. So often I see people "lift and shift" from a dedicated server model into the cloud and never convert their pets into cattle. This reduces flexibility later, not to mention makes it harder to respond to patching needs, scaling, and moving to optimize latency or costs.
- candiddevmike 4y agoCitation needed? There are tradeoffs to both, one is not always better than the other.
- paulryanrogers 4y agoWhat's the advantage of pets? Simplicity?
- gtirloni 4y agoA 128-core 4TB "pet" is much easier to manage than the equivalent Kubernetes cluster with same capacity. Not saying it's the way to go for every situation.
- paulryanrogers 4y agoDo you mean bare metal?
- jdub 4y agoIt should still not be administered as a pet. In fact, it's even more important when you have a single instance of some importance to make it entirely rebuildable and replaceable.
- vanviegen 4y agoIf you manage to run your entire service on a single (heavy weight) server, uptimes will usually be excellent. Reliability for a single server is high. It gets lower for each server you add, unless you're adding redundancy. But that adds complexity, and that hurts reliability too. My point: a single dedicated server may be a more reliable, simpler and much cheaper solution than the cloud provides.
- deleted 4y ago[deleted]
- tekla 4y agoReliability means jack shit when I have a system that auto-recovers on any failure. I literally do not care about any particular server. If it fails, it kills itself, and a new server comes up and serves traffic. No alerts, no pages, nothing.
- astrange 4y agoWhat happens if they fail in a loop and you end up in “livelock”?
- hackmiester 4y agoBut you are introducing the complexities of managing a fleet. This may or may not be an improvement in overall complexity, cost, or reliability. This decision should be made per-project, and neither option is a good fit for all projects.
- LawTalkingGuy 4y ago> the complexities of managing a fleet These things are inherent complexities in any situation with backup/DR. You've already got to figure this out. When one machine is degraded and another is trying to recover from db logs, that's a cluster. > This decision should be made per-project This is one of the least project-dependent things ever, akin to source control. If you don't have scripted installs you don't know what you're running and you can't reliably test it or upgrade it. You don't need to run a cluster, or have auto-scaling groups, just because you can.
- tekla 4y agoIts not a fair comparison to compare to k8s. I've replaced pets with ASGs that do nothing but sit there unless the EC2 instance fails. Its amazing because you no longer have to baby the pet, it just recovers in a nice clean state, same massive instance.. Pets are a artifact of laziness.
- gtirloni 4y ago> Pets are a artifact of laziness I was with you until this unnecessary point.
- ranger207 4y agoA cattle server needs no special maintenance. Yes, if you only have one server then by definition it'll require special maintenance, but that maintenance should be as generic as practicable so that if you ever need to set up a second one then you don't need to do anything special. That's the distinction between pets and cattle, not the tooling: you can configure cattle by hand if you have a repeatable and non-special cased setup script to follow, and you can use Ansible or Puppet or whatever to set up the unique environment your pet server needs. The important thing is that the setup process, maintenance, etc is standardized and documented. Automated is a bonus that should be relatively easy if you're doing it right
- hiAndrewQuinn 4y agoThere might be an earlier source, but I first ran across the pets versus cattle nomenclature in Tom Limoncelli's _Handbook of System and Network Administration_ - which is a really, really good read for anyone going deep into ops space (like a cloud engineer should be).
- r3trohack3r 4y agoAs an ex-FAANG engineer, this is FAANG advice. Pets are just fine. Most companies arent FAANG and don't need that class of solution. An R620 plugged into a switch in a colo, a bash script via cron, or a cloudflare worker are just fine for a lot of use cases. The only time it stops being fine is when you can't afford to do your pet -> cattle migration as you scale up. But I don't think this is a common death for companies. If you call "cattle" a cloudflare worker or lambda function - fine. But when we are talking about multiple redundant servers with load balancing across them, you really need to justify the cost of that vs the value you squeeze out. Sometimes you're squeezing the juice out of the rind.
- mr_toad 4y agoTreating servers as disposable is about more than just scale. It helps avoid creating snowflake servers, makes DR more predictable, and makes creating dev environments much easier.
- viraptor 4y agoAlso, just knowing what you're running at the time is great. Changing things manually means you'll forget some of them when you need to do a bigger change. Or when you're seeing up server no.2. We've got great tooling these days, so whether you're going with nixos, docker, puppet, shell scripts for installation, or something else, the cattle approach gives you benefits.
- fragmede 4y agoI'd say that until you're FANG-level or at least hockey-sticking, it's actually totally okay for some of your things to be pets and that cattleizing literally everything is actually a premature optimization. API boxes which you have 100 of? Absolutely cattleize. That one bespoke server that runs that one service that is totally weird? Just let it be weird for a while. Work on more important things.
- jon-wood 4y agoI think there’s degrees of cattleisation, I have in the past deployed crappy vendor software which needs a unique license key for each running instance. We discussed some sort of license service which would allow keys to be checked out when a container is started, but in the end settled on an auto scaling group with 1 instance for each running container with the key in an environment variable. That got us the comfort of knowing if the host machine died in the night the task would be rescheduled without a bunch of extra engineering to be able to scale arbitrarily when we only needed a couple of instances running.
- voiper1 4y agoSome replies are saying this is only for "as-scale/FAANG". It may only be absolutely necessary there, but it's helpful even for smaller folks. Over the years, even Debian LTS goes out of support and new features and software should be installed. There's moving systems, doing restores, things breaking and wanting to "reset" to a known working state. Any time you can do something simple with docker or even just (short) step-by-step build scripts, that's a huge win. I have playbooks for deploying a system, but with npm installs, bower installs, secrets to be hand copied from multiple places, etc, it feels more like pets and it's NOT simple to deploy.