3 ms·
I ran such a system in prod over 7 years with >5-9s uptime, multiple deploys per day, and millions of users interacting with it. Our deploy scripts were ~10 lin
by hellcow 3y ago
I ran such a system in prod over 7 years with >5-9s uptime, multiple deploys per day, and millions of users interacting with it. Our deploy scripts were ~10 line shell scripts, and any more complex logic (e.g. batching, parallelization, health checks) was done in a short Go program. Anyone could read and understand it in full. It deployed much faster than our equivalent stack on k8s.
k8s is a large and complex tool. Anyone who's run it in production at scale has had to deal with at least one severe outage caused by it.
It's an appropriate choice when you have a team of k8s operators full-time to manage it. It's not necessarily an appropriate choice when you want a zero-downtime deploy.
- freedomben 3y ago> It's an appropriate choice when you have a team of k8s operators full-time to manage it. Are you talking about a full self-run type of scenario where you setup and administer k8s entirely yourself, or a managed system or semi-managed (like OpenShift)? Because if the former then I would agree with you, although I wouldn't recommend a full self-run unless you were a big enough corp to have said team. But if you're talking about even a managed service, I would have to disagree. I've been running for years on a managed service (as the only k8s admin) and have never had a severe outage caused by K8s
- esafak 3y agoIs your short Go program public? I'm curious how you handled progressive rollouts, and automated rollbacks.
- hellcow 3y agoIt isn’t, sadly, but the logic is straightforward. Have a set of IPs you target, iterate with your deploy script targeting each, check health before continuing. If anything doesn’t work (e.g. health check fails), stop the deploy to debug. There’s no automated rollback—simply `git revert` and run the deploy script again.
- esafak 3y agoDid you manually promote deployments from one stage to another? This level of manual intervention is not sustainable if you deploy multiple times a day. How often did you deploy?