5 ms·
maybe scale is not needed ,but how do you achieve resiliency with bare metal VMs without adding LB and watchdog layers (which is what k8s is anyway)?
by saargrin 6y ago
maybe scale is not needed ,but how do you achieve resiliency with bare metal VMs without adding LB and watchdog layers (which is what k8s is anyway)?
- erikrothoff 6y agoWe use DigitalOcean's loadbalancer product, we've also tried Cloudflare's for pure HTTP loadbalancing. We also use a single Redis instance for job queues. We use Graphite and Grafana for monitoring system metrics (running a Bitnami Graphite/Grafana instance on AWS because we had credits) And for the rest I guess just keeping the services simple? When we do need to scale up and add a new web server or task runner, it takes about an hour of my day. One thing I realised with going bare-VM is that most services today are insanely stable. MySQL almost never crashes, Redis definitely almost never crashes, Rails/Passenger/Nginx never have any issues. The things that do happen is disks filling up, application bugs causing issues, or actual VM downtime, which is rare but happens when you have 30 VMs. With Docker or K8s it added a super complex layer that is in development and has issues. The 4 months we ran our web servers on K8s, I spent at least 1 month debugging issues that ended at an existing open ticket on Github.
- tlear 6y agoA lot of it been like this for a long time. Postgre, MySQL, Redis, Nginx are bulletproof solid. Sqlite might as well be a hammer and so many businesses could easily run on it, if only there was a way to drop a column :o Amount of $$ that was spent on Docker infra to run couple dozen servers at the last place I used to contract for.. oh and more fun when those machines have GPUs on them(a lot more fun) then they decided to support Singularity as well because.. I don’t know.
- mwcampbell 6y ago> When we do need to scale up and add a new web server or task runner, it takes about an hour of my day. So you manually create and configure your VMs? Do you have some kind of HA for your database? If a VM goes offline, how quickly can you replace it? Maybe you don't need to go back to Docker or Kubernetes, but at least consider using one of the hyperscale cloud providers, with its auto-scaling groups and multiple data centers per region, so you can have a system thatheals itself even while you're asleep or on a plane.
- erikrothoff 6y agoYes, I manually create and configure in the sense that I tweak a number in Terraform templates and manually run the ansible playbook for each new server. It's taken a lot of time to get to that level (I think keeping a setup of bash scripts would suffice in our case...) We run a read-replica on every database, so in case a hardware error occurs on the main database we can manually switch it over. It might mean up to an hour of downtime if the worst happens. Some data loss is OK and can be solved with manual customer support most of the time. It's also a lot cost effective than working towards a 100% SLA. Keeping the read-replicas alive is plenty pain enough! I can't imagine automating everything to auto-heal itself. (Sounds super fun though) Codifying the setup for auto-scaling would be a massive undertaking. Each new change then requires destroying VMs and bringing up new ones. That would then require a k8s-like layer of infrastructure for secrets, DNS, service discovery, not relying on ephemeral storage (which is a lot faster than volumes/block storage). I really love doing ops/devops, and would love to have the perfect setup which is 100% automatic and scaleable. Even now I have to stop myself from spending too much time scriptifying things that can just be run manually.
- saargrin 6y agoand hour of downtime and some data loss are not metrics acceptable to most businesses i know of and what if you have 10 customers joining every day? still gonna be running that ansible manually?
- erikrothoff 6y agoI guess you assumed 1 customer = 1 new server? For enterprise purposes where data siloes are important a different approach definitely makes sense. We have 300+ new users per day, so our manual system scales well.
- saargrin 6y agowell if you rely on SaaS solutions for LB and HA, thats fine less so if you're limited to airgapped/onprem or there are some other security or regulatory considerations
- secondcoming 6y agoAll cloud providers offer LBs with backend instance health checks. Custom scaling rules too.
- mrweasel 6y agoYou should add a load balancer, depending on what you do. Most load balancers will check in on the backend node, and disable them if they fail, rerouting traffic to other nodes. The load balancers we utilize will do failover between them selfs, and do it really well, as in "you don't notice". Many seem to underestimate the stability of modern virtualization, the build in redundancies, fail over feature and the capabilities of load balancers. I would guess that most Kubernetes clusters are built on virtual machines, not physical hardware. Meaning that you just now have layers of redundancy.