11 ms·
I went through sweat and tears with this on different projects. People wanting to be cool because they use hype-train-tech ending up doing things of unbelievabl
by ghomem 2y ago
I went through sweat and tears with this on different projects. People wanting to be cool because they use hype-train-tech ending up doing things of unbelievably bad quality because "hey, we are not that many in the team" but "hey, we need infinite scalability". Teams immature to the point of not understanding what LTS means have decided that they needed Kubernetes because yes. I could go on.
I currently have distilled, compact Puppet code to create a hardened VM of any size on any provider that can run one more more Docker services or run directly a python backend, or serve static files. With this I create a service on a Hetzner VM in 5 minutes whether the VM has 2 cores or 48 cores and control the configuration in source controlled manifests while monitoring configuration compliance with a custom Naemon plugin. A perfectly reproducible process. The startups kids are meanwhile doing snowflakes in the cloud spending many KEUR per month to have something that is worse than what devops pioneers were able to do in 2017. And the stakeholders are paying for this ship.
I wrote a more structured opinion piece about this, called The Emperor's New clouds:
https://logical.li/blog/emperors-new-clouds/ https://logical.li/blog/emperors-new-clouds/
- princevegeta89 2y agoYour first paragraph resonates strongly with what the folks have done at my startup......lol
- ghomem 2y agoMy thoughts and prayers :-\ Wish you a quick recovery!
- dijit 2y agoI'm with you, but for me Cloud does have one major benefit: If you use it as IaaS, it's a lot quicker to get prototypes working than if you use anything else, including VPS's from other providers. Google Cloud in particular has very few vectors for lock-in, and follows more principle of least surprise. But once you have prototyped, you should ask the question about rebuilding it somewhere that is cheaper. Near infinite scalability of disk drives is nice, and snapshotting, and cloud in general can allow you to extend your prototype into taking production load and allowing you to measure what you will need; but leaning in to "cloud magick" (cloud run, lambdas, etc) will consume almost as much time to learn and debug as just doing it the old school way anyway. In my lived experience.
- ghomem 2y agoI am not against the cloud. VMs are also cloud, unless you run them on your own servers. For instance, the Hetzner Cloud (mostly VMs, plus load balancers and disks) is so cheap and has such a nice CLI API that it competes aggressively with dedicated servers - I would definitely start any with VMs, not with iron. The biggest problem is the so called cloud native stuff which is both more expensive and more complex. There are contexts where it makes sense but for startups they are doing more harm than good.
- finaard 2y agoThing is, by the time the cloud native stuff makes sense most companies are at a scale where it'd be cheaper to just hire a good devops team, and start building your own cloud infra on own hardware.
- ghomem 2y agoProbably so. And that would be likely my approach at such scale. Still, my most benevolent interpretation of current reality is, rather than saying "that cloud native stuff is crap", accepting that there are cases where it may make sense. For instance, large companies might have trouble hiring a good ops team because they have in general trouble hiring and retaining talent (another conversation topic). Ops people are a scarce good because univs do not train people for that and most people prefer coding. I am leaving the work devops out because the market completely perverted its meaning. (my take on the devops funeral: https://logical.li/blog/devops/ https://logical.li/blog/devops/ )
- ghomem 2y agoReference: https://survey.stackoverflow.co/2022/#developer-profile-developer-roles https://survey.stackoverflow.co/2022/#developer-profile-deve... Only around 11% of the whole devs identify as devops specialist or cloud infrastructure engineer. This is why I am saying ops people are a scarce good (unfortunately) from a data driven perspective. Of course my daily life confirms it.
- hello0904 2y agoSerious question for you, why use Docker at all? You can just get rid of the clunky overhead. You mentioned Python backend, so literally just replicate build script, directly in VPS: "pip install requirements.txt" > python main.py" > nano /etc/systemd/system/myservice.service > systemd start myservice > Tada. You can scale instances by just throwing those commands in a bash script (build_my_app.sh) = You're new dockerfile...install on any server in xx-xxx seconds.
- ghomem 2y agoI mentioned Docker because it interests many developers but on VMs that I control I do not need Docker at all. Deploying with Docker provides host OS independence which is nice if you are distributing but unnecessary if the host is yours, running a fixed OS. For Python backends I often deploy the code directly with a Puppet resource called VcsRepo which basically places a certain tag of a certain repo on a certain filesystem location. And I also package the systemd scripts for easy start/stop/restart. You can do this with other config management tools, via bash or by hand, depending on how many systems you manage. What bothers me with your question is Pip :-) But perhaps that is off topic...?
- Gud 2y agoNo, you are tied to docker supported operating systems. Will not run on FreeBSD, for example.
- ghomem 2y agoI'll correct myself: s/host OS independence/a certain level of host OS independence And getting containers to run depends on the OS - if you don't control the host, leads to major ping-pongs. Even within Linux (Ubuntu, Debian, RHEL, etc) when you are distributing multiple related containers there are details to care about, not about the container itself but about the base OS configuration. It's not magic.
- ffsm8 2y ago> No, you are tied to docker supported operating systems No, you're tied to operating systems using a Linux kernel that supports the features necessary for running images.
- globular-toast 2y agoI feel like Kubernetes is always randomly mentioned in rants like this. Instead of saying your hardened VM has Docker you could have just said it has kubelet on it. Then instead of a bunch of ad hoc "docker services" you could pay pennies for a k8s control plane that gives you control over everything on those VMs. I fail to see how your way is anything but worse. The bad cloud infrastructure is when people try to use every single thing AWS sells and their whole infrastructure is at super high levels of abstraction that they could never migrate to another platform. K8s isn't that at all.
- ownagefool 2y agoIn think in either case, if you already have code that's done, using that is going to be less effort than switching. However, I ran kubeadm on a hetzner server and it's just sat chugging along forever basically. I use the cluster to run ephemeral apps where I build and deploy 1 golang service, a couple of node services in about 60 seconds ( with cache, obviously ). As someone old enough and skilled enough to do the same with puppet, why bother when it's simpler easier that even the kids who don't understand TLS can do it with k8s?
- zepolen 2y ago100% best comment in this thread. With k8s you get a way of saying 'WHAT YOU WANT' without 'HOW TO DO IT', and this is applies not only to the actual infra aspect, but the people maintaining it too. Any cloud platform and devops worth their salt can maintain a k8s system. Good luck finding someone to understand what that 'custom Naemon' plugin is doing.
- ghomem 2y ago> Good luck finding someone to understand what that 'custom Naemon' plugin is doing. You Kubernetes people get triggered very easily. I was already lucky to have found several juniors that worked in this kind of thing with minimal training. The 'custom Naemon plugin' is 30 lines of bash and you can adapt it to any monitoring system. Of course this is scary and complicated. I might consider switching to 'Kubernetes operators', which sounds simpler :-)
- a_c 2y agoApart from the operation side, there is a development side parallel too. Two examples that I came across - "Test" mean if it passes on CI, it is good. Failing to run test on local? Who do development on local anyway? - Teams so reliant on "AI" because this is the future of coding. "how to sort a list in python" became a prompt, rather than a lookup on the official documentation.
- hliyan 2y agoI started my career in a world where we did everything using shell scripts running directly on bare metal servers, usually running Solaris, and later SuSe or RedHat. I never understood the "how would you reproduce your setup without Docker (or X, where X is some other technology)". The scripts were deterministic. The dependency versions were locked. The configurations were identical. The input arguments were identical. The order of execution was identical. It all ran on a deterministic computational device. How could it not be reproducible?
- ookblah 2y agoreproducibility isn't just on your deployments, it's for development too. got old REAL fast when your fancy build doesn't work the same on every devs device or some one off issue with how your dev has setup their environment steals hours from everyone. it was a big reason why we moved to containers at the bare minimum, because its quick and easy to spin up and destroy and you are guaranteed what runs locally runs on prod. no more "well it worked on my system!".
- ghomem 2y ago>reproducibility isn't just on your deployments, it's for development too Absolutely. Adhoc configurations should be forbidden! It is easy to ensure dev env reproducibility when you run Linux. If you have config management your devs can have VMs that subscribe to the same exact configuration that the staging prod and dev environments have. They can literally have a deplpyment server in their machine, as a VM. Since the configuration is stored on a server and applied continuously, it is hard to screw it. You can achieve this with Docker as well, if the arrangement is not too complex. The problem, at least in my experience, comes when you start depending on several cloud native components where local emulations are always different from the real cloud env in tiny details that are going to screw the deploys over and over.
- ghomem 2y agoWell that's exactly the point! Creating complex cloud resources with, for instance, Terraform, is less reproducible than a shell script on an LTS system like Ubuntu or RHEL - that's because the cloud provider interfaces drifts and from time to time stops accepting the terraform manifests that previously worked. And to fix it, you have to interrupt your normal work for yet another unplanned intervention in the terraform code - this happened to my teams several times. This does not happen with Puppet + Linux, because LTS distributions have a long release cycle where compatibility is not broken. I tried to explain this topic in the article linked above. Not sure how far I succeeded.
- itronitron 2y agoI can't remember the last time I've seen a position description for a software developer (or anything tech related for that matter) that didn't include a requirement for skills in some cloud related tech. Sometimes the job descriptions are boastful in their reference to those technologies, and other times you can detect some level of despair.
- karmarepellent 2y agoNow I am curious: how do you detect despair regarding cloud tech in job descriptions?
- JamesonNetworks 2y agoI’ve just recently gotten into ansible and find myself building the same thing. I wrote a script to interact with virsh and build vms locally so I can spin up my infra at home to test and deploy to the cloud if and when I want to spend actual money. I’m still very much an ansible noob, but if you have a repo with playbooks I’d love to poke around and learn some things! If not, no worries, I appreciate your time reading this comment!
- zepolen 2y agoHow do you monitor this setup? How do you control access to this setup? How do you deploy on a different provider to Hetzner? How do you access logs on this setup? How do others maintain this setup? How do you run backups? How do you run cron jobs? How do you deal with an offline node? How do you expose a new ingress? How do you provision extra storage on this setup? If any of those is answered with 'something homegrown' or 'just write a script' then you have all the reasons k8s is worth it.
- pella 2y agoHetzner and Kubernetes are not mutually exclusive. - https://github.com/kube-hetzner/terraform-hcloud-kube-hetzner https://github.com/kube-hetzner/terraform-hcloud-kube-hetzne... - https://www.hetzner.com/hetzner-summit https://www.hetzner.com/hetzner-summit --> "Managed Kubernetes Insights and lessons learned from developing our own Kubernetes platform"
- ghomem 2y agoThe questions are short but the answers would be long. Puppet manages all fine grained OS resources (files, dirs, repos, cronjobs, sudo declarations, firewall rules, etc) and you aggregate those resources into classes which are then pushed to different machines. The classes are parametrizable for the differences between systems. If I was to write an idempotent script for each native resource I would finish in some years :-) You chose whatever monitoring system you like the most. For offline nodes you use whatever the level of criticity of your node justifies. This is something people struggle to understand: not every business needs 99.99% uptime. That said, I never had a downtime in Hetzner. On Digital ocean I had one short forced reboot in 4 years. YMMV so protect yourself as much as necessary. Deploying on a different provider than Hetzner is the same as deploying on Hetzner except the part of launching the machine which is trivial to script - the added value is making the machine work and Ubuntu/Debian/RHEL are the same everywhere. You don't have vendor lock in with this. If K8s works for you, enjoy it. Nobody is telling you to stop :-)
- karmarepellent 2y ago> while monitoring configuration compliance with a custom Naemon plugin. While I absolutely agree with you and your approach, would you mind elaborating what kind of configuration compliance you are referring to in this statement? I suppose you do not mean any kind of configuration that your Puppet code produces as that configuration is "monitored", or rather managed, by Puppet.
- ghomem 2y agoI don't mind elaborating - the fact that people are asking me questions reminds me that I need to invest a bit more effort on some articles. This case is actually pretty simple. Puppet applies the configuration you declare impotently when you run the Puppet agent: whatever is not configured gets configured, whatever is already configured remains the same. If there is an error the return code of the Puppet agent is different from that of the situations above. Knowing this you can choose triggering the Puppet agent runs remotely from a monitoring system, (instead of periodical local runs), collecting the exit code and monitoring the status of that exit code inside the monitoring system. Therefore, instead of having an agent that runs silently leaving you logs to parse, you have a green light / red light system in regards to the compliance of a machine with its manifesto. If somebody broke the machine leaving it in an unconfigurable state or if someone broke its manifesto during configuration maintenance you will soon get a red light and the corresponding notifications. This is active configuration management rather than what people usually call provisioning. Of course you need an SSH connection for this execution and with that you need hardened SSH config, whitelisting, dedicated unpriviledged user for monitoring, exceptional finegrained sudo cases, etc. Not rocket science.
- karmarepellent 2y agoThank you for your thorough explanation. Interesting to see that you basically use your monitoring system as a scheduler to run Puppet and it sounds beneficial to closely integrate it with your monitoring to have it all in one place. At my place of work we went the "traditional" way of running Puppet locally. It has been our experience that Puppet failures due to user misconfiguration or some such do not require our immediate attention (e.g. after hours), so we just check Puppetboard a few times per day to identify failing nodes. Another reason why we use Puppetboard to monitor Puppet nodes is that every alert that our Icinga monitoring system produces is automatically interpreted as an incident which needs immediate attention. We are currently in the process of changing that so we are able to process non-critical alerts in a saner way. Anyway, interesting to see how a fellow Puppet user manages their setup. Keep it up!