13 ms·
Genuine question: how many HN readers log on to boxes with user accounts that belong to humans, where some state may have been mutated? My experience of the la
by EngineerBetter 9y ago
Genuine question: how many HN readers log on to boxes with user accounts that belong to humans, where some state may have been mutated?
My experience of the last five years is so heavily weighted to (effectively) immutable infrastructure that checking to see who had been on a box hadn't event crossed my mind.
- sneak 9y agoYou are at the right hand side of the curve. Professionals do it that way, yes. Many smaller or less-well-managed organizations have a lot more warts.
- subway 9y agoeven on immutable hosts it isn't surprising to see somebody has ssh'd in for debugging purposes. they might be attaching a debugger, taking a core dump, or doing a number of other things that could cause a drastic perf hit while not actually mutating the host's state otherwise.
- jameshart 9y agoWhy take the risk of continuing to run a tainted host, though, if you can just tear it down and spin up a new, clean, untouched one? I think there’s another level we need beyond treating servers as pets or cattle, which is treating servers as wild animals; after you’ve captured one and interacted with it, you’ve doomed it because now it has the scent of humans on it.
- subway 9y agoSure, but that does nothing to prevent the person who invoke the debugger from inadvertently triggering prod alerts, possibly encouraging more folks to ssh in to see what's going on.
- gowld 9y agoGood hiring and training practices prevent people from invoking the debugger on live prod machines
- subway 9y agoThere are times when it has to happen. It certainly shouldn't be an every day occurrence, but even in the best environments, there are times when SHTF only under a production workload, and you need an understanding of why yesterday.
- avip 9y agoThe age of immutable inf. did not completely free us from occasional hotfixing production issues, because immutable deployment takes time, and if it's downtime you're accountable (as sysadmin/devops/whatever).
- gpmcadam 9y agoA stitch in time, saves `kill -9`
- Rowern 9y agoSo I am not the only one to have a slow eployment because re-building AMI and re-provisioning everything takes quite some times. I also had a difficult time to explain that a prod deploy of 30min (image creation, deploy with blue green) is normal for this kind of inf... Did you face the same thing?
- scrollaway 9y agoRebuilding AMIs is not a thing you should be doing every deployment. Sounds like you are on AWS so use proper containers on ECS or EBS. Docker itself caches pretty aggressively. Decompose your projects as well so that the independent parts build and deploy without rebuilding everything else in the project that hasn't changed. At the end of the day, if you're on continuous deployment, a commit should be rebuilding only what it touches. We have 4 min long deployments + 1.5 min tests and I definitely don't think we're optimizing aggressively.
- devonkim 9y agoThings get awkward when versioning standards require all application components to have the same version because several version numbers running around is cognitive overhead that engineers can’t afford in many situations. With more than about 10 components I’ve usually seen it turn into “deploy 10 services that have 10 changes, 9 of which have one commit that bump a version number up.” Many places still keep producing very stateful software (sometimes even very much by choice) that is better off managed through Puppet / Chef rather than an immutable containerized approach. If your software needs to take an hour and a half to shutdown, for example, you have to get a bit creative with your deployment strategies.
- markwillis82 9y agoWe have an immutable hosting platform with ansible, but some clients are hosted on their own boxes so we still end up logging into their hosting to sort out their issues (normally they have more problems than we do). But they aren't interested in investing in more infrastructure when a single LNMP server does the job.
- Fradow 9y agoOn my personal dedicated server, I don't have any kind of "modern" things, I do everything the old way, so I do log manually. It's a feature, it's a way for me to learn old-school sysadmin (and I have so little things on here anyway automation is not needed). For my startup, I often run bash on Heroku because I do migrations manually (again, it's a feature, I'm too inexperienced to have automated migrations that work everytime, I prefer to be already on it if it breaks). Sometime when something breaks I'll also poke around the filesystem (which is a copy, so no fear to break anything). Basically, I'd say the smaller your team, your uptime requirements and your traffic is, the less you need automation, and the more you are susceptible to login directly to a box (I combine all of that: team of 1, no uptime requirement, not enough traffic to even max the most basic server).
- raverbashing 9y agoNot to mention that while configuration management "magic" has taken over (Puppet/Chef - or even something higher level) you still need to set up those and you still need to know where to look when things go wrong.
- majewsky 9y agoCounterpoint: Even on my private VPS, I have all the configuration as code. It gives me peace of mind knowing that when a server comes crashing down for whatever reason, I can reinstall it and bring all services back online with not more than one hour of time invested. Time is precious. (Also, when someone asks me "how do you configure X", I can just link them to the corresponding place in my system configuration repo on Github.)
- closeparen 9y agoOur containers are immutable, but the bare metal hosts they run on aren’t.
- movedx 9y agoNo, but the "bare metal hosts" should be configured through configuration management and all logs sent "off site". I've been in the same boat and never had to log into a bare metal box. K8s and Ansible will do that for ya.
- freehunter 9y agoI work on those "off site" machines that collect your logs and I'm sitting on one as root right now. Ansible doesn't get you everywhere.
- movedx 9y agoYikes! We use time series databases, so we simply export the data from them to analytical software or our local machines. InfluxDB is a simple command to extract what I need and then work with it from there.
- closeparen 9y agoPuppet deploys a variety of host-level agents like L7 routing, dynamic config updaters, the scheduler agent (of course), metrics, logging, and tracing collection. Usually they work, but it’s necessary more often than you’d think to find out why one has fallen over and fix it, and to run diagnostic tools to investigate performance anomalies.
- deleted 9y ago[deleted]
- mrweasel 9y agoI do that daily. Most of the services we run consist of 1 - 8 application servers ( often closer to 2 than 8), so things like Docker doesn't make much sense. Even though we of cause try to automate as much as possible, using things like Ansible, we often log in to servers directly to verify changes. Database servers are normally manually managed, to some extend, so we'll always login directly. When I'm on-call and get an alarm on a server, checking that colleagues aren't logged in is normally the first thing you do. If a service fails it's more often than not someone doing maintenance, and they just forgot to tell you, or disable monitoring. And as someone else pointed out, also check disk usage. Hacker News can be a little blind to the fact that most software projects are rather small and modern servers are really powerful. It's actually pretty rare that someone manages enough infrastructure for a single service that logging isn't a viable option.
- bauerd 9y agoI share your viewpoint and am currently thinking through a similiar small-scale deployment. How do you handle logging? For single boxes I'd just use logwatch and friends, but for aggregating log output (both system/application logs) from a small-ish Consul cluster I feel the options are complete overkill. I've had a look at ELK, time-series backed stuff (e.g. Prometheus), but all I really want is log aggregation with search capability and optionally Regexp-based alerting. It seems to me Logstash provides what I want, but I sure as hell won't run 3 JVMs and a Redis instance to aggregate logs. tl;dr How to handle log aggregation within a small-scale cluster without losing your sanity?
- kaitnieks 9y agoFor simple, non-critical stuff this is the simplest solution I've come up with - if you don't need real-time aggregated logs (i.e. you only want to archive them) and you have log rotation configured, you can simply use cron with "aws s3 sync" or rsync for the logs folder.
- greenleafjacob 9y agoCheck out papertrail.
- drinchev 9y agoMostly `root` with ssh-keys for authentication. Some provisioning scripts ( ansible ) that run under root and pretty much that's it. I haven't seen human users in passwd, since virtualisation kicked in couple of years back.
- cup-of-tea 9y agoI do. In science we share clusters/supercomputers.
- lsc 9y ago>My experience of the last five years is so heavily weighted to (effectively) immutable infrastructure that checking to see who had been on a box hadn't event crossed my mind. In web scale production, you mostly have to worry about your fellow sysadmins troubleshooting the problem, and they mostly won't be too mad about you clearing state. But not everything is web scale production. Productionizing an app is a lot of work. Yes, yes, "Cloud is reliable!" and that's kinda true if you write your application to deal with any failure at any time. the reality is that the hardware can and will sometimes go away without warning, and you can't call up your oncall and tell them to head to the datacenter to fix it. all that recovery is done at the application layer; you had better hope you managed to make the data redundant enough. The whole idea behind the "cloud" is to take most of the lower level sysadmin work and make developers do it; and that's a fine way to some things, but... not so fine when it comes to other things. (This is the value proposition of a VPS system rather than a "cloud" instance. when the hardware a VPS is on goes down, some poor bastard's pager goes off, and they are expected to wake up, drag themselves to the datacenter and fix it. The gamble is "is being down for a few hours now and again, when you can reasonably expect to be brought back up in the same configuration cheaper than writing your software such that you can always start over on a new node?" The VPS is for the former, the "Cloud" for the latter.) Add to that, well, a lot of businesses need to run proprietary software. I personally, well, let us just say that for the last three years, the systems I am dealing with at work involve FlexLM. so corp and smaller sites are littered with one-off systems... systems where if they break, a pager goes off, and someone fixes the problem. Sometimes that works better than the "web scale" stuff... like I've yet to see a "web scale" posix-ish filesystem that meets the user expectations the way NFS does, and I've never seen a nfs server that didn't require someone on pager.
- chillidoor 9y agoI do this daily, mainly because most of our infrastructure isn't immutable yet :(
- shirro 9y agoWhen you have a small number of systems, a small team (possibly one person) and a relatively simple configuration it can be hard to justify spending time and money on a lot of deployment and automation. It isn't that deploying systems as you suggest isn't a better way of doing things. It is just the benefits are far greater at scale. At the scale I typically work it would be over-engineering most of the time.