3 ms·
No, what this means for DevOps is that you'd better create a foolproof method to insert your dotfiles into the distribution - and that it's more important than
by jaylevitt 14y ago
No, what this means for DevOps is that you'd better create a foolproof method to insert your dotfiles into the distribution - and that it's more important than ever to argue for individual user accounts, not role accounts, on production machines. For security and logging, of course.
Love,
DevOps
- moe 14y agomore important than ever to argue for individual user accounts, not role accounts, on production machines. For security and logging, of course. That's a terrible idea on so many levels, I hope you're being sarcastic.
- jaylevitt 14y agoYou're saying it's better for sysadmins to all log in as "prod" or "root"? How do you figure?
- moe 14y agoFirstly on properly automated hosts you very rarely have to login. And when you do then 'prod' and 'root' work just fine. Scattering user-accounts across machines is a maintenance headache (trust graph, sudo passwords or passwordless sudo, filesystem permissions) and encourages bad practices ("i'll just run this small script, quickly"). If you're serious about auditing then wtmp and notoriously incomplete shell-history files are not your tools either.
- ams6110 14y agoLogging in a root is an exceedingly BAD idea, regardless of how properly configured the host is.
- donavanm 14y agoLet me know how your compliance audits go. I'm sure the pci dudes won't mind that "root" logged in from 20 different hosts.
- moe 14y agoWhy would a human login at all? In serious deployments people don't login to production hosts outside of extra-ordinary incidents. This entire discussion is based on a broken premise.
- jaylevitt 14y agoIn single-server deployments, you're logging in because you're the only employee. In ten-server deployments, you're logging in because every machine is trying to take on a new role or aspect and you're still learning how to automate that. In thousand-server deployments you're logging in because the latest batch of Seagate drives has a statistically significant failure rate and you need to try some experimental firmware they wrote for you. When do you get "serious"? I don't always log into my production boxes... but when I do, I log in as 'jay'.
- moe 14y agoIn ten-server deployments [...] still learning to automate That's why your original comment tipped me off. It's hard to recover from bad patterns like the one you proposed, it's easier when you try to get as many things right as possible from the start. In thousand-server deployments you're logging in because I take it you have not worked in such an environment before. I can assure you nobody manually logs into hundreds of hosts to "try some experimental firmware" - or to do anything really. Beyond a couple dozen hosts pdsh and "knife exec" simply don't work anymore (rollback? what rollback?).
- jaylevitt 14y agoI take it you have not worked in such an environment before Does AOL count? I built the mail system. It was kind of a big deal. We pushed 4,000 TPS through servers less powerful than an iPhone. I also spent time at Akamai, though I admittedly worked on a tiny, isolated test lab (150 servers or so, a fraction of a percent of the "real" server base). I have worked with some amazing people, at all scales, and I have never seen a deployment that managed to completely avoid manual logins. No, you aren't logging into a few hundred hosts manually (well, there was this one shop... but let's not hold that up as exemplary). But you're probably logging into a few to test things out manually before you decide what your automated deployment script will be rolling out. Even when you do deploy a fix, with a large sysadmin team it's good to know WHO exactly ran that one-off deployment script. When I said "auditing", I wasn't even thinking of PCI, or auditing against malicious actions of any kind - just simple troubleshooting forensics. But we have a great, robust way to see exactly who modified a file and when; it's the file system. Why reinvent it? (I just spent the last day trying to figure out who edited a script. Why, "postgres" did.. of course it did.) Should you aim to do these manual config changes in the test lab first? Of course. Do you always succeed? No. Some things only get tested in the big lab. Reality intervenes. I try to plan for it. Mind you, I'm arguing this partly to see if I can be talked out of it. I think I believe it, but I've been wrong about way bigger things. If you've achieved my automation nirvana, and you roll things out to a few thousand servers, never manually, and never wondering who did it, tell me how it works.
- pyre 14y ago> properly automated hosts > Scattering user-accounts across machines is a > maintenance headache Seems to me that properly automated hosts could easily have user accounts automated too, no?
- bdunbar 14y agoAnd when you do then 'prod' and 'root' work just fine. They do until you have to show auditors who logged in and when. Scattering user-accounts across machines is a maintenance headache That's why God invented Active Directory.
- moe 14y agoThey do until you have to show auditors who logged in and when. SSH has logging. Also if you have audit-requirements then the login-log is normally the least of your worries. That's why God invented Active Directory. That must be one cruel god you have there...
- bdunbar 14y agoI agree that SSH logging is sufficient. That was not the finding of our auditor, per best practices, blah blah. Thank SOX for that one. That must be one cruel god you have there... I don't have to _deal_ with it directly, just use it to authenticate logins on (most) of my unix servers. Seen from afar, it's not that bad. Now, she is a jealous god, as perceived by her acolytes. Nobody really wants to hear about how, during last year's DR testing at IBM, my stuff was up and running waaay before AD was working.
- pyre 14y agoIf you're worried about logging who did what under the user, you can disable password access, and use ssh keys. IIRC, ssh can log which key was used to log in. I'm not sure how audit-able this would be (for example) in determining which user logged into 'prod' ran the malicious process.