3 ms·
Or you can signify the context using shell prompt and use kube rbac to give read only permissions instead of letting people mess with your prod cluster.
by longcommonname 6y ago
Or you can signify the context using shell prompt and use kube rbac to give read only permissions instead of letting people mess with your prod cluster.
- mlthoughts2018 6y agoManaging shell tooling to edit ps1 for this is a bad idea, because there are already many other much longer established mutations people use that for (like git branch or activated Python environment), kubernetes can’t come along decades later and suggest to use the tool effectively one has to break those long standing workflows. It’s always an available option for people who want it, but many many use cases are poorly served by it. The RBAC thing is different. I am 100% for those RBAC limitation as long as I have total authority to make a devops admin do what I need them to do right when I need them to do it. If my team is going to get alerted on build failures or service errors that require kubernetes mutation actions or exec to debug or resolve (99% of the time they do), then my team needs full admin access. If rbac prevents access, then you better route the alerts to someone else or give us authority to compel instant triage responses. In the 3 different companies I’ve worked in that use large kubernetes clusters, rbac has been a miserable problem across the board, because SREs don’t want to be responsible for triaging or resolving application team issues, and application teams will not honor alerts that are not actionable because of poorly conceived permissions issues where people mistakenly think they are setting useful access control policy but really they aren’t.
- DasIch 6y agoYou could have a mechanism to grant yourself temporary write access on an as-needed basis. At Zalando everyone has read access (except to secrets). You can request write access with a command line tool another employee can then approve it. There is also an option to request access with an incident ticket, in that case you immediately get write access without approval by someone else. Write access expires after 1h. Access to the underlying AWS account uses the same mechanism.
- mlthoughts2018 6y agoBut then when you’ve unlocked access, you could just make the same fat-fingered mistakes due to implicit use-context settings. You might use use-context to set yourself to that restricted production context at the start of an incident. 30 minutes later you resolved the incident, but the access is still active and whoops you delete a deployment you thought you were deleting from a stage experiment or something. All that does is add an extra hoop to jump through. People need to stop believing that adding extra bureaucratic hoops offers any type of safety - it doesn’t. If you want to restrict access then you need to actually restrict it and convert on-call alert responsibilities to a central devops team that, like it or not, is responsible for solving everyone else’s problems. If you don’t want a central devops team that can be paged and on the hook for everyone else’s systems, then you cannot have write access controls. There’s just no way out of this dilemma. Temporary access that can be self-granted == no access control. Temporary access that requires admin approval == admin team is on pager duty for every other team.