14 ms·
woah, wait..... It "geneerally works if" you " rebuild docker hosts on a daily or more frequent basis." Perhaps I'm misunderstanding, but needing to rebuild m
by bphogan 10y ago
woah, wait.....
It "geneerally works if" you " rebuild docker hosts on a daily or more frequent basis."
Perhaps I'm misunderstanding, but needing to rebuild my prod env several times a day seems pretty "not ready for prime time" to me.
That's like when we'd say that Rails ran great in production in 2005, as long as you had a cron task to bounce fastCGI processes every hour or so.
So, can you elaborate on why rebuilding the containers is good advice?
- jjn2009 10y agoThe host isn't the container itself. They want to re-provision the host likely not because of something wrong with the application but instead docker is in some state which is non-recoverable, or at least not recoverable by automatic means.
- jacques_chester 10y agoThe security exec at Pivotal, where I work, has been talking about "repaving" servers as a security tactic (along with rotating keys and repairing vulnerabilities).[0] The theory runs that attackers need time to accrue and compound their incomplete positions into a successful compromise. But if you keep patching continuously, attackers have fewer vulnerabilities to work with. If you keep rotating keys frequently, the keys they do capture become useless in short order. And if you rebuild the servers frequently, any system they've taken control of simply vanishes and they have to start from scratch. I'm not completely sold on the difference between repair and repave, myself. And I expect that sophisticated attackers will begin to rely more on identifying local holes and quickly encoding those in automated tools so that they can re-establish their positions after a repaving happens. But it raises the cost for casual attackers, which is still worthy. [0] https://medium.com/built-to-adapt/the-three-r-s-of-enterprise-security-rotate-repave-and-repair-f64f6d6ba29d#.k3dj6wlp1 https://medium.com/built-to-adapt/the-three-r-s-of-enterpris...
- tptacek 10y agoHaving everything patched as soon as patches are available (or within, say, 6 hours of availability, for "routine" patches, with better responsiveness for critical patches) is a win. The rest: not so much. Rebuilding continuously for security is not something I would recommend.
- jacques_chester 10y ago> Rebuilding continuously for security is not something I would recommend. So that I understand, could you elaborate? Particularly, do you mean "not recommend" as in "recommend against" or "not worth the bother"?
- tptacek 10y agoIt's not worth the bother. Apart from keeping patches up today --- which is a good idea --- it's probably not really buying you anything. It's not crazy to periodically rotate keys, but attackers don't acquire keys by, you know, stumbling over them on the street or picking them up when you've accidentally left them on the bar. They get them because you have a vulnerability --- usually in your own code or configuration. Rebuilding will regenerate those kinds of vulnerabilities. Attackers will reinfect in seconds.
- jacques_chester 10y agoThat was my hunch too. Thanks. I'll ask more about whether I missed something on the other side of the argument.
- tkiley 10y agoIt seems like it's good to be able to rebuild everything at a moment's notice after patching against a major exploit, though. You should have a fast way to rebuild secrets and servers after the next heartbleed-scale vulnerability.
- tptacek 10y agoBeing able to rebuild critical infrastructure from source, and know that you'll be able to reliably deploy it, is a _huge_ win for security. After a bunch of harrowing experiences with clients, I'm pretty close to believing "using packages for critical infrastructure is a bad idea".
- 10y ago
- nolok 10y ago> So, can you elaborate on why rebuilding the containers is good advice? While I sincerely hope I'm wrong, I assume it's because you reset the clock on the probability something goes very wrong.
- sp527 10y agoThe "have you tried turning it off and on again" of DevOps. It makes a surprising amount of sense though, as long as your service is truly stateless, the restart can be easily orchestrated, and it results in no difference in operational costs.
- creshal 10y agoIf it's stateless, then why does rebuilding it change anything about the frequency of bugs popping up?
- kbar13 10y agoif both the code and the infra it's running on is stateless, then yeah.
- XorNot 10y agoOoh I can answer this one: because ask people if their container root is writeable, and get amused at the blank stares you get back. I am currently fighting an ongoing battle at work to point out that the plans for our Mesos cluster have not factored in that the first outage we have will be when someone fills up the 100gb OS SSD because no one's given any thought to where the ephemeral container data goes.
- siliconc0w 10y agoWe see a couple of different bugs that are best solved by simply rebuilding the container host. To docker's credit these tend to decrease with high versions. We also see them mostly in non-prod environments where we have greater container/image churn. We use AWS autoscale and Fleet so containers just get moved to other hosts when we terminate them. We have actually thought about scheduling a logan's run type job that kills older hosts automatically - it's in the backlog.
- majewsky 10y agoBugs that can be solved by a rebuild are not restricted to Docker. We had an interesting week when the build was red all the time for various reasons, and then prod started failing. Usually we deploy once a day, and not deploying for a week caused several small memory leaks to turn into big ones.
- benologist 10y agoBecause then you always know you can always rebuild automatically and that's being tested constantly while developers work, a bit like how Netflix crashes everything all the time randomly to ensure they can always automatically recover from every dependency. It also naturally rewards optimizing around time-to-redeploy, probably a lot of benefits there.
- StreamBright 10y agoDepends, usually you have to be able to re-build your prod infra within minutes or maximum hours, otherwise you are doing devops wrong. The whole point of automation is reproducible infrastructure that you can stand up quickly. With stateless approach you can just do this. Why would you do that? Imagine an outage in one of the 3 datacenters you are running your infra in the same region. You need to move 1/3 of the capacity to the remaining 2 datacenters. This is not too much different to re-building it.
- raverbashing 10y ago"have to be able" is very different from "have to" Sure, you want to be able to deploy quickly. But if there's no reason to, then don't. And I would be very scared if Docker images had a 1 day uptime max
- cortesoft 10y agoI know you probably aren't trying to address all cases, but just because you can't re-build your prod infrastructure in minutes or hours doesn't mean you aren't doing devops right. Many larger companies can't do this; my company has 70+ datacenters with tens of thousands of servers. We can't re-build our prod infra in minutes or hours. We are still doing devops right :D Like I said, I know you aren't talking about my situation when you made your statement... I just get frustrated when people act like there are hard and fast rules for everyone.
- StreamBright 10y agoWell I am not talking about a 1M node outage. That you cannot fix with anything. I am talking about a maximum datacenter wide outage, that actually happens pretty often. Amazon has game days, Netflix has chaos monkeys for the same reason. Make sure that you can rebuild parts of your infra pretty quick.
- cortesoft 10y agoOh, of course. Our datacenters often go offline (both planned and unplanned), and we are always ready to handle that. We are pretty much constantly re-provisioning servers... with so many physical machines, hard drive and other hardware failures are a daily occurrence.