8 ms·
If you need a hot fix "RIGHT NOW" you might be doing something wrong in the first place. Being able to just ssh into a machine is one of the problems that we d
by tex0 6y ago
If you need a hot fix "RIGHT NOW" you might be doing something wrong in the first place.
Being able to just ssh into a machine is one of the problems that we did solve with containers. We didn't want to allow anyone to SSH into a machine and change state. Everything must come from a well define state checked into version control. That's where containers did help us a lot.
Not sure what you mean with lingering containers. Do you launch yours through SSH manually? That's terrible. We had automation that would launch containers for us. Also we had monitoring that did notify us of any discrepancies between intended and actuated state.
Maybe containers aren't the right tool for your usecase - but I wouldn't want to work with your setup.
Btw. most of this is possible with VMs, too. So if you prefer normal GCE / EC2 VMs over containers that's fine, too.
But then please build the images from checked in configuration using e.g. Packer and don't SSH into them at all.
- murkt 6y ago> If you need a hot fix "RIGHT NOW" you might be doing something wrong in the first place. Maybe I phrased this point badly. I haven't found myself to need a hot fix RIGHT NOW for a long time. What I had in mind was this: sometimes there is an issue that is very hard to reproduce locally. If a developer doesn't understand it from the get go and have to experiment a bit, it can be frustrating to commit, wait for deploy (even to stage environment), test what happens and repeat. > Not sure what you mean with lingering containers. Do you launch yours through SSH manually? Of course the application itself is run automatically, web server, workers, etc. However, we often need to do some bespoke data analysis, so we often ssh into a server, type `make shell` to launch a REPL and type/paste some stuff into it.
- compsciphd 6y agodocker exec -it <container id> kubectl exec -it pod -c container
- murkt 6y agoNot sure what do you want to say with this command.
- compsciphd 6y agoyou want to be able to "ssh in and make changes to figure out how to fix things". why doesn't that do that for you?
- reportgunner 6y agoSorry, but did you not forget the --rm option to remove the container once it dies ?
- AkshatM 6y agoI interpret it as a suggestion that you allow developers access to the running container in production to debug the issue There are more sane ways to do this than through raw kubectl access, of course: see e.g. telepresence (https://www.telepresence.io/ https://www.telepresence.io/)
- gnfargbl 6y agoI think you might be falling into a cycle that I often find myself falling in to. It goes like this: Oh, there's a bug. But it's obviously caused by _this line here_, so I won't even bother testing it offline, I'll just fix it and run CI and then... oh no, wait, it was more complex than I thought, but I'm in the change-CI-test cycle now so I'll keep doing that. And then, all of a sudden it's taken me three hours or more to fix what was in reality quite a small bug. The solution is to force yourself not to use the CI/container system at all during debugging, and instead to build the binary (or whatever) standalone. That's hard because you invariably aren't tooled up to run the component outside its deployment system so you have to do some extra tricks, but in the long tun it seems to be the way to go.
- murkt 6y agoOver the years I've learned to do this very rarely. However, I'm not the lone wolf in the woods, I also have other developers working for me. I can talk to them however long I want and motivate them to debug locally to understand the problem, but in the end they have to gain some experience with their own pain. I would prefer that their experience gaining would be faster, so it would cost me less money :)
- paloaltokid 6y ago> I would prefer that their experience gaining would be faster, so it would cost me less money :) Maybe your current strategy isn't working?
- one2know 6y agoThis is very common and just because it doesn't match your use case doesn't mean some businesses don't need a hot fix "right now." If you work in 24/7 ecommerce and your site is producing $60k per hour and there is a network failure that breaks something you need a hot fix right now, otherwise your 3 hour code review, build, q/a, deploy pipeline will cost the company $180,000.
- imdsm 6y ago> otherwise your 3 hour code review, build, q/a, deploy pipeline will cost the company $180,000 Taking code review away, because we all know we're not going to sit around waiting for a review while a patch needs to go out urgently, if your build, test, and deploy pipelines take three hours then you have some serious problems that you need to address, and containers aren't it. There are methods for handling hotfixes/patches to production quickly that work well in high volume sales website setups.
- collyw 6y ago> If you need a hot fix "RIGHT NOW" you might be doing something wrong in the first place. Without ant context that seems like an outrageously ignorant comment. (In saying that most companies probably are doing many things wrong).
- fizixer 6y agoAgreed. This confused the heck out of me. Almost sounds like script kiddies when they start learning programming for the first time in their lives they assume that the test of successfully learning how-to-program is that your code should compile on first try without errors.
- overgard 6y ago> In saying that most companies probably are doing many things wrong I think that sounds actually very plausible :)
- secondcoming 6y agoI don't know what I'd do without being able to ssh into VM instances. Whether it's for looking at various logs, the occasional core dump, or uploading a custom binary to test something, it's incredibly time saving.
- mistahenry 6y agoI’ve worked in a number of areas with basically no fail/delay SLAs. I think it’s naive to think “if you need a hot fix right now, you’re doing it wrong”...the number of times we needed to hot fix because of ourselves was very low. But when you’re in an integration heavy environment and one of the many moving parts (outside of your control) breaks, well thought out “put the fire out” stopgaps on the server consistently save the day (and the company money by not breaching the SLA)
- scarface74 6y agoRedeploy the older working version?
- mgkimsal 6y ago"outside your control" is key here. you're assuming a rollback would work. in many cases, some external system changes without your knowledge, and you're only seeing those changes on production. I've got a client that has data feeds from multiple vendors. some are pulls, some are... "hey, we'll FTP this file to you". the file format has changed - unannounced - at least 3 times in the past... 15 months. Then something breaks on production, but you don't know what. You need to get on that machine and take a look. "Redeploy the older working version" doesn't do anything except re-introduce more problems in these instances.
- scarface74 6y agoIsn’t the larger issue that your production environment can be brought down by bad user input?
- detaro 6y agoThere's breakage outside "brought down". A system that's running but doesn't produce outputs because the input data changed can be "broken" and violating its SLA too. And not really something you can design around outside "we won't promise anything", but then you loose to competitors that do take it on themselves to react quickly enough with hotfixes.
- soVeryTired 6y agoBut... you can ssh into a container and change state (depending on config options), can't you? I'm not sure I follow this response or the OP's complaint. I mostly write python code and one very nice pattern I've found is to run a container somewhere (either locally or on a server somewhere) with an open port. You can SSH into it and use the remote interpreter as your main project interpreter. That way your dev environment is 100% reproducible. VS Code an Pycharm Pro both support doing that.
- Galanwe 6y agoI think OP is working with a farm of containers that are spawned and destroyed dynamically. Which means SSHing on one and fixing a problem would not really help.
- eejjjj82 6y agoDepends on how you setup the container.
- mark_l_watson 6y agoFor small deployments, running a Common Lisp web app with remote repl access is great for finding problems. I wouldn't recommend this for high traffic apps but there are many use cases in this world for small user base, focused web apps where maximizing developer productivity is required for profitability.
- bartread 6y ago> If you need a hot fix "RIGHT NOW" you might be doing something wrong in the first place. Can we please grow up beyond this kind of comment? I suspect everybody knows that hotfixing production isn't an ideal thing to do, and the many reasons why that's the case, but lots of us nevertheless do it from time to time. Where I work now we've probably hotfixed production a handful of times over the past couple of years, amongst thousands of "proper" deployments. It's really not a big deal. We know so little of the context behind OP's issues that picking holes isn't helpful or informative. The specific question was about migrating from containerized infrastructure back to a more traditional hosting model. I for one would be interested in reading about any experiences people have in this area so it's quite frustrating to find this as the top comment when it's substantially off-topic.
- throwaway0a5e 6y ago>Can we please grow up beyond this kind of comment? No platform that assigns quantitative virtue points to comments based on how many people agree with that comment will ever not have an abundance of people lobbing low-effort quips and dog whistles designed to appeal to the majority.
- altcognito 6y agothis happens in the workplace as well though
- bartread 6y ago> this happens in the workplace as well though Exactly this. I'm fortunate in that where I work now it's pretty rare: the culture is fairly collegiate and friendly, and there's a strong sense of "we're all on the same side". Other places I've worked every meeting has been an exercise in point scoring, which is incredibly wearing and - try as one might - the culture does end up influencing one's own behaviour. You can obviously get preachy about how people should be stronger characters and not so easily influenced but the phrase "Bad company corrupts good character" exists for a reason. Unless you're Gandhi it's incredibly difficult for many people to resist the culture around them day in and day out without some positive reinforcement from the behaviour of others, especially when you're incentivised to do otherwise, and may be penalised for not doing so[0]. [0] The answer is, if you can, find another job. Not always an option, but a good idea if it is.
- rco8786 6y agoGive me a break. Everyone needs a hotfix occasionally. If your users are asking for something and your response is “you’re just doing it wrong!”, you’re probably the one who is wrong.
- jonfw 6y agoBut if your developers are asking for something and your response is "you're doing it wrong" (hopefully not verbatim), you've got a good mentorship opportunity on your hands
- zemo 6y ago> Being able to just ssh into a machine is one of the problems that we did solve with containers. you should be able to ssh into a machine and strace a process to see why something is going wrong. If your only solution is always "restart the container" or "revert to an old container" or "only deploy containers that are known to work" you're not actually debugging extant problems.