4 ms·
Likely they just had no need for it themselves. Might be that they didn't prioritize having a feature over the risk of someone taking over their hypervisors tha
by chousuke 6y ago
Likely they just had no need for it themselves. Might be that they didn't prioritize having a feature over the risk of someone taking over their hypervisors thanks to a buggy serial port emulator.
Pretty much all hypervisors support serial consoles, but usually those interfaces are limited to trusted admins. For something like AWS, they'll also have to connect it from the hypervisor hosts into their public UI, and they can't trust the users.
- bigiain 6y agoThey probably don't need it because they are much more likely to follow best practise "treat servers as cattle, not pets". If an instance wedges itself onto a state where I need console access, I'd just kill it and provision a replacement (ideally, my monitoring and automation will have done that already and not even have woken me up to tell me). I'm not sure I'd be at all comfortable having irreplaceable single points of failure in AWS. (Though I do recognise that people use it that way all the time...)
- sargun 6y agoHow do you debug broken instances?
- x3n0ph3n3 6y agoShut them down and mount the volume on a new host.
- somethingAlex 6y agoTypically the goal is to architect the system such that you don't really care. If it's stateful, there's some other replica. Promote that and then spin up a new replica from a backup and roll it forward. If it's stateless then just kill it and spin up another.
- sargun 6y agoBut what if there is a bug that’s repeatable / pops up on the regular?
- bigiain 6y agoSame as a herd of cattle, if a bunch of them get sick in similar ways you change your process to one where you can find out why. But just one? Shoot it and bury it. Then keep an eye out on the rest in case it's a developing pattern.
- kubanczyk 6y agoI'd use it to troubleshoot quirky AMIs which simply do not boot in a specific setting no matter how many times I try. Anything more esoteric than "a normal Ubuntu" can have a bug, e.g. it hangs with three network interfaces or similar.
- chousuke 6y ago"cattle vs. pets" misses some important nuance though, especially the way it's usually used to emphasize how you need to architect systems in a cloud environment that way because you have no way of fixing some issues. What it is actually saying is not that it's a good approach to just throw away servers when they start having problems. What's important is having that ability when it is necessary. Servers don't just randomly fail. If your instance goes down, sure, your automation will recover your system, but it will still be important to know why it failed, because it may be a symptom of a deeper issue. It can be due to a hardware issue, but even when that happens, you don't throw away good hardware; you fix it and the server can return to full operation. In the cloud, though, you have no idea what the hardware is doing. Maybe the instance failed because the underlying host failed; maybe it didn't. You should still find out.
- bigiain 6y agoSure, I totally agree about nuance and the "needing to architect it that way" meaning there. But once you have it architected that way, you then gain the ability to mostly ignore single failures, and only look for "deeper issues" if failures persist. Even at not-very-high scale, AWS instances _do_ "just randomly fail", at least for all practical interpretations. I don't run anything like FAANG scale, only hundreds of instances rather than thousands or millions, and I see at least a few "random failures" a year (not including spot instances terminating, which I see in clumps every month or so). I (almost) never try to repair broken a EC2 instance. Wherever I can, they'll be running totally stateless, and I just provision new ones and kill off old ones. I probably won't even bother investigating if it's a rare and singular problem on a known-reliable platform. If one instance wedges and gets replaced, I'll just have a note to investigate if it happens again any time soon. If we get a second failure, we'll go looking in logs and maybe keep and investigate the EBS volume. For platforms running new-ish code, procedures are different. If we see dead instances after deployments we obviously investigate the new code/config there. But a fair chunk of clients where I am only get 6 or 12 (or even 24) month backend update cycles, if I've got dozens of instances running the same code for months on end and _one_ dies, we just bury it and replace it, and keep a closer eye on the rest of the "herd" for a week or two.