4 ms·
I'm not going to pretend I know what Mach is, but around here, (big company that you're familiar with), rebooting/bouncing the servers is pretty much how issues
by mobilemonkey 16y ago
I'm not going to pretend I know what Mach is, but around here, (big company that you're familiar with), rebooting/bouncing the servers is pretty much how issues are dealt with. "Response times outside of SLA: bounced the server." "Database connections timing out: bounced the server." "Users experiencing high load times for pages: restarted JVMs. Then bounced the servers."
Root cause seems to be "server up too long."
- derleth 16y agoMach is a microkernel. That means it's a way to design an OS such that most of the things application programs rely on are outside the kernel, provided by servers that can be restarted and replaced without rebooting the whole system. For example, the majority of the filesystem is provided by a filesystem server, most of the networking code is in a networking server, and so on. You begin to see why a microkernel that needs to be bounced more often than, say, Linux, which is not a microkernel, begins to lose some of the appeal of having a microkernel in the first place. (Aside: Mac OS X is based on Mach, but practically all of the stuff applications see is provided by a single server which (to the best of my knowledge) can't be restarted without taking down the whole system. It's a hybrid design.) Also, just to confirm a stereotype, do you use Windows for your servers?
- stcredzero 16y agoAlso, just to confirm a stereotype, do you use Windows for your servers? Related tangent: There was a time when the the US military's COTS (Civilian Off The Shelf) initiative resulted in AEGIS missile cruisers running largely on machines running Windows NT. Yes, the ships responsible for screening carrier task forces against attack -- had to be rebooted on a weekly schedule. To be fair, the unmolested NT kernel is rock solid. I know people used to run RAID arrays with it, but for the desktop, Microsoft allowed the crossing of certain architectural lines.
- rst 16y agoAnd if the NT machines running your missile cruiser don't come back up, then you have real problems. During trials, the USS Yorktown had to be repeatedly towed into port after systems failures... http://lists.essential.org/1998/am-info/msg03829.html http://lists.essential.org/1998/am-info/msg03829.html
- VladRussian 16y ago>the unmolested NT kernel is rock solid. which, if memory serves me right, directly translates to "OS/2 kernel is rock solid"
- stcredzero 16y agoAlso closely related to "VAX kernel is rock solid"
- wmf 16y agoNo, I don't think NT is a derivative of OS/2. "Given these issues [with OS/2], Microsoft started to work in parallel on a version of Windows which was more future-oriented and more portable. The hiring of Dave Cutler, former VMS architect, in 1988 created an immediate competition with the OS/2 team, as Cutler did not think much of the OS/2 technology and wanted to build on his work at Digital rather than creating a "DOS plus". His "NT OS/2," was a completely new architecture." http://en.wikipedia.org/wiki/OS/2 http://en.wikipedia.org/wiki/OS/2
- deleted 16y ago[deleted]
- derleth 16y agoWindows NT is more closely related to VMS than OS/2, but most directly related to DEC's unreleased Mica OS for their unreleased PRISM CPU architecture. This is unsurprising: They're all designed by the same guy, Dave Cutler.
- rbanffy 16y ago> the unmolested NT kernel is rock solid Too bad it resents running programs on it so much...
- stcredzero 16y agoIt's not that. It's more boneheaded stuff like injecting components supporting the GUI into the kernel. When that happens, it doesn't resent those programs so much as that it's been violated and turned into a warped version of itself. Sort of like Jeff Goldblum's character in The Fly.
- patrickgzill 16y agoFrom what I understand the NT kernel was rock solid, if a little picky about hardware. The problem was that graphics performance suffered due to the strict isolation of the drivers from the kernel. MS relaxed those restrictions in later versions of their OS (either NT4 or Windows 2000), gaining better graphics performance at the cost of kernel stability.
- reeses 16y agoMach isn't a microkernel. For example, NeXTSTEP was based on Mach 2.5, which is a monolithic kernel. Mach 3.0 was the only microkernel version, and to my knowledge, no one has adopted it with any success.
- msbarnett 16y agoMach 3.0 was folded into XNU in roughly the same way that earlier versions of Mach were used, but with more of the BSD functionality ported down into Mach rather than hosted POE-style in kernel space. So 3.0 was adopted, with success, but not as a microkernal hosting user-space servers.
- glhaynes 16y agoDoes "POE" stand for something? POSIX Operating Environment, maybe? A quick Googling and a look at the Mach Wikipedia page doesn't reveal the answer.
- mobilemonkey 16y agoHmm, nope. Solaris of some flavor I think.
- rbanffy 16y agoOh boy... That's what you get when you assign the management of Solaris servers to MCSAs...
- josephcooney 16y agoSo if windows systems are unreliable it's windows fault, and if other OSs are unreliable it's also windows fault?
- rbanffy 16y agoNo. It's just that Solaris (and every other kind of Unix) admins know very well that they don't bounce servers when they don't know what's wrong - they find out what's causing the problem and correct it. There is a very rich set of tools to make that easy. Bouncing a server may even make the problem go away, but it does little to enlighten you on why it happened in the first place. And yes, If your MCSAs tend to bounce servers all the time, that's because their experiences have shown this works for Windows.
- josephcooney 16y agoThere are plenty of tools for troubleshooting problems on windows. And your statement that "every kind of unix admin knows not to do that" seems plainly at odds with the reality here.
- rbanffy 16y ago> seems plainly at odds with the reality here. You have a human resources problem, not a technical one.
- mobilemonkey 16y agoThanks for the great explanation, too, btw.
- deleted 16y ago[deleted]
- jwhitlark 16y agoAt my last sysadmin job, we had a policy of bouncing a server if it's uptime exceeded one year, not to clean up any resources, but to make sure the config in the files matched the config that was running. Of course, not only were we running Linux (gentoo, in fact), but a lot of our stuff was DJB services, including his service manager. I'd argue that a 5 why's analysis of "server up too long", leads to "server wasn't written well." YMMV
- mpk 16y agoNow this is an excellent reason and timeframe for doing periodic reboots. It's basically a scheduled sanity check for your infrastructure. It boils down to robustness and assurance. If parts of your system won't survive a reboot this is an unknown and therefore a risk. If a system can't survive a reboot it's better to find that out when it's during a scheduled reboot and personnel is on hand to deal with it rather than relying on on-call personnel responding to alarms from monitoring at some point. Apart from the repair time being longer when relying on on-call technicians, there is also the general observation that productive systems tend to become more valuable or have more exposure over time as their usage increases. A failure at a later point therefore also becomes more costly as time goes by.
- rbanffy 16y ago> Now this is an excellent reason and timeframe for doing periodic reboots. Most of the time you don't need to do complete reboots. Restarting services usually takes care of the majority of the sanity checks.
- mpk 16y ago> Restarting services usually takes care of the majority of the sanity checks. Sure, that usually takes care of most checks, but the point is to eliminate as much as possible unusual scenarios. A service-level restart will catch service config mismatches, but it won't catch, for example, manual interface or routing table updates. It also won't catch the 'fragile root' problem where the device that the kernel is loaded from has become corrupt. Bad things happen in the strangest places if you run enough systems.
- Twirrim 16y agoSee that's not actually solving the problem. Thats more like pretending that the problem doesn't actually exist. Instead of investigate and fix permanently so that customers don't experience problems they'd rather just keep letting the problems fester.
- doyoulikeworms 16y agoNot everyone is "pretending that the problem doesn't actually exist." There are limited resources available to any and every organization. Many times, a proper fix can cost more than a band aid, sometimes even over time. Here's a hypothetical, but not ridiculous, scenario. Which is a better way to spend resources? Track down and stamp out an extremely difficult resource leakage bug? Or simply bounce the server? The costs associated with the former may be HUGE. The costs associated with the latter may be small in comparison (cost of potential down time, minimal labor cost of bouncing the server(s), and costs associated with user reliability). I wouldn't always chalk it up to whatever it is that your comment implies (laziness? willful ignorance?).
- mobilemonkey 16y agoYeah I'm absolutely implying something by my comment. I understand the time:value relationship and can expect some degree of RCA to go away if rebooting a server solves a problem for X amount of time, where X is less than a few days. But if you're spending the time opening an alarm, calling people to a bridge, agreeing to bounce a server, and then bouncing it, at some point that equation comes out in favor of actually doing some sysadmin/RCA and making it so you're not bouncing a server every 48/72 hours. It also looks really JV to see those alerts day in/day out. Like, come on. We can't just fix the problem?
- NetMonkey 16y agoI really doubt anybody would accept that. The solution in many cases are simply to reboot the servers every night. Sure, we would all love to fix the bugs, but if there are no easy clues and a regular reboot fixes it - who really cares when there are features to build?
- deleted 16y ago[deleted]