12 ms·
Stop using Nagios (so it can die peacefully)
- gk1 13y agoYou should try Scalyr (https://www.scalyr.com/ https://www.scalyr.com/). It's easier than juggling between six different tools. It was built by ex-Google DevOps engineers for the same reason you made this Slideshare: the available tools suck. (Full disclosure: I'm working with Scalyr, but you should still try it.)
- fotcorn 13y agoI don't think something like this should be hosted/SAAS. I don't want to send my Gigabytes of logfiles (sometimes containing confidential informations ...) over the internet to some unknown entity with probably questionable security. My small startup company already has more than 10 (virtual) servers which would cost us 500 dollars a month to monitor which is more than the servers itself cost.
- gk1 13y agoCompletely valid points. Here's how we're dealing with each: 1. With our custom parser you can replace or delete confidential information from your logs before they're stored on our servers. 2. We're exploring a different pricing structure right now that would address this scenario. If this is your only hesitation, I hope you check us out anyway and contact us about pricing.
- j05h 13y agoCustom parser won't guarantee that all of the data is dumped...and some of it you may even want in the system. Also -- lots of companies are seriously concerned about pushing their data externally. Making it a hard sell. However, as a long tail service, this looks great.
- blueskin_ 13y agoDoes the parser run on the client? If not, it defeats its own purpose.
- snewman 13y agoYes, on the client. See "Redaction" at scalyr.com/agent.
- nknighthb 13y ago> With our custom parser you can replace or delete confidential information from your logs before they're stored on our servers. This is both error prone and utterly defeats the purpose. Why would I pay a bunch of money for somebody else to manage my logs when I'd just have to keep them all anyway so I can get at the unredacted versions when there's a problem?
- snewman 13y ago[Scalyr founder here] We take security very seriously, but let me turn this around into a question: what would it take for you to trust an external service to manage your logs? Some of the things we're doing: 1. SSL everywhere (including internal traffic between our backend servers). 2. We add a tag to the raw representation of every string value, so that we can verify that data never leaks across accounts. (This has never detected a problem -- except in tests, because yes, we do test it.) 3. Implementing in a "safe" language (Java), to rule out low-level buffer management bugs. 4. As Greg noted, we make it trivial for you to redact sensitive data before it leaves your server. We are sometimes asked for an on-premises installable version of our service. We don't provide that because we're using economies of scale on the backend to completely change the log management experience: when you give us a query, every CPU and spindle in our entire cluster is briefly devoted to that query. This means that you aren't limited to graphing predefined metrics; you can do ad-hoc exploration of your entire log corpus in on the fly. E.g. display a histogram of response latencies for all requests for url XXX on server group YYY in the last 48 hours, and expect a near-instantaneous response.
- mrweasel 13y ago>what would it take for you to trust an external service to manage your logs? I think that's a really weird question that completely fails to address the concerns that some people might have. We do logging of sales, profit margins and stuff like that. You can't have access to that because: "You're not us". If you can read our data, then we're not going to use your service and to do anything useful the logs you really do need read access. Of cause you might have no reason to spy on our data, but the only safety is that you promise not to. We could seperate logs for different things, so webserver logs go to you, but email logs goes to an internal system, but then we would need two systems.
- wdewind 13y agoDo you ever email about this data or put it into spreadsheets on Google's servers?
- 13y ago
- blueskin_ 13y agoAgreed. Contracting out your monitoring seems to me to defeat the point of running your own infrastructure.
- caw 13y agoThere's a distinct lack of images on that website. I for one would like to see how I would diagnose problems without opening 5 tabs. A "demo" mode would be even more helpful.
- snewman 13y agoThanks for the feedback! Demoing a log management product is a bit tricky, because it's hard to come up with a representative data set that isn't sensitive. FWIW, we have put together a demo based on data scraped from the Github API -- it's not a typical server log data set, but it can be fun to explore: https://www.scalyr.com/login?prefillEmail=demo-account%40scalyr.com&prefillPassword=demodemo&originalUrl=https%3A%2F%2Fwww.scalyr.com%2FlogStart https://www.scalyr.com/login?prefillEmail=demo-account%40sca...
- ceejayoz 13y agohttps://www.scalyr.com/dash?page=Github-Statistics https://www.scalyr.com/dash?page=Github-Statistics is broken. "Operation not permitted (Read Configuration permission required)"
- snewman 13y agoOop -- thanks! We'll sort this out, might take a little while though.
- angersock 13y agoYeah, no, not what I want to hear from the people hosting my logfiles. There's your answer to "Why won't you host your server logs (which are usually key for troubleshooting flaming boxen)?"
- snewman 13y agoAs I probably should have clarified, this issue is specific to a single page in the demo, which we do not even link to at the moment (aside from a couple of older posts on our blog). Yes, it is embarrassing, and I apologize. However, if this issue had been standing between a customer and their data, we would have scrambled instantly. Everything in life is a tradeoff. If you entrust your logs to us, you run the risk that we have an outage or failure of some sort. On the other hand, internal systems can fail as well. We hope to serve people who prefer not to carry the responsibility of maintaining their own monitoring infrastructure, and/or are interested in the features and performance we provide.
- samstave 13y agoThe price is absolutely insane and a non-starter.
- snewman 13y agoWe've been hearing that loud and clear. :) FWIW, the price actually works quite well for a lot of people. On a GB-for-GB basis, we're actually much cheaper than other hosted log management solutions -- we work hard on backend efficiency and we pass that along. But yes, if you're using small virtual servers then the pricing model breaks down. We originally went with this model to provide more predictability; log volume is often more volatile than server count. We've heard enough complaints that we've decided to move to more volume-based pricing model, we're just working out the details.
- samstave 13y agoThanks for the reply... as an OPs guy there are a ton of layered problems with running a highy elastic infra on something like AWS: 1. Dynamic registration of ephemeral systems with a monitoring platform. 2. Security monitoring of same 3. Meaningful graphing When we are optimizing the purchase of hundreds upon hundreds of spot instances daily, where we are looking at grabbin hosts for just a couple cents an hour, the model of per host fees for things like StackDriver, CloudPassage and your service makes per-host pricing completely a no-go. I don't have a good idea how these should be priced; but I think its important for people to understand all the other costs associated with having a solid management platform for your environment that covers all the bases and doesn't require another round of funding! :)
- kylek 13y agoMaybe I'm missing something but this just looks like log analysis (akin to Splunk) and not actually server monitoring? (active health checks, notifications, snmp, etc) Pricing seems wonky too...is it cheating the licensing model to aggregate logs on a single syslog server and submit from a single agent?
- gk1 13y agoRegarding the pricing, we're in the process of revising the pricing structure so that it's based more on volume instead of how many servers you have. In that case, you wouldn't need to cheat. :)
- quicksilver03 13y agoThis is interesting for us as well... We currently have "only" 50 servers (mostly virtual) that need to be monitored, a per-server pricing would push the cost far too high regardless of the quality of the solution.
- gk1 13y agoMay I email you about this? The pricing structure will be changed soon to accommodate cases like yours--there are many of them--so if that's the only thing keeping you from trying Scalyr then I'd love to chat.
- quicksilver03 13y agoWe've already started using a competitor's product (Logentries) and we're happy with it, so we're not looking to Scalyr or other log management solutions at the moment. Thanks for listening to feedback though!
- snewman 13y agoLog analysis is the heart of the product, but we also gather system metrics, provide notifications (scalyr.com/helpalerts), and we recently rolled out a basic active-checks feature (scalyr.com/helpMonitors). There's lots more to be done; for instance, we don't have any SNMP support today. But the vision is to be a full-spectrum tool, and we're actively working toward that. As for the licensing model: we're going to move to per-GB pricing anyway, so no worries there. If you'd like something more concrete today, e-mail us at contact@scalyr.com and I'm sure we can sort out any pricing concerns.
- Sanddancer 13y agoSo what happens when my email server goes down? One of the huge advantages nagios has is that there are plugins that'll send SMSes, plugins that'll send me phone calls. Hell, people have written scripts that let them call to get nagios alerts [1]. One of the huge advantages of having something hosted in-house, and with nagios, is that it can be configured to a level of precision I can't see your tool coming close to achieving. Yes, Nagios' configuration is ugly and occasionally requires sacrifices to elder gods. At the same time, I've never found any sort of monitoring/alerting I've needed done that it can't handle. As much as your service looks cool for a specific subset of monitoring, it is still missing half the hooks as to why nagios is stubbornly sticking around. [1] http://www.googlux.com/callnagios.html http://www.googlux.com/callnagios.html
- wdewind 13y agoStop using nagios, all you have to do is string together 6 random pieces of software, 2 of which don't exist yet!
- viraptor 13y ago2 of them are not available in nagios at all (graphing and anomaly detection) and 1 already sucks completely (UI), so I'm not sure this is a good way to look at the presentation.
- carterparks 13y agoMost nagios users configure the pnp4nagios plugin for graphing
- linker3000 13y ago...and some use Centreon, which bolts graphing and a better UI onto Nagios 'out of the box' http://www.centreon.com/ http://www.centreon.com/ ..or install via FAN: http://www.fullyautomatednagios.org/wordpress/ http://www.fullyautomatednagios.org/wordpress/
- kudu 13y agoIt isn't very clear to me which Nagios plug-ins are considered standard, and perhaps that sort of confusion is what creates all the FUD surrounding Nagios.
- malabar 13y agoUse http://mathias-kettner.com/check_mk.html http://mathias-kettner.com/check_mk.html Makes life so much better.
- drzaiusapelord 13y agoExactly. "Hey use Zabbix" is something only people who have never used Zabbix would recommend. Nagios, and a host of other popular software that's difficult to use, exist because the alternatives are so poor.
- blueskin_ 13y agoThat in one line: "But I don't like it because it doesn't do things my preferred way!". People use Nagios because it works, and it gets everything right, including config if you have any clue whatsoever how to set up a good object hierarchy. The only real problem with it is maintenance, an issue which Icinga resolved long ago.
- zwily 13y agoWe add/remove about 150 nodes every day to/from our monitoring system automatically via APIs. That use case has always sucked for me with Nagios. How would you do that?
- peterstjohn 13y agoI'd be interested to hear about that too - often Nagios can spend much of its time being reconfigured than actually monitoring…
- blueskin_ 13y agoIn my current job, it's done automatically with Puppet. My previous job was a lot smaller scale and hosts were manually added to a config file of hosts, but I had set up groups properly so they only needed a single hostgroup to inherit all their services, dependencies and contacts from there.
- darkandbrooding 13y agoAre you making the Nagios host pull (via NRPE) or are you asking the individual hosts to push (via NSCA)? I am trying to solve a similar problem. Given a dynamic population of hosts, each of which has a variable life span, I think that asking individual hosts to query their own state and then push that to a "monitoring receiver" is the most scalable, sustainable approach. At least, that's the theory I'll be testing this week.
- zwily 13y agoWe're not actually using Nagios. We use sensu because it was designed with this sort of dynamic environment in mind. (I'm trying to stay away from the "C" word. :)
- rjzzleep 13y agohas anyone tried icinga or opennms, and can comment on that ? https://www.icinga.org/nagios/feature-comparison/ https://www.icinga.org/nagios/feature-comparison/
- carterparks 13y agoI use icinga extensively. It has a better UI but still suffers from the same downfalls the presentation presents. With that said, the solution proposed seems to be incomplete and a step back from what icinga/nagios provide.
- comice 13y agoopennms is a bit of a culture shock if you're used to nagios. It works kind of inversely, in that it wants to auto-discover the servers and services to monitor itself. Frankly, it feels very much designed around snmp imo (which I'm not saying is a problem, but it's different to how we use nagios). It's also the opposite of nagios in that rather than lots of smaller moving parts, it is one big mega (java) process that does everything. Again, not necessarily a problem (though I happen to think so :), but different. I also found opennms to be VERY complicated. I suppose nagios is though, first time around. For some reason though, I really want to use opennms and keep going back to try it out, but eventually give up.
- seiji 13y agoOpenNMS is great to run in addition to a traditional monitoring system. Your traditional monitoring systems have hand-selected features to monitor and alert for. OpenNMS will just go out and discover everything you have (and graph everything without any intervention too). You probably aren't monitoring all the statistics on every interface of your switches (what? people have switches?), but just throw OpenNMS at your networking management subnet and it'll pick up everything for later review. You can use OpenNMS for alerting and inventory tracking, but I prefer more extensible tools for those. Just use OpenNMS as a largely hands-off sanity check of your existing monitoring and graphing systems.
- lafar6502 13y agoGreat idea. You can always use ugly Nagios for monitoring your great monitoring system built on top of RabbitMQ, Ruby, Elasticsearch, Redis and few other famous components.
- deleted 13y ago[deleted]
- js2 13y agoRejoinder - https://news.ycombinator.com/item?id=7340514 https://news.ycombinator.com/item?id=7340514
- gretful 13y agocame here to rebut, read your response, went away satisfied with it.
- ah- 13y agoI have wanted to try out collectd (https://collectd.org/ https://collectd.org/) for some time now, does anyone have experience with it?
- giulianob 13y agoI just set it up last week to push data to Graphite. It took a little bit of time to understand how to configure it and the docs have conflicting information in some places. Also, you will need to build it yourself or get a PPA if you're on Ubuntu 12.04 and want to use it with Graphite since the version that Ubuntu ships with doesn't support graphite. I haven't tried to write plugins for it but it comes with a lot out of the box and it's working well.
- seiji 13y agocollectd is great because takes very few resources (cpu/memory) to collect statistics very rapidly and you can log your results however you wish (back to graph collectors, straight to CSV for later processing, out to custom processes or network protocols for other services to consume). If you want to be lazy and not set up a complete graphing infrastructure, just run collectd, have it automatically log all your statistics, and use it with the bundled https://collectd.org/wiki/index.php/Collectd-web https://collectd.org/wiki/index.php/Collectd-web package to view your history when needed.
- jmccree 13y agoI use collectd and have found it to work rather well. The documentation could certainly use some work, but once you get figure things out it's pretty easy to configure. CPU usage is almost non-existent for the amount of data monitored. It's also pretty easy to write a custom plugin to collect whatever custom metrics you want. I use a script to expose the current values out as http/json for integration with Circonus for monitoring/alerts of key values. You can use whatever graphing tools you'd like, things are stored in rrd format. Recently I've been implementing salt stack, and collectd is really easy to automate config of. I've got salt fully configuring collectd, and then using the circonus api to setup monitoring rules automatically per applied states. It's a beautiful thing.
- rafekett 13y agowhile we're at it, let's let graphite die too. in a fire.
- markdennehy 13y agoSo... stop using a debugged and stable tool whose limitations and problems are well-known and understood and replace it with six bits of software duct-taped together, two of which aren't working yet (if they even exist), without any idea of how they interact when they hit edge cases. I mean, "don't use X, use ShinyX instead" is one thing (and most of the time it's a bad thing but it does occasionally turn up good ideas), but this is just So Much Worse...
- seiji 13y agoYou're presenting the MySQL argument. "Why should we switch since we know it fails in exactly these 1,000 different ways and we can fix these problems? Using something better has unknown failure scenarios!" Have you ever been woken up by a nagios page that automatically cleared after five minutes because the incoming queue was delayed past the alert interval? Have you ever had your browser crash because you click on the wrong thing in the designed-in-1996-and-never-updated nagios interface and had your browser crash because it dumps 500MB of logs to your screen? Have you ever had services wake you up with alert then clear then alert then clear again because some new intern configured a new monitor but didn't set up alerting correctly (because lol, they don't get paged, so who gives a flip if they copied and pasted the wrong template config, as is standard practice)? Have you had to hire "nagios consultants" to figure out how to scale out your busted monitoring infrastructure because nagios was designed to run on a single core Pentium 90? Being pro-nagios is like being pro-Russia, pro-North Korea, and pro-Rap Genius while arguing "but at least we know how bad they are and can keep them in line."
- volume 13y agoI think his point is there are tradeoffs, and I agree. On top of that, meaningful debate over what tool should be about what context you're in. This applies to the OP's slidedeck. To give context about my comment about context: * was nagios setup before you started the job? * did you setup nagios yourself? * is your internal process for managing nagios broken? * culturally do you work at a place where ops is an afterthought? * if Nagios is your technical debt do you have a way out? are you crushed by other commitments? Maybe it's more of a management/culture issue. ... hmm actually I should stop. From re-reading your comment, I can't tell how much of it is trolling (in a entertaining Skip Bayless, right wing radio, Jim Cramer kind of way).
- codingbeer 13y agoPresentations like this always reminds me why "DevOps" is very different from system administration.
- noja 13y agoSo use Shinken http://www.shinken-monitoring.org/ http://www.shinken-monitoring.org/, the Nagios rewrite in Python.
- scott_karana 13y agoShinken was specifically mentioned in the slides as just a "Nagios", and not solving the problem. Not sure whether that's true or not, but they did address it...
- noja 13y agoShinken is built to scale.
- rbc 13y agoI'll only address distributed monitoring. Use NSCA instead of NRPE. That bypasses the limitations of the Nagios active check scheduling. I have a wrapper that I use for that: http://rbcarleton.com/send_nsca_service_check.shtml http://rbcarleton.com/send_nsca_service_check.shtml Use some kind of automation system like CFEngine for the distributed scheduling. Some assembly required ;)
- cjlm 13y agoGood response article: https://laur.ie/blog/2014/02/why-ill-be-letting-nagios-live-on-a-bit-longer-thank-you-very-much/ https://laur.ie/blog/2014/02/why-ill-be-letting-nagios-live-...
- bassclef 13y agoI dropped nagios and switched to self hosted zabbix years ago..
- zimbatm 13y agoSensu is alright but it also has a few downsides: I'm not a huge fan of having yet another debian package with it's own version of ruby packaged. It does make the plugins easier to write though. Checks need to be installed on the client (like nagios). It means that some coordination is necessary when you want to add a new check on the server side. This is largely resolved when using a configuration management system but it doesn't seem clean to me. The sensu-community repo has a lots of checks which is great to get started, some of them need some ruby gem dependencies to work though. I had issues with malformed json config or rabbitmq disconnections which would crash the server. Because the debian packages uses the old sysvinit it wasn't restarting. Moved the init scripts to upstart and added json validation when generating the config and now it's fine.
- rbc 13y agoOne thing that is very strong is the need for a "host" with a Nagios service definition. This doesn't map so well on to environments like Amazon EC2 auto-scaling groups. You don't necessarily know the host names in advance. You wind up building Nagios plugins that can monitor a pool of hosts (using cloudwatch or whatever) and gives you some kind of aggregate status. It sounds like a kluge to push it into the plugin, but it does allow you to use the Nagios alerting, which is pretty well understood.
- madaxe_again 13y agoThe "main problem" with nagios is that it's configuration is godawful. The "main problem" with most nagios users is that they edit their configuration manually. Automate that shit. We use Nagios to monitor our infra (5000+ checks, hundreds of hosts), and chef maintains the config. Works without a hitch - and best of all - it's been running for years, and not once has anyone had to poke it with a stick. Yes, NetSaint is old, yes, the UI is worse to look at than Putin's crotch, yes, the plugin architecture is whimsical as all shit - but... IT WORKS AND YOU CAN RELY ON IT.
- roeme 13y ago^ This. I really don't understand how people claim that nagios can't scale.
- dbenhur 13y ago5000 checks and a few hundred hosts is a fairly tiny operation. At an order of magnitude above that you will start feeling real pain with Nagios.
- madaxe_again 13y agoYup, never claimed anything but - but it's big enough that automation is essential, lest you descend into some variety of lovecraftian horror-scape.
- k3oni 13y agoIt can't scale when we are talking 10000's of checks, triggers and hosts. This is why we moved to Zabbix: Number of items monitored: 77342 Number of triggers enabled: 25405 Try and scale Nagios to those values and let me know how it works. We tried and we let it go a few years ago.
- mcguire 13y agoOnce upon a time, back when I first installed NetSaint, the configuration was so bad that I learned OCaml to write a preprocessor to generate its configuration.[1] Nagios is the improvement! And yeah, it's horrible. [1] http://www.crsr.net/Software/nscc.html http://www.crsr.net/Software/nscc.html
- evantahler 13y agoI'm a big fan of monit (tried and true) + m/monit (web interface + more complex logging and anayltics) https://mmonit.com/ https://mmonit.com/
- deleted 13y ago[deleted]
- roeme 13y agoMy main experience from working with nagios for somewhere around 10 years now is that when people complain about it, they are either too lazy to read, or too inept to understand, the documentation or the architecture. (1) That being said, there _are_ limitations (but scaling is not one of them) to nagios, and the configuration is definitively not something you do in cute widdle config.yml. Combined with the recent negative developments with the corporation behind the Nagios trademark and the enterprise version - which the author fails to mention, and should be even more alarming - one should at least consider using and contributing to bareos, the (hopefully) true OSS fork of nagios (I will for future deployments). Oh hey, look at that, it would even pose the possibility to _improve_ the software. (I really don't see why an almost-complete rewrite of nagios should be necessary. Even after reading these slides(2)). (1) That includes the author. (2) Or rather especially after reading them.
- SmokeyMcPot 13y ago> one should at least consider using and contributing to bareos, the (hopefully) true OSS fork of nagios Bareos [1] is more an fork of Bacula [2], Isn't it? Did you mean Icinga [3]? [1] http://www.bareos.org/en/ http://www.bareos.org/en/ [2] http://www.bacula.org/en/ http://www.bacula.org/en/ [3] https://www.icinga.org/ https://www.icinga.org/
- roeme 13y agoOops, of course ! (Unfortunately, the edit allowed timespan on my comment has expired). The irony here is that with bacula a similar thing happened.
- martin_ 13y agoDo we use nagios to monitor the other 6 utilities? Or what about when the alerting gateway goes down?
- poulsbohemian 13y agoI have consulted on operational monitoring for many years, including with a customer that claims to have the largest deployment of Nagios anywhere (50K+ nodes). The author hits on many good points, right up until they suggest a solution. My advice to customers has long been that you can make any tool successful, but the tools are not what really matter. Too often I've seen customers invest $MM in tooling, and fail to understand that people and process around that tooling is the real challenge. Too often both the entrenched enterprise vendors AND startups in this space miss this too. When it comes to tooling, the problem that too many startups miss is that they repeat the patterns that the entrenched players formed decades ago, and fail to understand that the kind of monitoring tools like Nagios and its clones offer is but one piece of a comprehensive solution for all but the smallest of operations.
- rhizome 13y agoHah, you sure are a consultant: your last four sentences basically repeat themselves. ;)
- laichzeit0 13y agoNagios is pretty much a joke compared to most enterprise production monitoring tools, e.g. Wiley and Foglight. I always find it funny to read what people consider "monitoring" they're talking about a few disparate metrics and then complain that the alerting/paging sucks. If the tool can't trace a transaction end-to-end. I.e if a user visits your page or uses your application you need the ability to trace it from Http to EJB across any webservices and queues and ESB's right down to which queries were used in the database, if you can't do that you're using a shit monitoring suite. Knowing infrastructure metrics is useless without knowing if it's actually affecting end-users and in which use-cases.
- callesgg 13y agoI choose nagios cause there is currently nothing else on the market that is actually better. I think nagios is a piece of shit, but it is a working piece of shit.
- deleted 13y ago[deleted]
- nasalgoat 13y agoI hear a lot of people crapping all over Nagios, but none of the alternatives are any better. I recently had an opportunity to do a clean sheet build out for monitoring, so I evaluated Zabbix, Munin, and combos of statd/Graphite, etc. and none of them were better. That said, I have a stock Nagios base config that I can install and have monitoring in five minutes. The key to Nagios configs is to define hostgroups in one file, and then create config files for each host, assigning it to a group. Then you put the service definitions in a service file. Easy peasy.