7 ms·
I know people like wallboards and monitors but we found them anti-pattern. If you find yourself looking at a wallboard/dashboard, it should already be an automa
by CSDude 7y ago
I know people like wallboards and monitors but we found them anti-pattern. If you find yourself looking at a wallboard/dashboard, it should already be an automated alert.
- dexterdog 7y agoI find them great as a first look when there is a problem because you can often pinpoint the problem just by looking at the board.
- alex_d 7y agoWe are thinking about monetizing alert feature as a browser extension or PWA for mobile based on Monitoror Core API :) A wallboard is useful to monitor project builds, CI servers, even production things.
- sm4rk0 7y agoBut visualising the data and alerting are two different things.
- geofft 7y agoYes, which is why you shouldn't use wall boards for alerting, only for visualization. https://demo.monitoror.com/?configUrl=https://monitoror.com/assets/demo.monitoror.com-config.json https://demo.monitoror.com/?configUrl=https://monitoror.com/... is full of things that aren't visualizations at all (no graphs, no sense of whether things are abnormal but not past an alerting threshold, etc.) and are in fact alerts (the website is fine, one PR failed, the QA nodes are ... doing something but there isn't enough space to see what is wrong). If you want some graphs, great. If you want your team to look up every few minutes and poll some graphs (or worse, some colored rectangles) to figure out what they're supposed to be doing, consider that polling is usually the wrong approach. (To be clear, this is a criticism of the choice of demo data, not of the product overall. A product like this has its uses, but "our alerting system is people looking up at the TV" is not one of them.)
- OJFord 7y agoWhat would your alert be for # open PRs (an example in the demo linked from posted page)? How often would it fire? Whatever the answer, that's a different thing from this. Both have their place.
- CSDude 7y agoIf you just want to have a nice visualization to look at some numbers, fine. But, if you want to detect problems, it's ineffective. I saw too many companies do it to actually monitor the state of things and find out problems with charts, numbers, traffic lights etc.
- throw_away 7y agoYou can do both. Especially at the beginning of a system's lifecycle and you don't really understand its behavior yet. Lots of times, people wandering by have said hmm, that doesn't seem right… Later, as we learned more, these hunches evolved into more advanced automated alarms.
- virgil_disgr4ce 7y agoA "nice visualization" is not necessarily just a "pretty"/"shiny" thing to show off to people. Human beings are highly visual creatures with outstanding visual pattern recognition abilities. Maybe you personally don't get anything out of them but the value of visualization is proven. Here are a few sources to get you started: https://www.csgsolutions.com/blog/15-statistics-prove-power-data-visualization https://www.csgsolutions.com/blog/15-statistics-prove-power-...
- OJFord 7y agoBut that's my point, it isn't for alerting about problems, some things have a 'status' that might be interesting, but isn't a problem, or something to fire an alert on necessarily. You could have unintrusive notifications (inaudible etc.) to 'alert' to such statuses I suppose, if they were kept in view and not 'dismissed' (whatever that means for the medium they came in) - but then really you're just implementing a version of something like this Monitoror in your inbox, phone notification tray, Telegram channel, or whatever. You're not going to rip out logging, prometheus, or services' that this connects to own UI just because you have alerting, so I don't see why you would this. It's like prometheus & grafana for higher level stuff. (Of course you could use those tools for this sort of monitoring too, but that's not really the point.)
- AgloeDreams 7y agoYou know, for some people I think that's true and for others it's not. There is real value in making some data reactive rather than proactive in communication. Knowing current active traffic, open PRs, time til build is done, all that kind of stuff is 'I would like to see it/check it...but I do not want it to interrupt me.' People who deal with tens of interruptions at that level are clearly not very productive. On the other hand, for the site returning non-200 or for API issues, that should be an alert, for sure. Kinda surprised that Slack or MS Teams isn't in this market.
- wjossey 7y agoStrongly disagree. Understanding your metrics is a key part of so many roles, from devops, to product teams, to marketers... Yes, you should be automating alerts whenever possible. Yes, you should be putting up key metrics in a visible place so everyone can see how the product is performing. I can’t tell you how many times I caught an issue because I knew our metrics backwards and forwards, but it didn’t trip an alert threshold. Not every issue follows a pattern easily defined in a check, and human brains are incredible computers capable of helping to fill in that gap.
- geofft 7y ago> I can’t tell you how many times I caught an issue because I knew our metrics backwards and forwards, but it didn’t trip an alert threshold. So how many times was an issue missed because you weren't in the office, or because you were looking at your own screen and not dashboards at the moment? Humans are incredibly powerful, but our whole job as SREs is to make things reliable, repeatable, and scalable. We're doing an industry-wide migration from elegantly hand-crafted LAMP stacks running SSH to Kubernetes and infrastructure-as-code, not because you can't fix problems with SSH (you can, and you can usually fix them faster and better) but because you can't scalably fix problems with SSH. Similarly, if a human found an issue and alert didn't trip, I'd count that as a bug/missing feature in the monitoring. It's valuable while you're still small and working out your monitoring to keep a human in the loop - but at some point you need to get rid of that single point of failure. By all means, rely on a human to figure out where your alerting is lacking (just like you rely on a human to write the infrastructure-as-code), but you should eventually not rely on human intervention to actually keep incidents from happening.
- _jal 7y agoYou're both right. Instrumentation and alerts are vital - they leverage inhuman persistence, patience and low cost. But alerts do not substitute for a deep understanding of how your systems work. A number of the more useful "pre-crime" alerts we have derived from that - if I hadn't been elbow-deep in our systems long enough to notice certain behaviors have non-obvious second- and third-order effects downstream, we wouldn't have the alerts at all.
- scoutt 7y agoI don't think it is intended to be stared at it 8 hours-straight. I thinks it's more like a clock: you look at it several times a day, and not only when you hear an alarm.
- tilolebo 7y agoClocks are an anti pattern... Why would you want to have the time displayed permanently? It's such a distraction for developers. Just set automated alerts for lunch and end of day and that's it.
- dvtrn 7y agoClocks are an anti-pattern Well this is a first for me...
- scoutt 7y agoWell, I have a clock in the wall in front of me that permanently displays the time. I check the time several times a day, for example to check how much time left I have to do something before lunch or going home, or a having a meeting. I don't know why are we discussing the practical uses of a clock. I can't imagine a life where one is allowed to look at a clock only when an alarm or alert is triggered.
- gtyras2mrs 7y agoSounds like sarcasm to me. (Since folks somewhere above say that monitoring dashboards are an anti pattern and what you need is alerts)
- lioeters 7y agoCalendars, clocks, real-time notifications, and video chats are all anti-patterns, distracting developers from their zone of genius. Just send a concise email at the beginning/end of the day/week. (One can only dream..)
- jrockway 7y agoI think wallboards can be interesting. Do you want an alert if your site is suddenly trending on Twitter? If latency and error rates are good, probably not. Would you be interested if you walked by and noticed? Probably.
- wpietri 7y agoNot at all. Alerts serve a different purpose. One of the most important things a team needs over the long haul is a feel for their system. Many people refer to this as mechanical sympathy. And the way you develop that is long-term exposure to rich data. Alerts are the red and yellow lights on your dashboard. But you get mechanical sympathy by listening to the sound of the engine, feel of the road, and the smell of things when you take a peek under the hood. There are a lot of ways to achieve mechanical sympathy, of course. And information radiators are easily misused; you have to have the right information shown in the right ways for people to develop a correlative, intuitive understanding of what they've built. But nobody develops mechanical sympathy by looking at dashboard lights alone.
- TeMPOraL 7y ago> you have to have the right information shown in the right ways for people to develop a correlative, intuitive understanding of what they've built Lots of things have to be right for this to work, unfortunately, and company dashboards I've seen so far tend to be nowhere near it. For instance, the dashboard refreshed $PERIOD only makes sense if you're showing data that updates $PERIOD, and if you can respond to changes in that data $PERIOD. $PERIOD = "in realtime" or "every minute" or "hourly" or whatever is relevant in a given context. If you're looking at the dashboard much more frequently than the data changes, you're wasting time. If the data changes much more frequently than you're looking at it, you're likely to miss things, as 'geofft mentions elsewhere in the thread. And if you can't react to the data roughly as fast as it's updating, there's no point in looking at it so often. All those periods - recording, observing and reacting - must be roughly similar for the always-on dashboard to be useful, relative to generating reports every now and then. Panels full of lights and charts work on fighter jets or on the bridge of the Enterprise, because the pilots/crew are in a tight feedback control loop with their dashboards. (WRT. reacting in time, there are also error bars to consider. For instance, people on a diet are advised to weigh themselves weekly and not daily, because body mass varies by +/- 2kg during the day, so a naïve person checking weight daily would get fixated on those random oscillations. It's easier to tell regular people to reduce measurement frequency than to explain to them what a low-pass filter is and how is it relevant here. I have a feeling there's plenty of dashboard misuse that amounts to that too.) -- Speaking of the Enterprise and "getting the feel for the system", there's something that I'd like to try one day: make a monitoring tool that translates various system metrics into background sounds, creating an ambience similar to the one you hear on the Enterprise-D[0][1]. I feel a somewhat unobtrusive mix of background noises would be better to develop "the feel for the system" than a visual dashboard. Real-life examples of this are combustion engine's RPM, or spinning rust hard drives, if anyone still remembers those. -- [0] - https://www.youtube.com/watch?v=UKBvaOLDem0 https://www.youtube.com/watch?v=UKBvaOLDem0 - the bridge [1] - In Enterprise's engineering, there's a well-known pulsating sound of the warp core; I can't find a good enough YouTube video (whatever there is, apparently got broken by YT's audio compression). This background pulsing correlated to the speed Enterprise was traveling with.
- sunbear-lover 7y agoBy that logic a speedometer is an anti pattern and your car should just send up an alert when you're speeding... since when is getting accurate real-time information a bad thing?
- jedberg 7y agoThat’s... actually true. It doesn’t matter how fast you’re going unless you’re speeding. And it distracts you by making you look down. The only reason we don’t have that yet is because the car doesn’t know the speed limit everywhere all the time.
- timdorr 7y agoFunny thing, this is actually a feature in Teslas. You can set it to chime once the speed limit is exceeded (in areas where it knows the limit). Although, I've never seen anyone turn that on.
- kube-system 7y agoI strongly disagree. I can think of a ton of reasons why a driver may need (or even be legally required) to know their speed regardless of speed limit: * when speed restricted by equipment (trailer, temporary spare, etc) * when observing advisory speeds * when observing minimum speed requirements * as a reference for judging appropriate speeds under inclement conditions * as a reference for judging appropriate acceleration/deceleration rates when entering/exiting the roadway
- jedberg 7y agoOf course an alert system would have to be able to understand all those things. That's why we don't have that kind of system. A single number in isolation is rarely useful. Graphs with trends are useful. Alerts are useful. The only reason we don't have alert based speeds is because it can't get all the necessary information to make a useful alert, so we compromise by telling you the number. > as a reference for judging appropriate acceleration/deceleration rates when entering/exiting the roadway A perfect example of why a graph would be ideal here, not a single number.
- sirtoffski 7y agoI’ll chime in here to say we use both at work. In a NOC at a medium-sized ISP, we are getting hammered with alerts 24/7. Some are not urgent, while others need to be actioned much faster - I mean 100G transit link down is no good. We’d receive an automatic email about a large circuit going down, we’d also receive a ticket about it; sometimes people dont look at the tickets closely enough, other times people get distracted with other topics, issues, etc. Having a large screen with interface status monitoring has proven to be effective enough; for example, someone walks by the monitor and says “why is this thing red, is it supposed to be?... and we immediately know one of the larger interfaces is down. In an ideal world, we would not need it because every ticket will be diligently dealt with.... however in a real world, having a big red part of the screen flashing had proved quite effective.
- sparrish 7y agoIf you're getting alerts for non-actionable events, you need to do a better job of tuning your monitors and alerts. Alerts shouldn't be sent about anything that doesn't require an action.
- C1sc0cat 7y agoYeh any one from Google here if you like to allow us to fine tune the alerts from GSC - I came in today and found 87 non useful alerts in my inbox. I will have a look at the tool and have a play - I assume you can have multiple pages :-) Would be cool to monitor the looks at GA " 1 2 3 4 5 …. Many" sites I have an interest in
- sirtoffski 7y agoWell the thing is alerts are indeed for actionable events. For example many remote locations have an on-site battery backup, which would supply power in an event of loosing commercial power. Those are actioned in terms of notifying field teams and deciding whether a specific location needs to be placed on a generator. Imagine a hurricane disrupted commercial power grid and there are thousands of “site on battery” alerts; somewhere among them there is also an alert for OSPF down between two core switches. Having a monitor with a large red warning saying “Link X at location Y is down!” - is a pretty effective way to not miss important notifications. I mean playing devil’s advocate one might say “Then your alerts should have better filtering system with the important ones staying at the top of the page”... which is true. A lot of smart design features can render dashboards less relevant - however when there aren’t enough resources in a DevOps team to implement those solutions, a simple dashboard can go a long way!
- RBerenguel 7y agoI've spotted interesting "things" from idly looking at our dashboard while chatting with coworkers (and more than a few were interesting enough to warrant a lot of investigation and double-checking of metrics, providers and stack). They were not alert-able, or not very easily unless we wrote some complex time series analysis system for our internal metrics.