11 ms·
Show HN: I built an open-source tool to make on-call suck less
Hey HN,
I am building an open source platform to make on-call better and less stressful for engineers. We are building a tool that can silence alerts and help with debugging and root cause analysis. We also want to automate tedious parts of being on-call (running runbooks manually, answering questions on Slack, dealing with Pagerduty).
Here is a quick video of how it works: https://youtu.be/m_K9Dq1kZDw https://youtu.be/m_K9Dq1kZDw
I hated being on-call for a couple of reasons:
* Alert volume: The number of alerts kept increasing over time. It was hard to maintain existing alerts. This would lead to a lot of noisy and unactionable alerts. I have lost count of the number of times I got woken up by alert that auto-resolved 5 minutes later.
* Debugging: Debugging an alert or a customer support ticket would need me to gain context on a service that I might not have worked on before. These companies used many observability tools that would make debugging challenging. There are always a time pressure to resolve issues quickly.
There were some more tangential issues that used to take up a lot of on-call time
* Support: Answering questions from other teams. A lot of times these questions were repetitive and have been answered before.
* Dealing with PagerDuty: These tools are hard to use. e.g. It was hard to schedule an override in PD or do holiday schedules.
I am building an on-call tool that is Slack-native since that has become the de-facto tool for on-call engineers.
We heard from a lot of engineers that maintaining good alert hygiene is a challenge.
To start off, Opslane integrates with Datadog and can classify alerts as actionable or noisy.
We analyze your alert history across various signals:
1. Alert frequency
2. How quickly the alerts have resolved in the past
3. Alert priority
4. Alert response history
Our classification is conservative and it can be tuned as teams get more confidence in the predictions. We want to make sure that you aren't accidentally missing a critical alert.
Additionally, we generate a weekly report based on all your alerts to give you a picture of your overall alert hygiene.
What’s next?
1. Building more integrations (Prometheus, Splunk, Sentry, PagerDuty) to continue making on-call quality of life better
2. Help make debugging and root cause analysis easier.
3. Runbook automation
We’re still pretty early in development and we want to make on-call quality of life better. Any feedback would be much appreciated!
- LunarFrost88 2y agoReally cool!
- tryauuum 2y agoevery time I see notifications in Slack / Telegram it makes me depressed. Text messengers were not designed for this. If you get the "something is wrong" alert it becomes part of history, it won't re-alert you if it's still present. And if you have more than one type of alert it will be lost in history I guess alerts to messengers are OK as long it's only a couple manually created ones, and there should be a graphical dashboard to learn the rest of problems
- aray07 2y agoYeah, I agree that slack is not the best medium for alerts. I think we it has somewhat become the default in teams is that it makes it easy to collaborate while debugging. I don't know a good way to substitute that and share information. What strategies have you seen work well?
- tryauuum 2y agoI might have been lucky, most of my companies were big enough to have a dedicated person to watch the dashboard 24/7. And a human will mostly know when it's a good idea to escalate and wake up the rest of the team I have no idea what's the best setup for small companies
- stackskipton 2y agoWhy? We send alerts to Slack and Pagerduty. Slack is to help everyone who might be working, PagerDuty alerts the persons who are actually in charge of working on it.
- Aeolun 2y agoYeah, I think it’s convenient. We use email, but for the same thing. If I inadvertedly break something, I’ll have an email in my inbox 5 minutes later.
- northrup 2y agoTHIS. Whispering into a slack channel off hours isn’t a way to get on-call support help nor is dropping alerts in one. If it’s a critical issue I’m going to need a page of some kind. Either from something like PagerDuty or directly wired up SMS messaging.
- lmeyerov 2y agoBig fan of this direction. The architecture resonates! The base lining is interesting, I'm curious how you think about that, esp for bootstrapping initially + ongoing. We are working on a variant being used more by investigative teams than IT ops - so think IR, fraud, misinfo, etc - which has similarities but also domain differences. If of interest to someone with an operational infosec background (hunt, IR, secops) , and esp US-based, the Louie.AI team is hiring an SE + principal here.
- racka 2y agoReally cool! Anyone know of a similar alert UI for data/business alarms (eg installs dropping WoW, crashes spiking DoD, etc)? Something that feeds of Snowflake/BigQuery, but with a similar nice UI so that you can quickly see false positives and silence them. The tools I’ve used so far (mostly in-house built) have all ended in a spammy slack channel that no one ever checks anymore.
- shahargl 2y agohttps://github.com/keephq/keep https://github.com/keephq/keep (disclaimer - i'm the maintainer)
- RadiozRadioz 2y ago> Slack-native since that has become the de-facto tool for on-call engineers. In your particular organization. Slack is one of many instant messaging platforms. Tightly coupling your tool to Slack instead of making it platform agnostic immediately restricts where it can be used. Other comment threads are already discussing the broader issues with using IM for this job, so I won't go into it here. Regardless, well done for making something.
- aray07 2y agoThanks for the feedback. We want to get something out quickly and we had experience working with Slack so it made sense for us to start there. However, the design is pretty flexible and we don't want to tie ourselves to a single platform either.
- FooBarWidget 2y agoTry Netherlands. We're Microsoft land over here. Pretty much everyone is on Azure and Teams. It's mostly startups and hip small companies that use Slack.
- satyamkapoor 2y agoStartups, hip small companies, tech product based companies. Most non tech product based or enterprise banks in NL are on Teams
- Aeolun 2y agoI really feel like the world would be a better place if it was illegal to bundle Teams like this…
- darkstar_16 2y agoDenmark is the same. Only the smaller startups use Slack. Everyone one else is on Teams.
- october8140 2y agoAs Slack is not end to end encrypted we and I imagine many other companies cannot use it.
- solatic 2y agoIn my current workplace (BigCo), we know exactly what's wrong with our alert system. We get alerts that we can't shut off, because they (legitimately) represent customer downtime, and whose root cause we either can't identify (lack of observability infrastructure) or can't fix (the fix is non-trivial and management won't prioritize). Running on-call well is a culture problem. You need management to prioritize observability (you can't fix what you can't show as being broken), then you need management to build a no-broken-windows culture (feature development stops if anything is broken). Technical tools cannot fix culture problems! edit: management not talking to engineers, or being aware of problems and deciding not to prioritize fixing them, are both culture problems. The way you fix culture problems, as someone who is not in management, is to either turn your brain off and accept that life is imperfect (i.e. fix yourself instead of the root cause), or to find a different job (i.e. if the culture problem is so bad that it's leading to burnout). In any event, cultural problems cannot be solved with technical tools.
- aray07 2y agoI completely agree that technical tools cannot fix culture problems. However, one of the things that I noticed in my previous companies was that my management chain wasn't even aware that the problem was this bad. We also wanted to add better reporting (like the alert analytics) so that people have more visibility into the state of alerts + on-call load on engineers. What strategies have worked well for you when it comes to management prioritizing these problems?
- dennis_jeeves2 2y ago>However, one of the things that I noticed in my previous companies was that my management chain wasn't even aware that the problem was this bad. Isn't that a cultural problem?
- djbusby 2y agoShow them the costs! Wasted time, wasted resources, wasted money. Show the waste and come with the plan to reduce the waste. Alerts, on-calls and tests are all waste reduction. "We're paying down our technical debt"
- maximinus_thrax 2y agoNice work, I always appreciate the contribution to the OSS ecosystem. That said, I like that you're 'saying out loud' with this. Slack and other similar comm tooling has always been advertised as a productivity booster due to their 'async' nature. Nobody actually believes this anymore and coupling it with the oncall notifications really closes the lid on that thing.
- aray07 2y agoYeah, unfortunately, I don't think these messaging tools are async. During oncall, I used to pretty much live on Slack. Incidents were on slack, customer tickets on slack, debugging on slack...
- maximinus_thrax 2y agoThat is correct, they are not. My former workplace had Pagerduty integrated with Slack, so I get it...
- lars_francke 2y agoShameless question tangential related to the topic. We are based in Europe and have the problem that some of us sometimes just forget we're on call or are afraid that we'll miss OpsGenie notifications. We're desparately looking for a hardware solution. I'd like something similar to the pagers of the past but at least here in Germany they don't really seem to exist anymore. Ideally I'd have a Bluetooth dongle that alerts me on configurable notifications on my phone. Carrying this dongle for the week would be a physical reminder I'm on call. Does anyone know anything?
- michaelt 2y agoA candy bar cell phone, paid for by your employer and handed to whoever is on call. People who don't want it can just forward it to their phone.
- crawfishphase 2y agoin this case a satellite enabled candybar. the disaster recovery policy and budget should be applicable here. make sure its able to share xg and satellite tunnel for maximum value. ensure the reporting system is satellite enabled also. added points if its sending alerts 2 your handy byod. Disaster recovery is a big deal in 2024. All sorts of factors make satellite redundancy valuable in todays reality: Coworkers on a hike or a boat, random 0-day stuff, and war can cut your normal internet.. i have experienced all of these and only in the last 4 years and more than 1 time on each topic. Train your users to destroy it in case of war as its trackable by military tech. Put a sticker on it. Check out stackexchange for questions like this tho?
- Flop7331 2y ago[flagged]
- lars_francke 2y agoYup, this is something we've thought about as well. We're a remote company so everyone would need to get their own but this option is definitely on the table.
- dclowd9901 2y ago> It reduces alert fatigue by classifying alerts as actionable or noisy and providing contextual information for handling alerts. grimace face I might be missing context here, but this kind of problem speaks more to a company’s inability to create useful observability, or worse, their lack of conviction around solving noisy alerts (which upon investigation might not even be “just” noise)! Your product is welcome and we can certainly use more competition in this space, but this aspect of it is basically enabling bad cultural practices and I wouldn’t highlight it as a main selling point.
- aray07 2y agoYeah, thats fair feedback. The main aim was to reduce the alert fatigue for on-call engineers and provide a way to get insight into the alerts at the end of the on-call shift. This way there is data to make a case that certain alerts are noisy (for various reasons) and we should strive to reduce the time spent dealing with these alerts. Fixing some of them might be as easy as deleting them but for others might need dedicated time working on them.
- deleted 2y ago[deleted]
- DanHulton 2y agoI agree with your intent and desire, but the fact is that this problem _keeps happening,_ and we can't fix it by advocating "well, just do alerts better." There's a lot of cultural inertia at a lot of places that leads to creating too many, too low-signal alerts, and fixing that across an entire company -- or, hell, industry -- is a magnificently tall order. However, installing a tool to specifically reign in those garbage noisy alerts is a potentially easy, significant win for the time and mental health of on-call engineers. I mean, it sounds like you can then afterwards go in and identify the alerts that are just noise, and having that data means you can take action. Maybe contact the teams that are writing the noisiest alerts, or prepare some data-driven engineering standards for the company, whatever. But that still falls into "fix the culture", which is famously hard to do by fiat.
- 2y ago
- theodpHN 2y agoWhat you've come up with looks helpful (and may have other applications as someone else noted), but you know what also makes on-call suck less? Getting paid for it, in $ and/or generous comp time. :-) https://betterstack.com/community/guides/incident-management/on-call-pay/ https://betterstack.com/community/guides/incident-management... Also helpful is having management that is responsive to bad on-call situations and recognizes when capable, full-time around-the-clock staffing is really needed. It seems too few well-paid tech VPs understand what a 7-Eleven management trainee does, i.e., you shouldn't rely on 1st shift workers to handle all the problems that pop up on 2nd and 3rd shift!
- Aeolun 2y agoI guess 7-Eleven management trainees know that their company is just as replacable for their employees as their employees are to them.
- deepfriedbits 2y agoNice job and congratulations on building this! It looks like your copy is missing a word in the first paragraph: > Opslane is a tool that helps (make) the on-call experience less stressful.
- aray07 2y agoderp, thanks for catching. It has been fixed!
- sanj001 2y agoUsing LLMs to classify noisy alerts is a really clever approach to tackling alert fatigue! Are you fine tuning your own model to differentiate between actionable and noisy alerts? I'm also working on an open source incident management platform called Incidental (https://github.com/incidentalhq/incidental https://github.com/incidentalhq/incidental), slightly orthogonal to what you're doing, and it's great to see others addressing these on-call challenges. Our tech stacks are quite similar too - I'm also using Python 3, FastAPI!
- jobtemp 2y agoWhy not use statistics? Been reading about xmr charts recently on commoncog. That might help for example.
- aray07 2y agoThanks for the feedback! I saw the incidental launch on HN and have been following your journey!
- Jolter 2y agoI wouldn’t say it’s particularly clever. It’s a fairly obvious idea to anyone who has worked with alerting through IMs. What it is, is /difficult/, because you really really need to avoid false positives. Probably lots of hard work involved. So kudos for making this work (if it works)!
- david1542 2y agoI'm curious about incidental :) how are you going to compete with other, well established IM tools like rootly, incident.io, firehydrant.com?
- EGreg 2y agoOne of the “no-bullshit” positions I have arrived at over the years is that “real-time is a gimmick”. You don’t need that Times Square ad, only 8-10 people will look up. If you just want the footage of your conspicuous consumotion, you can easily photoshop it for decades already. Similarly, chat causes anxiety and lack of productivity. Threaded forums like HN are better. Having a system to prevent problems and the rare emergency is better than having everyone glued to their phones 24/7. And frankly, threads keep information better localized AND give people a chance to THINK about the response and iterate before posting in a hurry. When producers of content take their time, this creates efficiencies for EVERY INTERACTION WITH that content later, and effects downstream. (eg my caps lock gaffe above, I wont go back and fix it, will jjst keesp typing 111!1!!!) Anyway people, so now we come to today’s culture. Growing up I had people call and wish happy birthday. Then they posted it on FB. Then FB automated the wishes so you just press a button. Then people automated the thanks by pressing likes. And you can probably make a bot to automate that. What once was a thoughtful gesture has become commoditized with bots talking to bots. Similar things occurred with resumes and job applications etc. So I say, you want to know my feedback? Add an AI agent that replies back with basic assurances and questions to whoever “summoned you”, have the AI fill out a form, and send you that. The equivalent of front-line call center workers asking “Have you tried turning it on and off again” and “I understand it doesn’t work, but how can we replicate it.” That repetitive stuff should he done by AI and build up an FAQ Knowledge Base for bozos and then only bother you if it came across a novel problem it hasn’t solved yet, like an emergency because, say, there’s a windows BSOD spreading and systems don’t boot up. Make the AI do triage and tell the differencd.
- snihalani 2y agocan you build a cheaper datadog instead?
- david1542 2y agoYou have plenty of options. some of them are open source https://signoz.io/ https://signoz.io/ https://coroot.com/ https://coroot.com/ Did you search tools that are cheaper than DD?
- mikeshi42 2y agowe've leveraged Clickhouse/S3 to build a cost effective alternative to Datadog at https://hyperdx.io https://hyperdx.io (OSS, so you can self-host as well if you'd like)
- protocolture 2y agoI feel like this would be a great tool for people who have had a much better experience of On Call than I have had. I once worked for a string of businesses that would just send everything to on call unless engineers threatened to quit. Promised automated late night customer sign ups? Haven't actually invested in the website so that it can do that? Just make the on call engineer do it. Too lazy to hire off shore L1 technical support? Just send residential internet support calls to the the On Call engineer! Sell a service that doesn't work in the rain? Just send the on call guy to site every time it rains so he can reconfirm yes, the service sucks. Basic usability questions that could have been resolved during business hours? Does your contract say 24/7 support? Damn, guess thats going to On Call. Shit even in contracting gigs where I have agreed to be "On Call" for severity 1 emergencies, small business owners will send you things like service turn ups or slow speed issues.
- nprateem 2y agoThat's why it's always at least double time for call outs
- Flop7331 2y agoIs this for missile defense systems or something? What's possibly so important that you need to be woken up for it?
- fragmede 2y agoOn the Internet, with the war in Ukraine, that's entirely possible!
- kgeist 2y agoSystem goes down or degrades in some other way at night and important customers with a different timezone get angry, threatening to leave? (happened with us a few times) But I would'nt use LLM for it due to hallucinations
- Flop7331 2y agoWhat's so important that your customers in other timezones feel like waking you up? If they're ready to walk that fast, you can't trust them not to ditch you for an alternative as soon as they find one.
- kgeist 2y agoWell, I guess it depends on the business. I forgot to mention that we're B2B. For example, suppose a large food chain or a major bank has an important exam scheduled for their employees on a specific day. If our platform has a blocking bug, no one can proceed (some may be sitting in the class) because the developers are too busy sleeping. Some of our clients are also airplane pilot certification authorities, which have stricter requirements. When there's an alert, you never know if it affects small clients or large clients too. You don't have to be a missile defense system to require a stable system where devs respond quickly...
- Flop7331 2y agoBut those are the kinds of scenarios where I imagine the sun still came up if comparable disruptions occurred prior to our current era of constant connectivity. We're too invested in the myth that our special problem can't wait half a day.
- throw156754228 2y agoI don't want to be relying on another flaky LLM for anything mission critical like this. Just fix the original problem, don't layer an LLM into it.
- aray07 2y agoI agree - fixing the original problem is the main motivation. We wanted to provide that awareness because a lot of teams arent fully aware how bad the problem might be (on-calls change weekly, there might be a bunch of other issues)
- amelius 2y agoA central service might be in a better position to classify messages compared to lots of individual agents.
- throwaway984393 2y agoDon't send an alert at all unless it is actionable. Yes, I get it, you want alerts for everything. Do you have a runbook that can explain to a complete novice what is going on and how to fix the problem? No? Then don't alert on it. The only way to make on-call less stressful is to do the boring work of preparing for incidents, and the boring work of cleaning up after incidents. No magic software will do it for you.
- voidUpdate 2y agoFiltering whether a notification is important or not through an LLM, when getting it wrong could cause big issues, is mildly concerning to me...
- mads_quist 2y agoFounder of All Quiet here: https://allquiet.app https://allquiet.app. We're building a tool in the same space but opted out of using LLMs. We've received a lot of positive feedback from our users who explicitly didn't want critical alerts to be dependent on a possibly opaque LLM. While I understand that some teams might choose to go this route, I agree with some commentators here that AI can help with symptoms but doesn't address the root cause, which is often poor observability and processes.
- Jolter 2y agoTelecoms solved this problem fifteen years ago when they started automating Fault Management (google it). Granted, neural networks were not generally applicable to this problem at the time, but this whole idea seems like the same problem being solved again. Telecoms and IT used to supervise their networks using Alarms, in either a Network Management System (NMS) or something more ad-hoc like Nagios. There, you got structured alarms over a network, like SNMP traps, that got stored as records in a database. It’s fairly easy to program filters using simple counting or more complex heuristics against a database. Now, for some reason, alerting has shifted to Slack. Naturally since the data is now unstructured text, the solution involves an LLM! You build complexity into the filtering solution because you have an alarm infrastructure that’s too simple.
- samcat116 2y agoThe alerts being sent to Slack are normally from one of those alert databases (such as Prometheus and AlertManager). Slack isn't the source of truth for them, just a notification channel.
- Jolter 2y agoOh, Prometheus is good for metrics but it doesn’t hold alarms in the Fault Management sense, though. It only keeps the metrics and thresholds, checks for threshold violations, and then alerts via some mechanism. If it were an alarm database, an operator would be able to 1. Acknowledge the alarm 2. Manually clear an alarm that was issued in error. Without those mechanisms, alarm handling becomes really difficult for an ops team, because now all you have is either a string of emails or a chat log.
- drivers99 2y agoThe wikipedia for page for fault management has a "see also" for alarm management, which looks extremely relevant as well.
- Arch-TK 2y agoWe could stop normalising "on-call" instead.
- sir_eliah 2y agoCould you please elaborate?
- Arch-TK 2y agoYes, increasingly companies are pretending like their SAAS needs to run with at least 99.999% uptime and so are insisting that all their engineers/programmers/whatevers must therefore be happy to be on-call on a rota for no extra pay because of vagueness in their contracts. Meanwhile they either have a global workforce so don't actually need to have anyone on-call or only have customers in countries they have employees in. It's bullshit. Either companies should be up front about this when hiring or it should be optional and paid. Or, they can use their engineering talent, just like telecoms companies have been doing for ever, to engineer their products to be more resilient and automate failure cases so proper remediation can wait until working hours.
- T1tt 2y agohow can you prove it works and doesnt hallucinate? do you have any actual users that have installed it and found it useful?
- T1tt 2y agois this only on the frontpage because this is an HN company?
- topaztee 2y agoco-founder of merlinn here: https://merlinn.co https://merlinn.co | https://github.com/merlinn-co/merlinn https://github.com/merlinn-co/merlinn We're also building a tool in the same space with the option of choosing your own model (private llms) + we're open source with a multitude of integrations. good to see more options in this space! especially OS. I think de-noising is a good feature given alert fatigue is one of the repeating complaints of on-callers.
- aflag 2y agoIt feels to me that using LLM to classify alerts as noisy is just adding risk instead of fixing the root cause of the problem. If an alert is known to be noisy and have appeared on slack before (which is how the LLM would figure out it's a noisy alert), then just remove the alert? Otherwise, how will the LLM know it's noise? Either it will correctly annoy you or hallucinate a reason it figures that alert is just noise.
- aray07 2y agoyeah, thats the goal of adding the context and the report - to hopefully bring awareness to the team that this alert should be removed. My rationale for flagging the alert was to help prioritization for the on-call (lets say there are multiple alerts going off at the same time)
- ozim 2y agoThat’s a people problem and you cannot fix people problems with tech. If no one cares to do the good job of managing alerts putting AI in front of it will not change that.
- jrochkind1 2y agoAn AI could help bring to your attention alerts that need managing. I like it for this better than for someone in the moment of receiving an alert deciding whether or not to pay it attention.
- ozim 2y agoIf someone ignores alerts they will keep ignoring them but now you automated part of ignoring with "AI" and human at the end still will ignore alerts the same. Writing it out makes me laugh because that's like something from Douglas Adams stories. Automated Ignoring System along with Infinite Improbability Drive.
- jrochkind1 2y ago
- c0mbonat0r 2y agoif this is open-source project how are you planning to make this a sustainable business? also why the choice of apache 2.0
- Terretta 2y agoNote that according to StackOverflows dev survey, more devs use Teams than Slack, over 50% were in Teams. (The stat was called popularity but really should have been prevalence, since a related stat showed devs hated Teams even more than they hated Slack.) Teams has APIs too, and with Microsoft Graph working you can do a lot more than just Teams for them. More importantly, and not mentioned by StackOverflow, those devs are among the 85% of businesses using M365, meaning they have "Sign in with Microsoft" and are on teams that will pay. The rest have Google and/or Github. This means despite being a high value hacking target (accounts and passwords of people who operate infrastructure, like the person owned from Snowflake last quarter) you don't have to store passwords therefore can't end up on Have I Been Pwned.
- nprateem 2y agoAlmost all alerting issues can be fixed by putting managers on call too (who then have to attend the fix too). It suddenly becomes a much higher priority to get alerting in order.
- ravedave5 2y agoThe goal for oncall should be to NEVER get called. If someone gets called when they are oncall their #1 task the next day is to make sure that call never happens again. That means either fixing a false alarm or tracking down the root cause of the call. Eventually you get to a state where being called is by far the exception instead of the norm.
- fnimick 2y agoI wish everyone shared your philosophy! I once worked at a company where it was expected to get 10+ pages per day, and worse, a configuration error by a customer success team would trigger an engineering page because the error handling didn't distinguish between a config problem and an actual system issue. It was insane.
- henryfjordan 2y agoDepending on the stakes this is a pretty dangerous attitude. The goal for oncall is to keep the website working, and if you're tuning for "never get paged" then you'll necessarily miss an incident eventually.
- cdchn 2y agoIf you make your goal as high availability as possible, and you only get paged on outages, then your goal should be to never get paged. You should be building resilient architectures, not being on firewatch duty.
- henryfjordan 2y agoThis is a classic developer vs business incentives misalignment. Developers don't want to ever be paged because they don't want to be bothered, but the business might be perfectly happy to pay you to be on firewatch duty. Consider a "low traffic" alert, how can you tell the difference between a slow period at 3am on a holiday vs a true outage? You can't without someone getting up and testing if the site is still up. (Maybe you can automate that check but there's always edge-cases you can't automate). OP seemed to suggest it's better to disable the alarm than to just suffer the false alarm every now and then. I doubt very much that the people paying you for the on-call service would agree though.
- jedberg 2y agoPeople do not understand the value of classifying alerts as useful after the fact. At Netflix we built a feature into our alert systems that added a simple button at the top of every alert that said, "Was this alert useful?". Then we would send the alert owners reports about what percent of people found their alert useful. It really let us narrow in on which alerts were most useful so that others could subscribe the them, and which were noise, so they could be tuned or shut off. That one button alone made a huge difference in people's happiness with being on call.
- tracker1 2y agoI worked on a small team that covered a relatively big site where there were so many alerts it was simply hard to track... They were all sent over email to a group list and most would just delete. I spent about 3 months, each day trying to triage into buckets based on activity and dealing with whatever was causing the most alerts each day. Some came down to just tamping out classes of 4xx errors that should never have been in the email/alert system to begin with. Others came down to indexes to reduce load/locking/contention on some db tables. Others still were much harder to dig into. Will say at the end of the 3 months, there was only a trickle of emails a day and the notifications were taken much more seriously after not being so overwhelming as to being ignored altogether. edit: This was just the first thing I did each day was deal with one problem, then moving to new feature work... It wasn't assigned, as the company would always prioritize new feature work, it was just something I did for my own sanity.
- aray07 2y agoYeah, that was one of the goals we had. We try to classify when an alert comes up and let the engineer give us feedback. We use that to generate a report so that teams have visibility into which alerts are causing the most amount of noise.
- Jolter 2y agoIf each and every alert has an owner, you’ve solved half of the cultural problem already. Good on you!
- 7bit 2y ago> * Alert volume: The number of alerts kept increasing over time. It was hard to maintain existing alerts. This would lead to a lot of noisy and unactionable alerts. I have lost count of the number of times I got woken up by alert that auto-resolved 5 minutes later. I don't understand this. Either the issue is important and requires immediate human action -- or the issue can potentially resolve itself and should only ever send an alert if it doesn't after a set grace period. The way you're trying to resolve this (with increasing alert volumes) is the worst approach to both of the above, and improves nothing.
- makmanalp 2y agoUnderrated oncall problem that needs solving is scheduling IMHO: - We have a weekday (2 shifts) / weekend (1 slightly longer shift including friday morning to allow people to take long weekends) oncall rotation as well as a group-combined oncall schedule which gets finnicky. - When people join or leave the rotation, making sure nothing shifts before a certain date or swapping one person with another without changing the rest and other things are a massive pain in the butt - Combine this with a company holiday list - usually there's different policies and expectations during those. - Allow custom shift change times for people in different timezones. - We have "oncall training" / shadowing for newbies, automate the process of substituting them in gradually, first with a shared daytime rotation and then on their own etc. - Make oncall trades (if you can't make your shift simpler) Gripes with PD: - Pagerduty keeps insisting I'm "always on call" because I'm on level N of a fallback pager chain which makes their "when oncall next" box useless - just let me pick. - Similarly, pagerduty's google calendar export will just jam in every service you're remotely related to and won't let you pick when exporting, even though it will in their UI. So I can't just have my oncall schedule in google calendar without polluting it to all hell.
- aray07 2y agoThanks for the feedback! I completely relate to PD scheduling issues and something that we want to take a look at as well.
- jpb0104 2y agoI love this space; stability & response! After my last full-time gig, I was also frustrated with the available tooling and ONLY wanted an on-call scheduling tool with simple calendar integration. So I built: https://majorpager.com/ https://majorpager.com/ Not OSS, but very simple and hopefully pretty straightforward to use. I'm certainly wide open to feedback.
- asdf6969 2y agoI don’t really understand the use case. If there’s a way to programmatically tell that it’s a false alarm then there must also be a way to not create the alert in the first place I’ve never seen an issue that’s conclusively a false alarm without investigating at all. Just delete the alarm? An LLM will never find something like another team is accidentally stress testing my service but it does happen Another perfect example is when the queen died and it looked like an outage for UK users. Can your LLM read the news? ChatGPT doesn’t even know if she’s alive I expect you will need AGI before large companies will trust your product.
- CableNinja 2y agoI get your sentiment, but theres another side of this coin that everyone is forgetting, hilariously. You can tune your monitoring! Noisy alert that tends to be a false positive but not always? Tune alert message to only send if the issue continues for more than a minute, or if the check fails 3 times in a row. Theres hundreds of ways to tweak a monitor to match your environment. Best of all? It takes 30 seconds at most. Find the trigger, adjust slightly, and after maybe 1-2 tries, youll be getting 1 false positive sometimes, and actual alerts when they happen, compared to 99% false alerts, all the time. Oh and did you know any monitoring solution worth its salt can execute things automatically on alerts, and then can alert you if that thing fails? Also, Slack is not a defacto anything. Its a chat tool in a world of chat tools