22 ms·
Grafana Labs launches free incident management tool in Grafana Cloud
- matryer 4y agoI work at Grafana, so AMA about the tool :)
- buro9 4y agoI also work at Grafana Labs, and could just Slack you... but as you've asked... Incident is available in the free tier, that's awesome... are there any limitations on that at all? Is the free tier version of Incident as fully featured as the paid tiers?
- farhan0410 4y agoYes Grafana Incident is available (fully featured) in the free tier of Grafana Cloud
- matryer 4y agoHello :) The free version is fully featured, so you just limited to number of users in the free tier (three).
- deleted 4y ago[deleted]
- solarkraft 4y agoIs there a chance we'll see it open sourced / a self hosting option?
- matryer 4y agoNo plans currently, but a self hosted option seems reasonable. Although, most people like their emergency tech not hosted on their own tech :)
- aalbertson 4y agoBeing not hosted on my same tech is one thing, still being self hosted so I can externalize it for a federal implementation is another. Definitely needs to be self hosted for ALL components. :)
- matryer 4y agoYeah makes sense. I'll add your vote to the list :)
- jrockway 4y agoI dunno, I don't really mind self-hosting monitoring infrastructure. I basically pay for a website uptime checker to check that Alertmanager is working. If Alertmanager is down, obviously you have to manually check to see what else is down, but it doesn't fail open. I wrote a little glue to make this straightforward for anyone else who uses Prometheus/Alertmanager: https://github.com/jrockway/alertmanager-status https://github.com/jrockway/alertmanager-status This ensures that the website check checks the health of the whole alerting pipeline; Prometheus has an always firing alert, Alertmanager is set to send that alert to alertmanager-status, and alertmanager-status starts failing its external health check if it isn't seeing that alert firing at the configured interval. If one of [Prometheus, Alertmanager, alertmanager-status] fails, then your website health check fails.
- trog 4y ago> No plans currently, but a self hosted option seems reasonable. Although, most people like their emergency tech not hosted on their own tech :) FWIW our core application is hosted in AWS but we maintain our own Grafana infrastructure independently. So it's not hosted on our own tech, per se, though we're still responsible for keeping it online. This looks great & would also love to see a self-hosted option. Honestly the more stuff like this that gets rolled into the OSS Grafana it actually makes me both more likely to try it and then more likely to eventually end up on the managed Grafana Cloud, as I will inevitably get sick of trying to maintain our own separate infra & the business case for centralising in Cloud makes more and more sense.
- mrtimbo 4y agoAny plans for Teams integration? We recently switched all our bots over from Slack.
- matryer 4y agoYou don't need Slack to use the tool, but yeah, a Teams integration is on the list, and will probably drop early next year.
- prepend 4y agoAny plans to have this FedRAMP certified so it can be used in US federal government incident management?
- SkoChippy 4y agoNo immediate plans for FedRAMP, but this may be on the roadmap in future quarters.
- DrRobinson 4y agoIt looks really cool and since we already use Grafana it would be a good fit for us, but for on call purposes Slack isn't very useful. If we were to migrate from PagerDuty to this, we would need an app that can override do not distrub and wake people up. Do you have any plans for any such app?
- drc 4y agohi, disclaimer I work at Grafana, We have plans to build a native mobile app for ios & android for OnCall that would let you achieve this over the next few months. OnCall is a separate product from Incident. It's available via OSS and Cloud. Incident and OnCall work well together, or you can use either as standalone!
- thayne 4y agoI'm a little curious how this will work with self hosted OnCall. Will the user need to set up their own push notification accounts for apple and Google, will it used a centralized service from Grafana, will it have a background service that polls the hosted OnCall service, or something else?
- jjtang1 4y ago
- nikolay 4y agoAll good except that Grafana Cloud is super expensive when you consider it per metric. This probably is the most expensive service per bit of data!
- bcjordan 4y agoInteresting, I'm using Grafana Cloud for just a few Prometheus metrics at the moment and have found it reasonable so far so am interested in what scale up looks like. I'm curious—what other sorts of services are you referring to in your comparison?
- divygoel 4y agoHey @nikolay! I work at Grafana Labs & focus on pricing - would you be up for a 15 minute chat to discuss this further? If so, feel free to either drop me a note at divy.goel@grafana.com or let me know how best to reach you :)
- gaffneyc 4y agoWe recently swapped our metrics to Grafana Cloud and were really surprised (despite being documented) that pricing is based on samples per minute not metrics series. So, for example, if we send a metric every 15s (the Prometheus default) then we get charged as if that were four separate metrics. Support was very helpful explaining everything and they reversed the charge but it still feels weird.
- oxfordmale 4y agoRecently you pulled one of our Grafana plugins without notice as you had upgraded this to an Enterprise license. We are more than happy yo pay, however, pulling production support without reaching out to negotiate a license is a d** move. We suffered an outage of several days while scrambling to get the payment approved. Luckily we didn't suffer any major outage in that window.
- igetspam 4y agoAre you planning any posts on comparing your new incident tool to other services? We currently use incident.io and are happy with it but we pay a lot for Grafana Cloud right now so it's worth considering if we can reduce spend elsewhere. Edit: We're happy with incident.io but free is compelling if the product is good and having a single view for observability is useful
- farhan0410 4y agototally understand that and great shout. happy to pull something together for ya if there are particular workflows you are most interested in comparing farhan.manjiyani@grafana.com
- matryer 4y agoThey're both great tools :) Lots of similarities, and plenty of differences.
- sjwhitworth 4y agoHey, incident.io CEO here. Glad to hear you’re happy with the product. The people at Grafana are great - congrats on the launch! Will have to take the product for a spin sometime :)
- farhan0410 4y agoThanks Stephen (product marketing lead for Incident) - we are also big fans of what y'all are building!
- altdataseller 4y agoAnyone replacing PAgerduty with this?
- buro9 4y agoI think that would be Grafana OnCall https://grafana.com/products/oncall/ https://grafana.com/products/oncall/ Internally we (Grafana Labs) already have replaced PagerDuty and are using it for our teams running critical systems.
- deleted 4y ago[deleted]
- donavanm 4y agoHow do you/users programmatically quantify MTTR (and related metrics) per incident, or in aggregate? Although it shades towards problem management this would seem necessary to achieve the claim of “reduces mean time to repair (MTTR).” Bonus questions, are you tracking or driving improvement in the related times for detection/response/mitigate/recover? Disclosure: Principal at AWS currently in a similar apace. Though I ask in a personal capacity and interest.
- matryer 4y agoHey, thanks for your question. The tool keeps track of declaration and resolution times by watching when the status is changed. It also lets you manually specify when the incident really started, and when it actually ended. We can use this data to measure a few things, and watching how this changes over time helps us figure out if we're getting better or worse, on average. We want to be careful what we incentivise by default, and we're actively working on this area. The data is going to be available for people to build their own visualisations (in Grafana). I'd be very interested to hear your thoughts too?
- donavanm 4y agoSorry for the delay, was traveling on holiday. In short yeah, using the incident status as the implied times makes sense for the bulk of cases. Totally agree on picking out signal from the users inherent actions, but allowing them to provide more specific data when they know better. Digging in a little further Im personally interested in moving past the incident data and inspecting the incoming alert(s) and related telemetry/metric/alarm data. For example think of the alarm definitions like “five 1m datapoints with a value above 0.1.” There’s a good argument to count impact (and incident duration) from that first datapoint > 0.1. Then theres the delta from metric processing to alert to incident creation. On the backend theres frequently a delta between mitigating impact and actual incident resolution, again I think getting back to the source alarm/alert/metric data would get us a more accurate view of operations and customer impact.
- CSMastermind 4y agoSeems great if you're already on the Grafana platform. One thing I'd say is that I find the "react with a robot emoji on slack to add information to the timeline" as a little kluge, hopefully that's not the only mechanism for doing that. Also does this tool have a postmortem workflow? I didn't see one in the documentation and that seems like an important part of the incident response process.
- drc 4y agoThanks for feedback re: robot emoji. You can also use a backslash command if you want to add a new piece of text to the timeline from Slack. re: postmortem workflow. The timeline view is built to help postmortems, one of the ways we're doing this is making it easy to paste the info from the timeline as rich text or markdown from the timeline into your post mortem workflow. You'll see that on the top right of the timeline view. We have a lot more ideas in this area and will be investing in this. Curious if there is specific features you'd like for postmortems?
- lstamour 4y agoAnother suggestion re gathering data or threads from Slack: using Message Shortcuts for greater visibility/discoverability? https://api.slack.com/interactivity/shortcuts https://api.slack.com/interactivity/shortcuts Might need to combine this with Slack modals for adding details (if it lets you do this)
- oxfordmale 4y agoAre you going to upgrade this feature to an Enterprise license one day and then revoke access without a grace period? This happened to one of our Grafana plugins and resulted in a several day outage while we scrambled to sort out the payment. As your company has shown zero respect for its customers, I will not be using any of your systems.To be absolutely clear it is fair to charge for any of your products, however, if you change it from freemium to paid you can't just pull the plug without reaching out.
- RichiH 4y ago
- aglazer 4y agoThis looks great and cool to see more innovation in the space. We've been using Rootly https://rootly.com https://rootly.com and love it.
- jjtang1 4y agoThank you for the kind words, Aaron. Been a pleasure partnering with the Taplytics team! We work with 100s of companies like Canva, Grammarly, OpenSea and others to help build a consistent incident response process on Slack if you're interested. Happy to give you the no-BS sales demo. FWIW - we are big fans of Grafana and have a native integration (think automatic Grafana metric/dashboard snapshots into #incident channel.
- xwowsersx 4y agoI'm looking into Grafana Cloud currently. We run a few services across 4 different environments. I'd like to have a single place to view metrics as well as response times for various API endpoints, metrics related to RDS, etc. Also interested in incident management tool. We have around 30 EC2 instances running currently but will be scaling that up further. What can I expect in terms of total pricing? Perhaps hopping on a call with someone for Grafana would make the most sense?