5 ms·
Is there a common framework how to come up with the values (e.g. 99.9%) set as SLOs in the first place? I currently want to do that for a service built on seve
by phkx 5y ago
Is there a common framework how to come up with the values (e.g. 99.9%) set as SLOs in the first place?
I currently want to do that for a service built on several cloud services with their respective SLAs and my approach is to go through the combined probabilities to get an effective error rate. That‘s a bottom-up approach.
I‘d combine that we a top-down derivation of what SLOs are required from the business side. If the first number doesn’t fulfill the business requirements with some buffer, we‘ll need to redesign. How do others do it?
- LeonB 5y agoThe SLO should be tied to customer satisfaction as much as possible. I think you need to collect a bunch of metrics (potential SLIs) for a period of time and then say — in which windows were things bad enough that they caused real issues? And of the metrics we gathered which ones actually show the issue? Then use those metrics as your SLIs and set the SLOs such that they reflect that reality.
- phkx 5y agoI like the practicality of analyzing where the customer had pain and adjusting SLOs accordingly. Our system is not open to customers yet, so we‘re lacking historic data with real load (besides system tests, but they currently don‘t include edge cases/chaos). Also, we‘ll probably need to have SLAs from the start, which would be derived from the SLOs, so I need something beforehand. —- SLA: service level agreement, values of KPIs promised to customers SLO: service level objective, internal target values for those KPIs, typically slightly more demanding than the SLA SLI: service level indicator, measured values of the KPIs to check against SLOs/SLAs
- LeonB 5y agoIt’s good to gather things now and record metrics that are meaningful. SLOs will come later. Or, if you create SLOs now expect you’ll have to create whole new SLOs when things begin. But in the meantime you can get slick ci/cd, unit, integration, end to end and performance tests. Maybe feature flags, containers and orchestration, with failover, self-healing, global distribution, DORA metrics again, monitoring, alerting, dashboards, visualisations - and then you’re ready for hello world. Or find features customers like.
- bradknowles 5y agoIf you haven’t publicly launched yet, then it’s hard to work backwards from what the customers actually want. What you may have to do is try to theorize who your main customers would be and what they would want, and then work backwards from there. In such a case, I would encourage you to set the SLAs relatively lower, at least until you can gain some real knowledge from actual customers. You might also want to approach some prospective customers and bring them into a private preview version (protected by NDAs), so that you can start gathering that data earlier. Also, ship early. Ship an MVP earlier than you thought possible. It’s okay if it’s all held together with spit and bailing wire, at least you’ll be able to start gathering data sooner, so that you can start the real business of actually building what the customers truly want.
- Raed667 5y agoWe have been experimenting with using Web-Vitals [0] as SLOs. The idea is that they represent the closest user satisfaction. If anyone else has come up with better metrics I'd love to hear about your feedback. [0] https://web.dev/vitals/ https://web.dev/vitals/
- Jensson 5y agoStep 1: Measure you current uptime. That is your baseline SLO, to not make things worse. Step 2: Evaluate how much money you would earn by increasing uptime, and evaluate how much the processes required to do that would cost. Incrementally raise SLO levels as long as you benefit from it.