Subcapability 03 of 05 · Platform Engineering

SRE & Reliability

SLOs and error budgets that translate reliability from a technical conversation into a quantitative contract, giving the board the decision model it needed to invest in availability with defined criteria.

What is at stake

The system went down. The team wakes up at 2 AM. Fixes it under pressure. Nobody documents what happened. The following week, the same pattern. Without an SLO that defines what is acceptable, without an error budget that measures how much instability the business can absorb, each incident starts from scratch and the company pays the same downtime cost indefinitely.

What it is, in practice

Reliability is an economic decision. Going from 99.9% to 99.99% availability costs exponentially more. Without SRE, the board never knows whether it overpays or underpays for reliability, and the trade-off becomes a fight between dev and ops without a shared language. Downtime has a direct price. One hour of unavailability in a mid-market operation represents tens of thousands in uncaptured revenue and trust recovery costs that appear in no incident report. Organizations that actively manage error budgets reduce downtime-related costs by up to 40%, according to reliability engineering practice analyses.

How we work

Measurable gains

What changes in the result when this subcapability matures.

Frequently asked questions

How do you define the right SLO for a service?

Start from business impact, not technical capability. The right question is: what percentage of unavailability or degradation does the user notice and the business consider unacceptable? 99.9% means just under 9 hours of downtime per year. 99.99% means under one hour. The cost difference between the two is exponential. The right SLO balances the cost of unavailability against the cost of additional reliability.

What is an error budget and how does it change the dynamic between dev and ops?

Error budget is how much instability the SLO permits in a period. If the SLO is 99.9% availability over 30 days, the error budget is 43 minutes of downtime. While the budget is available, the team can deploy with confidence. When the budget is depleted, the priority shifts to stability. This mechanism eliminates the fight "dev wants speed, ops wants stability" because both work with the same explicit budget.

Does blameless postmortem work in a culture where the first reaction is to find someone to blame?

The change starts with the document structure, not with the culture. A postmortem template that asks for systemic root causes, event timeline and prevention actions, with no field for a responsible name. That structure already directs the conversation to the system. The culture follows when the process produces better results: the same incident stops repeating.

How do you calculate downtime cost to justify SRE investment?

A defensible approximation starts with transaction volume per hour on the service and the conversion rate or average value per transaction. Multiply by the percentage of users affected during the downtime. Add the support cost generated by the incident. For B2B services, the cost of breaching contractual SLA enters the calculation. That number, presented to the board alongside historical incident frequency, justifies SRE investment with criteria that IT budgets rarely have.

How long does SRE take to produce visible results?

First improvements in restoration time appear within weeks, with structured runbooks and on-call rotation with clear accountability. SLOs connected to the board and error budgets that change deployment behavior take one quarter to establish. Sustained reduction of recurring incidents through structured postmortem is visible within three to six months.

Want clarity on where to invest first?

A complete technology capability assessment with an evolution roadmap connected to financial result.