Subcapability 04 of 04 · High Performance Teams

Engineering Metrics & DORA

The four engineering metrics that correlate with organizational performance and shift management from narrative to data, connecting delivery velocity to financial results.

What is at stake

The board asks how long it takes for a change to reach the user. Engineering has no answer with a number. It knows it is "a few weeks." It knows it "depends." Without a measurement of delivery flow, every technology investment decision is made based on the intuition of who requests and the intuition of who approves. The cost of that imprecision goes into the IT budget that grows without clear results.

What it is, in practice

Engineering without measurement is management by opinion. Story points, velocity and burndown measure effort, not outcome. The board decides based on timeline, cost and margin. The gap between what engineering reports and what the business needs to know is where technology investment decisions lose credibility. More headcount looks like the solution when nobody measures where time is being consumed.

How we work

Measurable gains

What changes in the result when this subcapability matures.

Frequently asked questions

Why these four specific metrics and not others?

Because longitudinal research across software organizations in different sectors, sizes and geographies identified these four as the ones that correlate most consistently with organizational outcomes: revenue growth, profitability, customer satisfaction and employee satisfaction. Other metrics like story points and velocity measure activity, not outcome. The four engineering metrics measure the delivery flow of value and system stability, which are the two factors with a direct connection to business results.

How do you collect the four metrics without a manual process?

Most pipeline and incident management tools already have the necessary data. Deployment frequency and change lead time come from pipeline history with commit timestamp and deploy timestamp. Change failure rate comes from crossing deployments with incident openings in the same time windows. Restoration time comes from incident management with open and close timestamps. Instrumenting existing tools is the first step, not buying a new observability tool.

What performance range do you expect from mature teams?

High maturity teams deploy multiple times a day, have change lead time under one hour, change failure rate below 5% and service restoration time under one hour. Low maturity teams deploy monthly or less frequently, have lead time in weeks, failure rate above 45% and restoration time of a day or more. The distance between the two extremes is one to two orders of magnitude and is reducible with systematic work on process, platform and topology.

How do you prevent the team from optimizing the metric instead of the outcome?

By using the metrics as diagnosis, not as individual targets. The risk of turning the metric into a target is that the team finds ways to hit the number without achieving the outcome. Deployment frequency that rises because the team started making smaller and safer deployments is good. Deployment frequency that rises because the team artificially split deployments to count more events is bad. Reading the four metrics together, with output quality context, is what keeps the diagnosis honest.

How long does it take to have the metrics installed and visible to C-Level?

Basic instrumentation with automatic collection of the four metrics in two or three pilot teams can be operational in four to six weeks. The C-Level reading dashboard with translation to financial impact takes another two to four weeks of configuration and calibration. Within one quarter, C-Level has the first real baseline and the technology investment argument shifts from narrative to data.

Want clarity on where to invest first?

A complete technology capability assessment with an evolution roadmap connected to financial result.