An engineering indicator can rise and, at the same time, hide that the business got worse. That is why the best technology capability metric is not the one that climbs fastest. It is the one that maps to a business number, resists being gamed, and is read at the system level.
The AI era made this distinction urgent. In the DORA report, every 25% increase in AI adoption came paired with individual productivity gains and, simultaneously, a drop in stability and delivery throughput, an effect attributed to larger change batches. In a controlled study with experienced developers working on mature repositories, participants took roughly 19% longer when using AI, even though they predicted a 24% gain and believed afterward they had been 20% faster. The individual indicator improved. The system result did not.
Below are the three tests a metric must pass: mapping to a business result, resisting gaming, and being read at the system level. Then, why benchmarking against the market has become a ceiling too low to serve as a target.
Require the metric to map to a business number
A metric is only worth what it predicts about revenue, cost, or risk. The problem is not a shortage of metrics. It is the missing layer that translates a technical indicator into a business number. A study of value stream practices found that fewer than 15% of organizations connect flow metrics to business outcomes such as revenue, cost, or retention. Those that do outperform the rest in both agility and profitability.
In practice, this means tying each indicator to a declared financial consequence. Change lead time connects to time-to-revenue. Change failure rate connects to incident cost and risk. Deployment frequency connects to market learning velocity. An indicator that does not close this loop is operational data, not a capability metric. This translation is the same exercise covered in how to measure technology ROI and in financial impact of technology on the business.
Read at the system level, never the individual level
A capability metric describes the delivery system, not the person. The DORA metrics, lead time for changes, deployment frequency, time to restore service, and change failure rate, separate elite from low-performing teams by orders of magnitude. Read at the system level, they indicate whether the organization learns fast and recovers well. Read at the individual level, they become surveillance and lose meaning.
The SPACE framework, academic in origin, was designed precisely to prevent single-dimension measurement and recommends never using a single isolated indicator. The rule is straightforward. Combine indicators from different dimensions and read the set at the team and flow level. The controlled study from the AI era proves the cost of ignoring this. Individual productivity rose while system throughput fell, and anyone watching only the individual did not see it. The team structure that this design demands is explored in high-performance technology teams.
Choose the metric that resists being gamed
Easy-to-collect indicators tend to be easy to manipulate. Lines of code, commit count, open pull requests, and story points measure activity, not value, and create perverse incentives. Teams learn to move the number without moving the result, and trust in the measurement collapses. Industry surveys show that roughly two thirds of developers do not believe the metrics used to evaluate them reflect their actual contribution.
AI compounds the problem. When code volume grows artificially, commit and line counts look excellent while value decelerates. The same vice appears in internal platforms, where counting delivered components says little about result. The test is straightforward. For each metric, ask how a team acting in bad faith would inflate it without delivering value. If the answer is easy, the metric is fragile. Choosing indicators that resist gaming is part of what is covered in how to measure technology maturity and in the engineering effectiveness theme.
Why benchmarking against the market became a low ceiling
The evidence closes the argument. Beating the market average has become easier because the average fell. In the DORA report, the high-performance group shrank from 31% to 22% in a single cycle, while the low-performance group grew from 17% to 25%. More organizations are deteriorating in software delivery, not improving.
The consequence for the board is direct. A deflated external benchmark turns relative position into a comfortable illusion. A company can rise in the ranking while its own capability stagnates, simply because the field regressed. The only honest measure is whether the capability moves a specific business number over time, not whether it beats an average that is falling. The best technology capability metric, therefore, is the one that survives three questions. Does it map to revenue, cost, or risk? Does it resist being inflated without delivering value? Is it read at the system level, not the individual level? An indicator that fails any of the three does not deserve its place on the dashboard.





