SMDigital Book
Chapter 16A metric is an argument, not a fact
© Bhumaha Solutions Private LimitedAuthor: B. Thirumoorthy
16

Part VI — Governing Performance and Investment

Measuring the Whole System

What should leaders measure to improve the factory without distorting it?

11 minute read2,302 wordsPublished · Edition 1.0

A metric is an argument, not a fact

Software organizations can count almost anything: commits, tickets, review time, deployment frequency, incidents, availability, cost, customer behavior, survey responses, exceptions, rework, and learning actions. The availability of a number does not make it a valid representation of production health.

The central question is:

What should leaders measure to improve the factory without distorting it?

The bounded proposition is:

A measure becomes governable when its construct, unit, decision, observation window, provenance, distribution, balancing evidence, burden, gaming risk, and review trigger are explicit.

This chapter uses flow, quality, reliability, value, risk, cost, and learning as seven coverage lenses. They are prompts for finding blind spots and tensions. They are not a validated taxonomy, composite score, weighting system, benchmark, or maturity sequence.

SPACE provides the software-specific foundation for this caution. It treats developer productivity as multidimensional and distinguishes individual, team, and system perspectives. It explicitly rejects the idea that one activity measure can represent productivity.C16-A01 EngThrive, a current Microsoft-authored preprint, offers a different first-party design using speed, ease, quality, and thriving, combining telemetry and surveys.C16-A02 Their difference is useful: even interested software frameworks select constructs for particular purposes. Neither validates this book's seven lenses or a universal dashboard.

C16.1 — Begin with the decision

A measurement system should start with a decision record, not a data source.

The record asks:

  • What decision will this evidence inform?
  • Who is authorized to interpret and act on it?
  • What construct is being represented?
  • What is the unit: person, team, Production Line, service, portfolio, or organization?
  • What population and period are included?
  • What alternatives are being compared?
  • What missing evidence could reverse the interpretation?
  • What action is allowed if the signal changes?

“Measure cycle time” is not yet a decision design. Cycle time may inform work-in-progress control, a service expectation, a constraint investigation, a staffing choice, or an investment question. Each decision needs a different population, window, decomposition, and balancing evidence.

The same value can support opposing interpretations. A shorter elapsed time may reflect removed waiting, smaller work, changed classification, excluded difficult cases, reduced assurance, or genuine improvement. A higher deployment frequency may reflect smaller safe changes, automated noise, a different product mix, or work split to satisfy a target. Interpretation depends on the construct and the production context.

DORA's current guidance illustrates this bounded approach. It frames delivery measures at team level, advises organizations to choose a framework that fits their goals, and records changes in metric definitions over time.C16-R04 DORA is an interested framework owner, not independent proof that adopting its measures improves every organization.

C16.2 — Keep levels distinct

Individual, team, line, service, portfolio, and organizational evidence are related, but they are not interchangeable.

An individual may complete little visible code while resolving an architectural constraint, mentoring a team, containing an incident, or preventing unsafe work. A team may move quickly while a shared approval queue limits the Production Line. A line may improve delivery while service reliability declines. A portfolio may report favorable averages while one region, user group, or operational role bears the burden.

Aggregation removes detail. That can be useful when the removed variation is irrelevant to the decision. It becomes dangerous when the variation contains the mechanism, affected population, or unequal effect being investigated.

Before aggregation, record:

  • the underlying units and inclusion rules;
  • the distribution, not only the average;
  • the effect of missing or censored records;
  • material changes in work type or complexity;
  • whether definitions are comparable across groups;
  • who becomes visible or invisible at the chosen level; and
  • which actions the aggregate is allowed to trigger.

No formula converts person-level activity into team productivity or team delivery into organizational value. Measures can be related through explicit hypotheses, but the relationship must be tested rather than assumed.

This is also a governance boundary. Operational telemetry collected to improve a production system does not automatically become suitable for employee ranking, compensation, surveillance, or discipline. Those uses change incentives, consent, burden, access, and required validity. They require separate authority and evidence.

C16.3 — Use seven lenses to look for omissions

The seven lenses are deliberately plural.

Flow

Flow evidence asks how demand moves from recognition to operated outcome. It may include elapsed time, waiting, work in progress, arrival and completion patterns, rework, blocked time, and constraint location. Chapter 9 owns the mechanics. Here the measurement question is whether the selected evidence represents end-to-end movement or only local activity.

Quality

Quality evidence asks whether the produced capability remains fit for its declared users, obligations, and operating conditions. Defects are one signal. Security, accessibility, correctness, maintainability, usability, and conformance can reveal different failures. Chapter 13 owns Quality at the Source; this lens prevents faster movement from being interpreted without evidence about the work produced.

Reliability

Reliability evidence asks what happens in operation: availability, degraded service, failed change, recovery, recurrence, and dependency behavior. A release process can complete successfully while the operated outcome remains fragile. Definitions, observation windows, and affected service boundaries matter.

Value

Value evidence asks whether the operated capability changes an outcome that matters to users, the mission, or the institution. Adoption, completion, satisfaction, task success, avoided burden, and outcome evidence may help, but none is automatically value. Value depends on whose objective is represented and what alternative would have occurred.

Risk

Risk evidence asks what uncertainty and consequence remain, for whom, and under which conditions. Counts of findings, controls, or exceptions are activity evidence. Risk interpretation requires exposure, consequence, evidence quality, treatment, residual uncertainty, and an authorized decision. Chapter 14 owns gate and exception mechanics.

Cost

Cost evidence asks which resources the system consumes over the relevant lifecycle. Labor, infrastructure, supplier, delay, assurance, recovery, coordination, migration, and retirement costs can fall on different owners and periods. Chapter 17 will connect these observations to economic choice. This lens only prevents a local saving from being reported as whole-system efficiency without examining displaced cost.

Learning

Learning evidence asks whether production experience changes future work. A retrospective, action item, or training event is activity. Retained learning requires a changed mechanism, owner, effective version, follow-up evidence, and a decision to retain, revise, replace, or retire the change.

The lenses overlap. A recovery event may affect reliability, cost, risk, value, flow, and learning. The purpose is not to count it seven times. The purpose is to make the intended interpretation explicit and expose evidence that a single headline would hide.

Evidence Status — Bounded. Current evidence supports multidimensional, decision-linked measurement and the hazards of single scores, incentives, and unexamined aggregation. The seven lenses are an author-synthesized coverage checklist. Independent longitudinal validation and distributional evaluation remain future evidence.

C16.4 — Pair every signal with a plausible failure

A balancing measure is not a ritual second number. It represents a plausible way that movement in the primary signal could mislead or cause harm.

If the decision concerns delivery time, inspect work size, excluded work, quality, assurance, operational outcomes, and burden. If it concerns reliability, inspect suppressed demand, reduced change, user-visible degradation, recovery cost, and manual work. If it concerns adoption, inspect task success, accessibility, coercion, displaced channels, and affected non-users.

The pairing follows a hypothesis:

desired movement
→ possible mechanism
→ plausible failure or displacement
→ balancing evidence
→ decision boundary

The balancing evidence can contradict the desired movement. That is its job. A dashboard that only confirms the preferred story is not a control system.

Distribution is part of balancing evidence. An average improvement can coexist with deterioration for a region, role, service, disability group, employment class, supplier, or on-call population. Where the data cannot support a valid group comparison, the limitation should be recorded rather than replaced with an overall average.

Burden also belongs in the evidence. Collecting, correcting, explaining, and responding to measures consumes work. Surveys can create fatigue. Telemetry can create privacy and surveillance exposure. Manual classification can shift administrative load to the people being measured. Measurement cost is not proof that the measure is wrong, but it is part of its fitness for a decision.

C16.5 — Expect behavior to respond

Measures do not remain external to the system. Visibility, targets, rewards, penalties, rankings, and executive attention change behavior.

The Wells Fargo sales-practice record is an extreme, non-software countercase. Regulator findings show how sales targets and incentives contributed to misconduct while aggregate reporting obscured the relationship between activity and customer outcome.C16-R01 It does not establish how often software metrics are gamed or predict comparable severity. Its authorized role is narrower: when a measure becomes a consequential target, the measurement system itself becomes part of the production design.

GAO's review of six public-sector performance-management pilots documented gaming and perceived fairness problems in bounded historical settings.C16-R03 A separate GAO grants review advised testing measures and data before connecting them to rewards or penalties and described reasons to revise accountability mechanisms as goals, technology, behavior, and decision needs change.C16-R02 These are not software-production outcome studies. They strengthen the general control questions:

  • Can the measured actor materially influence the outcome?
  • Can the signal be improved without improving the intended construct?
  • Which work becomes unattractive, invisible, or displaced?
  • Who can challenge the data, definition, or interpretation?
  • What happens when the measure is wrong?
  • Is the measure being used for a decision it was not designed to support?

Gaming is not always deception. People may reclassify work, split items, avoid difficult demand, delay recording, optimize the observation window, or redirect effort toward what the system rewards. Some responses reveal an ambiguous definition or harmful incentive. Investigation should distinguish manipulation, adaptation, learning, and legitimate disagreement.

Separate sources, separate roles

SPACE supports multidimensional and multilevel reasoning. It does not validate the seven lenses.

EngThrive illustrates a current organization-authored design combining telemetry, surveys, outcome-oriented measures, and a thriving guardrail. It remains a first-party preprint and does not supply independent causal or complete distributional validation.

DORA illustrates an evolving team-level delivery framework. It does not authorize individual ranking or a universal factory-health score.

Wells Fargo supplies an extreme regulator countercase about targets, incentives, and misleading aggregate signals. GAO supplies independent public records about testing, gaming, fairness, and revision outside software production.

These records answer different questions. Their populations and values are never pooled.

C16.6 — Give every measure a lifecycle

A measure should have an owner, definition, source, version, decision, review trigger, and retirement path.

Review is required when:

  • the decision or organizational objective changes;
  • the measured work, technology, or population changes;
  • data quality or comparability degrades;
  • behavior adapts to the signal;
  • burden or unequal effects emerge;
  • the measure loses sensitivity or becomes saturated;
  • a better source becomes available;
  • the interpretation repeatedly requires exceptions; or
  • the measure no longer changes any authorized decision.

These are triggers, not a universal review interval or retirement threshold.

At review, the owner can:

  • retain the measure and its current use;
  • revise the construct, definition, source, window, or interpretation;
  • restrict the measure to a narrower decision or population;
  • replace it with stronger evidence;
  • retire it while preserving the historical definition needed to interpret prior decisions.

Retirement does not erase the record. Historical decisions may depend on the definition and data version that applied at the time. Chapter 15's traceability discipline therefore applies to measures too.

The decision-linked measurement test

Leaders can ask:

  1. Which decision and authorized user need this evidence?
  2. What construct does the measure represent, and what does it not represent?
  3. What is the unit, population, observation window, and comparison?
  4. Which definition, source, transformation, missingness rule, and version produced it?
  5. Which of the seven coverage lenses are material to this decision?
  6. What plausible failure, displacement, burden, or gaming response accompanies the desired movement?
  7. What does the distribution reveal that the aggregate hides?
  8. Can affected people challenge the data and interpretation?
  9. What action is authorized, and what evidence would stop or reverse it?
  10. Which trigger will cause retention, revision, restriction, replacement, or retirement?

The test is a design aid, not a universal schema. A routine operational decision may need a small record. A consequential portfolio, workforce, safety, or regulatory decision needs stronger evidence, independent challenge, and recourse.

Figure F16.1 production specification: Decision-Linked Measurement Architecture

Figure F16.1 — A named decision leads to explicit construct and unit, seven unweighted coverage lenses, evidence provenance and distribution, balancing evidence and burden review, and a retain/revise/restrict/replace/retire decision. The figure is a bounded architecture, not a score or causal model.

What to remember

Production health is plural. One signal cannot represent the whole system.

Begin with the decision, not the available data.

Keep people, teams, lines, services, portfolios, and organizations distinct.

Use flow, quality, reliability, value, risk, cost, and learning as coverage lenses, not a composite score.

Pair every desired movement with a plausible failure, burden, distributional effect, or gaming response.

Label interested sources and bounded countercases.

Review and retire measures when their decision relevance, validity, behavior, burden, or context changes.

Measure to improve a named decision—not to manufacture a favorable number.

Continue the argument

From evidence to economic choice

Measures reveal conditions, constraints, outcomes, burdens, and uncertainty. They do not decide where scarce capacity should go. Chapter 17 turns from observation to choice: how to compare immediate demand, risk, discovery, lifecycle cost, option value, and compounding Factory Assets without pretending that every consequence has one monetary equivalent.

C16-A01: [C16-A01] Forsgren et al., “The SPACE of Developer Productivity,” ACM Queue 19(1), 2021. C16-A02: [C16-A02] Houck et al., “EngThrive: Make It Fast and Easy to Do Great Work,” public preprint, 2026. C16-R01: [C16-R01] CFPB, 2016 Wells Fargo sales-practice testimony; SEC, 2020 enforcement record. C16-R02: [C16-R02] U.S. GAO, Grants Management, GAO-06-1046, 2006. C16-R03: [C16-R03] U.S. GAO, Performance Management, GGD-98-162, 1998. C16-R04: [C16-R04] DORA, “DORA’s software delivery performance metrics,” current guide, retrieved 2026.

End of Chapter 16
Its construct, unit, decision, observation window, provenance, distribution, balancing evidence, burden, gaming risk and review trigger are explicit.
Return to contents