SMDigital Book
Chapter 19Begin with a question, not a transformation
© Bhumaha Solutions Private LimitedAuthor: B. Thirumoorthy
19

Part VII — Introducing and Scaling the Model

Diagnose Before You Design

How can an organization identify the few system conditions that most constrain improvement?

10 minute read2,125 wordsPublished · Edition 1.0

Begin with a question, not a transformation

Organizations rarely begin change with no explanation. They begin with too many explanations.

Releases are slow because the platform is weak. Quality is poor because teams lack skill. Demand is fragmented because governance is unclear. Costs are rising because automation is insufficient. A new operating model, tool, structure, or transformation programme arrives already attached to the diagnosis.

That sequence is dangerous. When the preferred solution defines the problem, evidence becomes a sales aid. The organization gathers examples that justify the design, translates disagreement into resistance, and treats activity as validation.

The central question is:

How can an organization identify the few system conditions that most constrain improvement?

The bounded proposition is:

Diagnosis is a versioned inquiry record: a bounded question, separately labelled evidence streams, competing constraint hypotheses, disconfirming observations, declared uncertainty, and a next discriminating inquiry.

It is not a maturity score. It is not a search for one universally true root cause. It is not permission to redesign the organization.

The 2026 Magenta Book recommends defining an intervention and its context, synthesizing existing evidence, surfacing assumptions, examining alternative mechanisms and negative programme theory, and testing theories through multiple evidence sources.C19-S01 Its domain is public-policy evaluation, not software production. The transferable discipline is narrower: make the explanation visible enough to challenge.

C19.1 — Bound the diagnostic question

“Diagnose the software organization” is not an answerable question. It has no defined outcome, population, time, boundary, or decision.

A useful inquiry starts with a concrete tension:

  • Which condition most constrains the reliable completion of regulated changes for this service family?
  • Why does customer demand wait between authorization and production in this line?
  • Which shared capability contributes to repeated recovery delay across these products?
  • Why has adoption of this Factory Asset stalled among the teams it was intended to serve?

The question should name:

  • the outcome or decision the inquiry serves;
  • the demand, services, products, or Production Lines included;
  • the actors, systems, suppliers, and controls included;
  • the relevant time window and geographic scope;
  • the people affected by the present condition and a possible intervention;
  • the exclusions; and
  • the authority that will decide what happens next.

This boundary prevents a local observation from becoming an organization-wide verdict. It also reveals where the inquiry lacks access. A study of tool telemetry may omit outsourced work. A study of successful releases may omit abandoned demand. Interviews with managers may omit the people doing support, assurance, accessibility work, or unpaid coordination.

The boundary is not fixed forever. It is versioned. If evidence shows that the suspected problem crosses a supplier, policy, data, or organizational boundary, expand the inquiry explicitly and record why.

C19.2 — Keep four evidence streams distinct

Diagnosis needs triangulation, but triangulation does not mean mixing everything into one score. Four evidence streams answer different questions.

Production records

These include Work Orders, queue states, changes, deployments, defects, incidents, assurance decisions, dependency records, asset use, support requests, and outcome evidence. For each record, state its source, population, time window, definitions, missingness, access restrictions, and known instrumentation changes.

Production records show what the system recorded. They do not automatically show why something occurred, which work was invisible, how people experienced it, or whether the measured event caused the outcome.

Direct observation

Observation can expose waiting, handoffs, workarounds, coordination, repeated clarification, and unofficial control work that systems do not record. GAO's Agile Assessment Guide includes value-stream mapping, direct observation, and stakeholder participation among its assessment practices.C19-R01

Observation is situated. Record who observed, for how long, in which setting, with what access, and how their presence may have changed behavior. An observer sees a sample interpreted through their own concepts. The note should separate what was seen, what a participant said it meant, and what the observer inferred.

Stakeholder experience

People can explain intent, consequences, incentives, history, and unrecorded work. Include different roles and affected groups, not only sponsors or the loudest team. Record how participants were selected, which incentives or power relationships may shape testimony, where accounts disagree, and whose perspective is missing.

Experience is evidence. It is not ground truth. Repeated agreement may reflect a shared condition, a shared story, or a shared incentive.

Obligations and context

Law, policy, contracts, architecture, funding, geography, labor conditions, accessibility, public value, suppliers, and prior commitments constrain what the system can do. Some “inefficiencies” are deliberate protections. Some “local” delays are consequences of decisions elsewhere.

Keep these streams separately labelled before comparing them. Agreement across them can strengthen a provisional explanation. Conflict between them is often the most valuable finding.

Process mining reveals recorded behavior

Process mining can reconstruct paths, variation, rework, waiting, and conformance from event logs. It is attractive because it appears to convert a complex organization into observable flow.

The IEEE Process Mining Manifesto makes the prerequisites visible. Useful analysis depends on the trustworthiness and completeness of event records, clear semantics, appropriate case and activity concepts, correlation, privacy, and data quality.C19-S03

Before using a process-mining result, ask:

  1. What is the case? A request, ticket, change, incident, customer, or release?
  2. What counts as an activity, and who defined it?
  3. Which time is recorded: start, completion, ingestion, approval, or later correction?
  4. Which events are absent, duplicated, reordered, aggregated, or generated automatically?
  5. Did definitions, tools, teams, or instrumentation change during the window?
  6. Which work happens in conversation, documents, spreadsheets, supplier systems, or human judgment outside the log?
  7. Which people are represented, and which may be exposed by the analysis?

A process map is therefore a claim about recorded event behavior under declared semantics. It is not the whole production system. It cannot, by itself, infer motivation, experience, incentives, causality, or the value of the work.

Sparse data should remain sparse. Do not fill an empty interval with certainty, silently treat missing events as no work, or compare small populations as if their differences were stable. The diagnostic record should state what the data can and cannot discriminate.

C19.3 — Build competing constraint hypotheses

After establishing base observations, formulate more than one explanation.

Suppose a regulated change waits forty days between acceptance and release. Plausible hypotheses might include:

  • demand enters without an authorized outcome or usable acceptance evidence;
  • a scarce assurance capability creates a control queue;
  • the line depends on a shared environment with unstable availability;
  • large batches create late discovery and rework;
  • teams delay because incentives punish visible failure more than waiting;
  • the recorded waiting time is partly an instrumentation artifact.

Each hypothesis needs five fields:

  1. Mechanism: how the condition could produce the observed outcome.
  2. Expected observations: what else should be visible if it is materially constraining.
  3. Rival explanation: another mechanism that could produce the same pattern.
  4. Disconfirming evidence: an observation that would weaken or reject the hypothesis.
  5. Scope: where and when the explanation is expected to hold.

NASA investigation procedure provides a useful independent discipline. It describes establishing base facts, posing hypotheses, testing them against evidence, using elimination where justified, and acknowledging cases in which causal factors cannot be adequately determined from available evidence.C19-R02 NASA's definitions also treat evidence as material used to support or refute a hypothesis or finding, across documentary, interview, demonstrable, and physical forms.C19-R03

This is a safety-investigation context, not a software transformation recipe. The transferable lesson is epistemic: an explanation earns attention by surviving attempts to reject it, not by fitting a familiar framework.

Rejection is a result

A diagnostic programme often rewards confirmation. Teams are expected to find a transformation-sized problem, consultants are expected to justify an intervention, and sponsors are expected to demonstrate progress.

That incentive makes rejected hypotheses essential evidence.

For every hypothesis, preserve:

  • the evidence sought;
  • the result;
  • whether the hypothesis was retained, revised, rejected, or left unresolved;
  • who challenged the interpretation;
  • which uncertainty remains; and
  • what decision changed because of the result.

Rejection reduces the search space. Revision improves the mechanism. An unresolved result prevents false confidence. None is a failed diagnostic.

The Magenta Book's theory-based approach emphasizes coherence, sufficiently specific evidence, triangulation, alternative causes, critical reflection, peer review, and external scrutiny.C19-S01 Chapter 19 adopts those controls without claiming that they reveal a uniquely true cause.

Consensus is not the exit gate. A room can agree because participants share a mental model, power structure, vocabulary, or desired solution. The question is whether the retained explanation accounts for bounded observations across more than one evidence stream and remains credible after a documented challenge.

Maturity models are prompts, not verdicts

A maturity model can provide vocabulary, remind investigators to inspect neglected capabilities, and support a structured conversation. Its danger begins when ordinal categories become a composite diagnosis.

A single score can:

  • hide strong and weak evidence behind one number;
  • imply equal intervals between levels;
  • combine constructs that answer different decisions;
  • reward visible artifacts over operating behavior;
  • invite benchmark comparisons without a valid reference population;
  • turn a conditional path into a universal sequence; and
  • make the recommended solution appear to follow mechanically from the score.

If a maturity framework is used, treat each item as a prompt. Record the underlying observation, source, boundary, uncertainty, and relevance to the diagnostic question. Do not average the items. Do not infer that a higher level is always better. Do not use the result to rank teams or individuals.

The purpose is not to describe everything the organization lacks. It is to identify which provisional condition is most worth discriminating next.

The diagnostic record

A compact record can make the inquiry inspectable:

Question and boundary

State the outcome, decision, scope, time, affected groups, exclusions, inquiry owner, and deciding authority.

Evidence ledger

For each production record, observation, stakeholder account, and contextual obligation, record provenance, definitions, sampling, access, missingness, interpretation, and limitations.

Hypothesis tree

List competing mechanisms. For each, state expected observations, rival explanations, disconfirming evidence, and scope.

Challenge log

Record dissent, conflicts between streams, observer and participant limitations, privacy or surveillance risks, and interpretations changed through review.

Disposition

Mark each hypothesis retained provisionally, revised, rejected, or unresolved. Do not erase earlier versions.

Next inquiry

Choose the smallest evidence-gathering action that can discriminate between the leading explanations. That may be a targeted observation, a repaired definition, an expanded sample, an independent review, or a bounded live test.

The record should be proportionate. A local reversible question may need a short inquiry completed in days. A safety-, rights-, or enterprise-consequential decision may require independent competence, broader participation, controlled access, and stronger assurance.

Figure F19.1 production specification: Production-System Diagnostic and Constraint Tree

Know when to stop diagnosing

Diagnosis can become avoidance. Leaders may keep gathering data because action carries risk, because no explanation is perfect, or because the inquiry itself creates status.

Stop the current diagnostic cycle when:

  • the decision and boundary are explicit;
  • relevant evidence streams have been examined proportionately;
  • leading rival explanations have been tested;
  • missingness and dissent are visible;
  • no unresolved gap changes the safety or legitimacy of the next step; and
  • a small, reversible action can generate more discriminating evidence than further desk analysis.

The result is not “the organization has been diagnosed.” It is:

Within this boundary and evidence window, this hypothesis is sufficiently credible to justify this next bounded test, subject to these uncertainties and stop conditions.

The 2026 Test and Learn guidance positions small real-world tests, explicit assumptions, mixed methods, and adapt, scale, or reconsider decisions as parts of iterative learning.C19-S02 Chapter 20 owns the design and operation of that first Production Line. Chapter 19 hands it a hypothesis—not a mandate.

What to remember

Begin with a bounded question, not a preferred transformation.

Keep production records, direct observation, stakeholder experience, and context separately labelled.

Treat process mining as evidence about recorded events, with explicit semantic and missing-data limits.

Generate competing explanations and state what could reject each one.

Preserve rejected, revised, and unresolved hypotheses.

Use maturity frameworks as prompts, never as composite verdicts.

End with the next smallest discriminating inquiry or bounded test.

A diagnosis is credible when it can show how it might be wrong.

Continue the argument

From diagnosis to a first Production Line

The inquiry has produced a bounded hypothesis and its uncertainties. Chapter 20 turns that hypothesis into a live, reversible test through the first Production Line.

C19-S01: [C19-S01] HM Treasury, Magenta Book: Central Government guidance on evaluation, 2026. C19-S02: [C19-S02] HM Treasury, Test and Learn, 2026. C19-S03: [C19-S03] IEEE Task Force on Process Mining, Process Mining Manifesto, 2011. C19-R01: [C19-R01] U.S. GAO, Agile Assessment Guide: Best Practices for Adoption and Implementation, GAO-24-105506, 2023. C19-R02: [C19-R02] NASA, Procedures and Guidelines for Mishap Reporting, Investigating, and Recordkeeping, NPR 8621.1. C19-R03: [C19-R03] NASA, NPR 8621.1D Appendix A.

End of Chapter 19
A versioned inquiry record containing a bounded question, separately labelled evidence streams, competing hypotheses, disconfirming observations, uncertainty and a next inquiry.
Return to contents