Part VII — Introducing and Scaling the Model
Launch the First Production Line
How should leaders turn the model into a credible, bounded experiment?
The first line is a test, not a miniature destiny
A diagnosed constraint does not authorize a transformation. It authorizes a question.
Chapter 19 ended with a provisional statement: within a defined boundary and evidence window, a particular mechanism is credible enough to test. Chapter 20 turns that statement into operation.
The temptation is to call almost anything a pilot. A new tool is demonstrated to a cooperative team. A showcase receives exceptional leadership attention. Easy work is selected. Specialists shield the effort from ordinary dependencies. Early movement is announced as proof. The temporary arrangement then becomes the model everyone else is expected to copy.
That is not a credible first Production Line.
The central question is:
How should leaders turn the model into a credible, bounded experiment?
The bounded proposition is:
The first Production Line is an inspectable feasibility-and-learning intervention: meaningful demand, a diagnosed mechanism, explicit alternatives, real operating conditions, declared opportunity cost, pre-agreed decisions, and a legitimate path to continue, reframe, or stop.
It does not prove Software Manufacturing. It does not identify the effect of every mechanism inside a coherent line. It does not authorize federation or scale.
The 2026 Test and Learn guidance recommends making assumptions explicit, agreeing decision criteria before testing, documenting adaptations, interpreting early findings cautiously, and using proportionate evidence under real conditions.C20-S01 Its policy context does not prescribe a software operating model. Its transferable discipline is to design the test around a decision before sunk cost and advocacy reshape the criteria.
C20.1 — Select meaningful, bounded demand
The first line must be meaningful enough to reveal the production system.
If it serves only synthetic tasks, friendly internal users, or demand with no material dependencies, the line may test a tool while hiding the conditions that matter. If it begins with the most irreversible, safety-critical, rights-affecting, or enterprise-wide demand, learning may expose people to unacceptable consequence.
Selection therefore balances:
- a real user or operational outcome;
- enough end-to-end scope to expose coordination, controls, support, and evidence;
- a demand volume that can produce interpretable observations;
- known dependencies that the line must actually navigate;
- safeguards proportionate to consequence;
- reversibility or containment where uncertainty is high;
- access to affected users and operators;
- a decision owner able to continue, reframe, or stop; and
- enough duration to encounter ordinary variation without inventing a universal timebox.
Record why this demand was chosen and what was excluded. Identify whether the line is unusually visible, well staffed, colocated, technically modern, politically protected, or relieved from shared obligations. Those conditions do not invalidate the test. Hiding them does.
Selection also creates an ethical boundary. Name the people who could carry additional effort, interruption, surveillance, migration risk, service degradation, or exclusion. Define recourse, rollback, incident authority, and stop conditions before exposure begins.
GOV.UK's beta guidance provides a bounded operational example: begin with limited real-user exposure, learn from feedback, ensure the team and support staff can sustain iteration, address accessibility and whole-journey conditions, and keep a legacy service operating during transition where necessary.C20-R02 This is not a universal phase model. It illustrates how “real” and “bounded” can coexist.
Establish the decision before the baseline
A baseline is useful only when it serves a decision.
Begin by recording:
- the diagnosed mechanism;
- the outcome or decision the line will inform;
- the observations expected if the mechanism is material;
- evidence that would weaken it;
- the line boundary and affected groups;
- the alternatives; and
- the decision date and authority.
Then define the baseline.
The baseline includes more than current performance. It should describe:
- business as usual — what is expected if current arrangements continue;
- do minimum — the smallest change that could meet the bounded objective;
- definitions, population, time window, trends, seasonality, and missingness;
- current demand, waiting, completion, quality, recovery, user, support, risk, cost, and workforce evidence relevant to the decision;
- known changes already in motion; and
- how evidence will remain comparable if instruments or definitions change.
The Green Book treats business as usual as an active benchmark, requires genuine alternatives, and defines opportunity cost as the value of the next-best use of resources.C20-S02 Chapter 20 borrows these questions, not the UK's full public-appraisal machinery.
Without alternatives, the first line is compared with an imaginary world in which nothing else could have improved. Without trends, regression and seasonal variation can look like impact. Without missingness, silence becomes success.
C20.2 — Record the intervention as a changing system
A Production Line is not a single treatment. It combines demand, people, Work Orders, Work Centers, Work Rooms, Factory Assets, environments, controls, evidence, governance, and feedback.
That coherence is the point operationally. It is also an attribution limit.
Before operation, record the intended mechanisms:
- which demand and authorization changes;
- which work states or handoffs change;
- which capabilities or assets are introduced;
- which policies, Quality Gates, or evidence requirements change;
- which roles and decision rights change;
- which support, migration, or training is supplied;
- which existing systems remain; and
- what the line is expected to make easier, safer, faster, clearer, or more valuable.
During operation, version this record. A line learns and changes. Record:
- intervention version and date;
- who authorized the change;
- evidence and rationale;
- unusual staffing or executive escalation;
- expert support not available to other lines;
- workarounds and temporary exemptions;
- instrumentation changes;
- unresolved dependencies;
- incidents, complaints, accessibility failures, or excluded users;
- work displaced or transferred; and
- the forecast that will later be compared with the outcome.
Test and Learn warns that adaptation can obscure which version generated a result and can overweight early positive findings.C20-S01 Traceable changes allow learning without pretending the intervention remained fixed.
Operate the whole line under real conditions
A credible line must encounter ordinary production.
That includes:
- real demand and authorization;
- discovery and changed understanding;
- integration with existing services and data;
- accessibility and different user conditions;
- security, privacy, assurance, and legal obligations;
- release, operation, support, recovery, and retirement implications;
- shared platforms, suppliers, and scarce specialists;
- planned and unplanned work;
- staffing changes, leave, and on-call burden; and
- the offline or human channels surrounding the software.
GAO's Agile Assessment Guide describes incremental development with continuous evaluation of functionality, quality, and customer satisfaction.C20-R01 It supports the limited idea that useful software evidence is generated through repeated operational encounters. It does not prove that Agile or a Production Line reduces risk in every context.
Real operation does not require careless exposure. Use progressive access, parallel running, feature controls, rollback, monitoring, supervised work, limited cohorts, or other safeguards appropriate to the consequence. A first line that cannot be operated safely at any useful boundary may reveal that the organization needs a different test or prerequisite capability.
Do not protect the showcase from every shared dependency. If a central assurance queue, supplier interface, legacy environment, or scarce platform capability constrains ordinary work, hiding it defeats the inquiry. Record whether the line resolves, routes around, or receives exceptional priority over the constraint.
Measure feasibility, not victory
The evidence panel should match the decisions.
Feasibility
Can the line operate as designed? Are roles, skills, environments, support, controls, and evidence available? Which mechanisms are stable, fragile, or dependent on extraordinary help?
User and operational experience
Can affected people complete the whole journey? Which workarounds, errors, delays, burdens, exclusions, and support needs appear? Who is not represented?
Flow and quality
How do bounded demand, waiting, completion, rework, failure, recovery, and predictability compare with the declared baseline and trend? Chapter 16's measurement boundaries apply.
Risk and assurance
Which hazards, incidents, exceptions, control burdens, residual uncertainties, and recourse records appear? Did safeguards operate?
Economics and opportunity cost
What capacity, money, assets, and leadership attention were used? Which work was delayed or displaced? What maintenance and support obligations were created? Which costs moved to another group?
Workforce and distribution
Who gained agency or skill? Who absorbed coordination, surveillance, interruption, or emotional burden? Which regions, roles, users, or suppliers experience different effects?
Learning
Which hypothesis was retained, revised, rejected, or left unresolved? Which forecasts were wrong? Which new evidence changes the next decision?
This is not a composite pilot score. Unlike constructs should not be averaged into green. Material safety, rights, or legitimacy failures may stop a line even when flow improves. A feasible line may still lack evidence of effectiveness.
CONSORT's extension for pilot and feasibility trials distinguishes feasibility from effectiveness, asks for prespecified progression criteria and unintended consequences, and permits proceed, amend, or do-not-proceed outcomes.C20-A01 A Production Line is not a randomized health trial. The transferable control is that a small test should answer feasibility questions rather than overclaim an effect it cannot estimate.
C20.3 — Decide before sunk cost decides
Agree the decision conditions before operation.
Continue the bounded line
The line can keep operating within its current scope when safeguards work, the operating design is sufficiently stable, the evidence supports continued learning or service, and the opportunity cost remains authorized. Continuation does not prove effectiveness or transfer.
Reframe and retest
Revise the diagnosed mechanism, selection, intervention, boundary, evidence, or safeguard when the test exposes a wrong assumption or unstable design. Record what changed. Begin another bounded evidence cycle.
Stop and preserve learning
Stop when consequence exceeds authority or safeguards, affected users are harmed, opportunity cost is no longer justified, the mechanism is rejected, feasibility is absent, or the line survives only through unsustainable exception. Preserve the evidence, assets, obligations, and retirement work. Stopping is not erasing the test.
Evaluate wider use
Authorize a separate inquiry into transfer when the line is stable enough to study and the organization has a reason to consider other contexts. That inquiry needs new populations, dependencies, costs, power effects, and counterfactual design. It is not a synonym for scale.
Test and Learn recommends that continue, adapt, or stop decisions be assessed against criteria agreed in advance to reduce over-commitment.C20-S01 The criteria are guides exercised through governance, not automatic thresholds that replace judgment.
The pause is part of the evidence
The Department of Veterans Affairs began deploying a modernized electronic health-record system in 2020 and extended it to additional sites. In April 2023, VA paused further deployments after veterans and clinicians reported that the system was not meeting expectations. VA then focused on improvements at initial sites before planning later deployments.C20-R04
GAO reported incremental configuration, safety, performance, and support work, while also identifying unresolved configuration requests, weak cost and schedule evidence, and a missing baseline and target for one system-impact metric.C20-R03
This is not a first Software Manufacturing Production Line, and the records do not establish that pausing caused improvement. The case demonstrates narrower lessons:
- live users can invalidate rollout assumptions;
- deployment progress is not the same as readiness;
- a pause can be an accountable decision rather than failure to execute;
- improvement activity does not remove unresolved economic or measurement gaps; and
- resuming deployment requires fresh readiness evidence.
Evidence Status — Bounded. Current evidence supports explicit feasibility questions, real-world testing, prespecified decisions, alternatives, opportunity cost, and a legitimate pause or reframe. It does not establish the effectiveness of Software Manufacturing or predict transfer from one Production Line.
The first-line record
Before launch, ask:
- Which Chapter 19 hypothesis is being tested?
- Which decision will the evidence inform?
- Why is this demand meaningful, bounded, and ethically acceptable?
- What is excluded, and how does selection limit interpretation?
- What are business as usual and do minimum?
- Which definitions, trends, and missing evidence form the baseline?
- Which mechanisms, roles, assets, controls, and supports will change?
- Which affected people, safeguards, recourse, rollback, and stop authority apply?
- Which capacity and work are displaced; what is the next-best use?
- Which feasibility, user, operational, quality, risk, economic, workforce, distributional, and learning evidence will be gathered?
- Which continue, reframe, stop, and review conditions are agreed now?
- What evidence would justify only a new inquiry into wider use?
Figure F20.1 production specification: First-Line Launch and Learning Framework
What to remember
Select real demand without choosing either a toy or an unacceptable hazard.
Disclose selection, exclusions, unusual support, and showcase conditions.
Compare with business as usual and do minimum.
Record opportunity cost and displaced work.
Operate the whole line under real conditions with proportional safeguards.
Version the intervention, including workarounds and extra support.
Decide continue, reframe, or stop against criteria agreed before operation.
Treat wider use as a new evaluation question—not proof of scale.
The first line succeeds when it produces an honest decision, including the decision to stop.
From one line to federation
One bounded line can reveal feasibility in one context. Chapter 21 asks how several lines can share capability and policy without turning a central factory into the next constraint.
C20-S01: [C20-S01] HM Treasury, Test and Learn, 2026. C20-S02: [C20-S02] HM Treasury, The Green Book, 2026. C20-A01: [C20-A01] Eldridge et al., “CONSORT 2010 statement: extension to randomised pilot and feasibility trials,” BMJ 355:i5239, 2016. C20-R01: [C20-R01] U.S. GAO, Agile Assessment Guide, GAO-24-105506, 2023. C20-R02: [C20-R02] GOV.UK Service Manual, “How the beta phase works.” C20-R03: [C20-R03] U.S. GAO, Electronic Health Record Modernization: VA Is Making Incremental Improvements, but Much More Remains to Be Done, GAO-25-108091, 2025. C20-R04: [C20-R04] U.S. GAO, VA Electronic Health Record Modernization: Critical Actions Needed to Support Accelerated System Deployments, GAO-26-108812, 2025.