Part IV — Compounding Capability
Factory Intelligence and the Digital Twin
How can production evidence support explanation, prediction, experimentation, and learning?
Evidence becomes useful when it changes a named decision
Digital Workers, people, Work Centers, Quality Gates, releases, and operated systems all produce evidence. More evidence does not automatically produce more understanding. A trace can show that two events occurred close together without showing why. A graph can connect a decision to an artifact while preserving an incorrect assertion. A simulation can calculate a precise answer for a model that is unfit for the decision.
The central question is:
How can production evidence support explanation, prediction, experimentation, and learning without being mistaken for causal knowledge?
The bounded answer is:
A Software Manufacturing system needs governed evidence connecting events, provenance, structure, decisions, interventions, and outcomes; no single analytical method supplies that understanding alone.
Factory Intelligence is the decision capability created when governed evidence can be interpreted, challenged, used, and revised for a named production question. It is not a dashboard, data lake, model, graph, or product category. A Digital Twin can be one decision-specific representation inside that capability. It is neither mandatory nor inherently superior to a simpler record, process view, observability query, causal study, or simulation.
This chapter separates six complementary evidence domains: process mining, observability, provenance and knowledge representation, causal inference, simulation, and data governance. It then connects them through a decision-specific architecture. The architecture is author synthesis. Combining the domains does not turn correlation into causation, eliminate missing data, guarantee decision quality, or establish a universal platform.C12-F01C12-F02C12-F03C12-F04C12-F05C12-F06
C12.1 — Begin with the consequential question
Evidence architecture should start with a decision, not a collection technology. “Collect all engineering data” has no stopping condition and no defensible completeness test. “Should this rollout proceed for the declared service cohort?” identifies an owner, alternatives, urgency, affected system, evidence tolerance, and observable outcome.
A decision specification should name:
- the question and accountable owner;
- the system, population, cohort, or time boundary;
- the available alternatives, including no action;
- the consequence of being wrong or late;
- the evidence and uncertainty required;
- who may access, challenge, approve, appeal, or stop;
- the intervention or action that follows; and
- the outcome and revision trigger that close the learning loop.
The specification prevents evidence availability from becoming decision quality. A richly instrumented system may still lack the measure that distinguishes the alternatives. A sparse but purpose-built observation may be sufficient for a bounded choice. The owner is accountable for using evidence; they are not assumed omniscient.
The NASA Software Engineering Laboratory provides a historical boundary. It connected production questions, project evidence, analysis, owners, feedback, standards, tools, and training across one flight-dynamics environment. The participant record describes sustained study and improvement work, but also incomplete data, months-long feedback, collection burden, changing organizational conditions, and eventual decline. It demonstrates that an evidence capability is operated and maintained, not merely installed. It does not establish a universal outcome magnitude, modern platform design, or causal effect.
C12.2 — Keep six evidence domains distinct
The six domains overlap in their inputs but answer different questions. Their non-duplication matters because an attractive representation can otherwise imply more knowledge than the method supplies.
Process mining: what path was recorded?
Process mining discovers observed patterns, checks conformance, and supports enhancement from event records. In software production it can expose variants, loops, waiting, rework, and deviations—if cases, activities, timestamps, identities, and semantics are usable.
The mined path is not automatically the intended process, the complete process, or its cause. Missing events, inconsistent identifiers, changed tools, parallel work, and selective logging shape the result. A path frequency does not establish worker value, customer outcome, or why the path occurred. Governance and provenance must describe the event source; causal methods are needed for effect claims.C12-F01
Observability: what is the running system doing?
Traces, metrics, logs, and baggage provide different runtime signals. They can expose state, request paths, resource behavior, errors, saturation, or correlated changes. Sampling, cardinality, schema, clock semantics, instrumentation gaps, and collection cost remain visible constraints.
Telemetry describes recorded behavior. It does not establish the full production state, user value, or the cause of an outcome. A correlated signal can be operationally useful for diagnosis or stop decisions while remaining insufficient for causal attribution.C12-F02
Provenance and knowledge representation: how are records related?
Provenance models can represent entities, activities, agents, derivation, attribution, and influence. Knowledge graphs can connect Work Orders, decisions, artifacts, releases, services, controls, evidence, owners, and outcomes across heterogeneous systems.
A graph edge is an assertion. It does not prove that the assertion is true, current, complete, lawful, or causal. Useful representation therefore preserves source, time, confidence, coverage, access, and transformation history. It should distinguish recorded, derived, inferred, disputed, missing, and retired relationships.C12-F03
Causal inference: did an intervention change an outcome?
Causal inference begins with an intervention, outcome, population, time, alternatives, assumptions, and counterfactual question. Randomization, natural experiments, quasi-experimental designs, causal models, sensitivity analysis, and triangulation can strengthen identification under different conditions.
Temporal order and correlation are inputs, not conclusions. An intervention may coincide with a recovery because task mix, staffing, traffic, system state, or another change also moved. Causal claims must state the design, assumptions, competing explanations, uncertainty, and transfer boundary.C12-F04
Simulation: what might happen under a modeled condition?
Simulation explores possible behavior under declared assumptions and inputs. It can compare capacity conditions, policy changes, failure propagation, queue behavior, or operating strategies without immediately changing the real system.
Credibility is specific to intended use. Verification asks whether the model was implemented correctly; validation asks whether its relationship to relevant reality is adequate for the decision; uncertainty quantification and use history bound interpretation. A validated model is not reality and does not remain valid outside its acceptance domain.C12-F05C12-R02
Data governance: may this evidence be trusted and used?
Data governance assigns stewardship and lifecycle controls for provenance, quality, access, security, privacy, retention, preservation, and disposition. It makes data ownership and fitness challengeable.
Governance does not guarantee correctness, representativeness, lawful implementation, or good judgment. Available data is not automatically fit for a new purpose. A compliant retention decision can still preserve biased or incomplete evidence. Governance protects interpretation by exposing ownership, conditions, limits, and recourse.C12-F06
The practical map is simple:
| Question | Primary domain | What it cannot establish alone |
|---|---|---|
| What work path was recorded? | Process mining | Intent, completeness, cause, value |
| What is the running system doing? | Observability | Causal explanation or total state |
| How are evidence and decisions related? | Provenance/graph | Truth, freshness, causality |
| Did an intervention change an outcome? | Causal inference | Transfer beyond its design |
| What might happen under an option? | Simulation | Empirical outcome truth |
| Can evidence be retained and used? | Data governance | Substantive correctness |
C12.3 — Build a traceable evidence-to-decision chain
A decision-specific evidence architecture connects nine stages:
events
→ provenance and access
→ quality and missingness
→ method-specific views
→ validation and uncertainty
→ named decision and owner
→ intervention or no action
→ observed outcome
→ revision, suspension, or retirement
Events enter with source, time, scope, and collection conditions. Provenance shows their origin and transformations. Access controls protect sensitive evidence without pretending that restricted evidence does not exist. Quality checks expose completeness, freshness, semantic drift, sampling, and missingness.
Method-specific views then answer the question appropriate to their domain. Process views reconstruct recorded flow. Observability views examine runtime behavior. Graphs connect evidence. Causal designs test effects. Simulations explore modeled options. None is promoted into a universal analytical engine.
Validation records whether the selected evidence and method are fit for this decision, including uncertainty and challenge. The named owner decides among action, no action, request for more evidence, controlled experiment, escalation, or stop. The resulting intervention and observed outcome feed back into the evidence, model, policy, or decision boundary.
No-action is a real decision, not a gap. So are “insufficient evidence,” “narrow the question,” and “retire the representation.” A decision architecture that records only actions and successes cannot learn from restraint, false alarms, abandoned models, or failed assumptions.
Figure F12.1 production specification: Decision-specific evidence architecture
Figure F12.1 — Governed events pass through provenance, access, quality, missingness, method-specific views, validation and uncertainty into a named decision. Intervention or no-action produces observed outcomes that can revise or retire the evidence system. The six domains remain complementary rather than becoming one causal engine.
Azure Gandalf: a bounded evidence-to-operational-decision case
Microsoft Azure’s Gandalf system provides the principal production case. The NSDI 2020 paper is peer reviewed and organization-authored. It describes a service operating for more than eighteen months across Azure data- and control-plane rollouts.C12-A05
The disclosed chain is:
runtime signals
→ anomaly, correlation, and impact assessment
→ proceed or stop decision
→ automated stop, ticket, and owner intervention
→ reported model and caught-failure outcomes
Gandalf used logs, counters, and process events with anomaly, correlation, and impact models to support rollout decisions. The system could stop a rollout and notify or ticket the owning engineering team. The paper reports precision and recall under its definitions and 155 critical data-plane failures caught during an eight-month outcome window.
These records establish bounded model performance and operational use. They show that telemetry and analysis can support a named proceed-or-stop decision, intervention, owner notification, and measured outcome in one production context.
They do not independently prove the causal effect of Gandalf on Azure safety. The system’s creators authored the paper and data. False alarms and data/model limits remain. The public record does not provide complete lifecycle cost, privacy and employee-monitoring outcomes, decision-revision history, retirement evidence, or independent organizational corroboration. Gandalf does not implement or validate all six evidence domains and is not evidence that a Digital Twin is required.
GitHub’s first-party eBPF circular-dependency report offers a current supporting contrast: dependency observations informed a staged intervention and reported stability and recovery observations. It lacks a population, comparator, outcome distribution, full cost, privacy analysis, and independent review. It remains a separate bounded case, not a merged cohort.C12-R04
Evidence status — bounded. Gandalf supports a production telemetry-to-decision-to-intervention chain with disclosed model and caught-failure outcomes. Independent causal validation, complete economics, privacy outcomes, revision history, and transfer evidence remain collection priorities for future editions.
C12.4 — Choose representation fidelity for the decision
“Digital Twin” is used across many fields with heterogeneous definitions. In this chapter it means a governed representation of the minimum structure, state, behavior, and constraints sufficient for a named decision. The word does not certify synchronization, completeness, prediction, causality, or value.
Decision fidelity asks: What is the least costly representation that is credible for this purpose?
Level 0 — Do not build
Do not build when the decision, owner, observable boundary, validation method, lawful evidence use, or lifecycle value cannot be established. A direct query, static record, controlled experiment, or human inquiry may be more appropriate.
Level 1 — Descriptive record or static view
Use a bounded record to locate, reconstruct, or explain state. It needs provenance, scope, update date, and known omissions. Stop when it cannot answer the named question.
Level 2 — Synchronized descriptive model
Use a refreshed representation to monitor a bounded process or detect divergence. It needs a defined update mechanism, latency, completeness, and correspondence tests. Suspend it when freshness or fidelity cannot be validated.
Level 3 — Predictive or comparative model
Use a model to forecast a stated outcome or compare options. It needs holdout or backtesting, calibration or error evidence, uncertainty, drift monitoring, and a decision protocol. Narrow or stop when error exceeds the decision’s tolerance.
Level 4 — Experimental or prescriptive model
Use a model to recommend or test bounded interventions only with explicit causal assumptions, simulation or experiment validation, reversibility, independent challenge, and accountable authorization. High consequence or uncertainty may require controlled real-world learning instead.
These are decision-fidelity levels, not maturity stages. An organization does not improve by moving upward. A Level 1 record may be exactly right for an audit question. A Level 4 model may be unjustified, unsafe, or too expensive. Outcomes can move the representation down, up, into suspension, or into retirement.C12-A01C12-A02C12-A03C12-A04C12-R01
Figure F12.2 production specification: Decision-fidelity framework
Figure F12.2 — Representation fidelity is selected for a named decision and accepted validation boundary. “Do not build” is mandatory; higher levels do not imply maturity, improved outcomes, or entitlement to automate decisions.
Govern evidence without turning workers into instrumentation
Production evidence can expose employee activity, communication, location, identity, performance proxies, or protected information. A technically available signal is not automatically appropriate for individual evaluation, surveillance, or a new purpose.
The architecture must record purpose, lawful and ethical basis, affected parties, access, retention, contestability, and disposal. Separate operational diagnosis from worker-performance judgment. Prefer aggregate or minimized evidence where it answers the decision. Preserve appeal and correction where evidence can affect people. Security controls protect access; they do not resolve power, fairness, purpose, or consent.
Neither Gandalf nor the other bounded cases establishes privacy or workforce outcomes. This chapter treats privacy and employee monitoring as governance requirements and unresolved evidence needs, not case-derived benefits.C12-F06
The decision-evidence test
Leaders can test a proposed evidence capability with nine questions.
- Which decision will change? Name owner, alternatives, scope, urgency, and affected parties.
- Which domain answers each question? Keep process, runtime, provenance, causal, simulation, and governance claims separate.
- What is recorded, derived, inferred, sampled, missing, or disputed? Do not hide evidence states.
- What validation makes the method credible for this use? State assumptions, error, uncertainty, and challenge.
- Who may access, decide, intervene, appeal, or stop? Make authority and recourse explicit.
- What action or no-action follows? Prevent analysis from becoming an unowned report.
- Which outcome will be observed? Avoid replacing decision outcomes with data-volume or dashboard-use counts.
- What will cause revision, suspension, or retirement? Evidence capabilities have lifecycles.
- Would a simpler representation answer the question? “Do not build” and “build less” are valid outcomes.
This is a decision framework, not a universal architecture, data completeness threshold, real-time mandate, cost optimum, Digital Twin certification, or guarantee of good judgment.
What to remember
Factory Intelligence begins with a consequential decision and a governed evidence chain, not with data accumulation.
Process mining, observability, provenance and knowledge graphs, causal inference, simulation, and data governance answer complementary questions. None substitutes for the others.
Observation is not causality. Provenance is not truth. Correlation is not causal identification. Simulation is not reality. Governance is not correctness. Evidence availability is not decision quality.
Azure Gandalf demonstrates a bounded production chain from telemetry and analysis to a stop-or-proceed decision, intervention, notification, and reported outcomes. It does not validate a universal evidence architecture, Digital Twin, portable economics, privacy outcome, or independent causal effect.
Representation fidelity belongs to a named decision. “Do not build,” “build less,” suspend, and retire are legitimate choices.
Build the smallest governed evidence capability that can support and revise a named decision.
From intelligence to quality at the source
Governed evidence matters only when the production system uses it. Chapter 13 turns from evidence architecture to Quality at the Source: placing prevention, testability, secure and operable defaults, accessibility, fast feedback, and ownership near the point where work is created while preserving independent challenge.
C12-F01: [C12-F01] IEEE Task Force on Process Mining, Process Mining Manifesto. C12-F02: [C12-F02] OpenTelemetry, signals and specification documentation. C12-F03: [C12-F03] W3C, PROV-O: The PROV Ontology. C12-F04: [C12-F04] National Academies, Advancing the Framework for Assessing Causality of Health and Welfare Effects. C12-F05: [C12-F05] NASA-STD-7009, Standard for Models and Simulations. C12-F06: [C12-F06] NIST SP 1500-18r2, research data framework and data-governance guidance. C12-A01: [C12-A01] Guinea-Cabrera and Holgado-Terriza, “Digital Twins in Software Engineering—A Systematic Literature Review and Vision,” 2024. C12-A02: [C12-A02] Dalibor et al., “A Cross-Domain Systematic Mapping Study on Software Engineering for Digital Twins,” 2022. C12-A03: [C12-A03] Kimmel et al., “Digital Twins for Software Engineering Processes,” ICSE NIER 2025. C12-A04: [C12-A04] Samoud, Aissat, and Bordeleau, “A Model-Driven Digital Twin for the Systematic Improvement of DevOps Pipelines,” 2026 preprint. C12-A05: [C12-A05] Li et al., “Gandalf: An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud Infrastructure,” USENIX NSDI 2020. C12-R01: [C12-R01] NIST IR 8356, Security and Trust Considerations for Digital Twin Technology, 2025. C12-R02: [C12-R02] Shao, Hightower, and Schindel, “Credibility Consideration for Digital Twins in Manufacturing,” 2023. C12-R04: [C12-R04] GitHub, “How GitHub uses eBPF to improve deployment safety,” first-party engineering report. : [A28] Basili, Caldiera, and Rombach, “The Software Engineering Laboratory: An Operational Software Experience Factory,” ICSE 1992. : [A29] Basili et al., “Lessons Learned from 25 Years of Process Improvement: The Rise and Fall of the NASA Software Engineering Laboratory,” ICSE 2002. : [S16] ISO/IEC/IEEE 24748-7000:2022, addressing ethical concerns during system design. : [S19] NIST SP 800-160 Vol. 1 Rev. 1, Engineering Trustworthy Secure Systems, 2022.