Part II — Defining the Production System
The Software Factory as a Socio-Technical System
What is the minimum complete model of a Software Factory?
When architecture became accountability
Chapter 4 defined Software Manufacturing as a discipline and named its object: the whole production system through which intent becomes trusted software capability and evidence becomes learning. That definition creates an immediate practical problem. A system cannot be governed merely by declaring that it exists. Leaders need to know what belongs inside its boundary, what remains outside, how the parts relate, and what evidence would show that the arrangement is working.
The danger is to answer too quickly. In software, the word factory often points to the most visible technical machinery: source repositories, build services, deployment pipelines, cloud platforms, test automation, and operational tooling. Those capabilities matter. Yet a technically sophisticated path to production can coexist with weak ownership, unsafe work, fragmented demand, invisible supplier dependencies, poor recovery, or evidence that never changes a decision. The machinery can function while the production system fails.
DBS offered a public example of that distinction in 2023. The bank described a long movement from a monolithic mainframe toward a cloud-native, microservices-based architecture. Management associated that shift with faster execution, but also with a more complex infrastructure requiring greater operational rigor and oversight. During a year of repeated digital-service disruptions, the technical architecture could not be separated from change management, resilience, incident management, technology-risk governance, third-party systems, customer journeys, and board accountability.
The response crossed those boundaries. DBS reported creating a near-production test environment for critical services, strengthening monitoring and diagnostic capability, revising operational playbooks, clarifying disruption ownership with service providers, increasing technology-risk oversight, and forming a board risk subcommittee focused on technology. The Monetary Authority of Singapore imposed a six-month pause on nonessential activities and later decided not to extend it after reporting substantive progress.
This record does not establish one cause for the disruptions, nor does it validate the remedies. DBS management is an interested narrator, and the independent technical review was not published in full. The case is useful for a narrower reason. Once software became operationally consequential, no single technical component or department described the object under scrutiny. Architecture, people, suppliers, environments, controls, evidence, customer outcomes, and external authority had become one management problem.
Calling the deployment platform “the factory” would therefore have hidden much of what mattered. Calling the technology organization “the factory” would have been too broad in some respects and too narrow in others. The relevant boundary ran through business journeys, engineering capabilities, third-party relationships, operational command, risk decisions, and evidence of service behavior. It was defined by responsibility for a recurring outcome, not by a reporting line.
That is the question of this chapter:
What is the minimum complete model of a Software Factory?
A Software Factory is the bounded socio-technical production system that puts Software Manufacturing into operation by connecting people, Production Lines, Work Centers, Work Rooms, Factory Assets, Digital Workers, controls, and evidence.
Its minimum model has seven functions: outcome flow, reusable capability, governed context, reusable assets, accountable actors, controls, and evidence. The functions are author synthesis. Research and standards support their constituent roles and the risks of omission; they do not prove that seven is the only legitimate count or that these names are universal. Another model may combine the parts differently. It is complete only if it preserves the relationships each function makes visible.
The central reframe is simple:
The Software Factory is not its machinery. It is the governed relationship between demand, capability, context, actors, controls, assets, and evidence.
Draw the boundary around the outcome
Every system description begins by choosing a system of interest. The choice is consequential because a boundary determines which relationships appear internal, which appear as interfaces, and which disappear from attention. A boundary that follows the organization chart may be administratively convenient while cutting through the actual flow of software production. A boundary that includes the entire enterprise may be so broad that it explains nothing.
ISO/IEC/IEEE 42010 provides useful discipline without prescribing the answer. It distinguishes an entity of interest from the description of its architecture and requires attention to stakeholders, concerns, viewpoints, and correspondences. Its scope can include products, service lines, enterprises, product families, and systems of systems. That breadth matters here: there is no reason to assume that a Software Factory must match a legal entity, a technical platform, or a department.
ISO/IEC/IEEE 15288 adds lifecycle discipline at the system level. It makes acquisition, supply, development, operation, support, and retirement visible across technical, management, agreement, and organizational processes. It does not define a Software Factory, and a standard does not prove performance. It does support the claim that a governable production boundary must survive beyond development and must account for relationships with acquirers, suppliers, operators, maintainers, and affected stakeholders.
The right starting question is therefore not “Which tools do we own?” or “Who reports to the chief technology officer?” It is: For what recurring class of demand are we accountable for an operated software capability and its evidence? The answer may be a regulated payments journey, a spacecraft flight-software family, a national identity capability, an internal risk service, or a connected industrial product. The boundary should include the relationships that materially condition that outcome.
This does not pull the entire institution inside. Corporate strategy, general finance, human resources, procurement, legal work, facilities, and physical operations may all influence production. They enter the factory boundary only where their decisions directly constrain, authorize, supply, operate, or learn from the recurring software outcome. Otherwise they remain external stakeholders connected through explicit interfaces.
An interface is not a reason to ignore the outside party. Regulators can impose conditions without participating in daily production. Suppliers may operate critical capabilities without belonging to the accountable institution. Users produce operational evidence without sitting inside an engineering organization. Open-source communities can shape foundational assets while retaining their own authority and incentives. The boundary must show these relationships without pretending the factory owner controls every actor.
The same principle applies inward. A single team may perform several functions. A small institution may combine architecture, engineering, testing, release, and operation in a handful of people. A large enterprise may distribute the same functions across business units and suppliers. Completeness describes necessary relationships, not organizational bulk. Seven functions must never become a prescription for seven departments, seven platforms, or seven process stages.
This distinction prevents two opposite errors. The first is technical reduction: defining the factory as automation and infrastructure while treating people, authority, risk, and learning as external complications. The second is organizational inflation: defining almost every enterprise activity as part of the factory until no boundary can guide a decision. Draw the boundary around the recurring outcome, then name every material interface that crosses it.
Why the system must be socio-technical
The socio-technical premise is older than modern software. Trist and Bamforth's study of changing coal-mining work showed that technical redesign and social organization interacted. The relevance is not a direct analogy between mining and software. The industries, hazards, labor arrangements, and technologies differ radically. The transferable insight is that changing machinery changes coordination, discretion, skill, identity, and responsibility; the technical subsystem cannot be optimized independently from the organization of work.
Baxter and Sommerville later brought this problem closer to software and systems engineering. They argued that organizational change and technical system development are often handled through separate traditions, even though large systems depend on their interaction. They also recorded a practical obstacle: socio-technical approaches struggle when they do not fit procurement, project management, and established engineering processes. “Joint design” can become an aspiration with no place in the decisions that actually shape work.
A Software Factory must therefore describe both the means of production and the conditions under which people and software-based actors use them. A tool changes who can act, what they can see, which errors become possible, how quickly feedback arrives, and where authority moves. A policy changes the viable technical routes. A supplier contract changes access to evidence and the time needed to correct failure. A test environment changes what can be learned before exposure. A staffing decision changes which controls are real and which exist only on paper.
NIST SP 800-160 makes a parallel point through trustworthy systems engineering. Trustworthiness cannot be attached after construction as a single control. It must be engineered across lifecycle activities, system elements, stakeholder needs, requirements, architecture, interfaces, verification, validation, operation, and assurance. The publication is security-centered and does not define this model, but it reinforces the need to connect technical structure, responsible actors, controls, and evidence through time.
The result is not a call for consensus on every decision. Socio-technical does not mean that all stakeholders govern all work, or that expertise and authority disappear into workshops. Consequential systems require clear ownership, independent challenge, and sometimes nonnegotiable constraints. The point is to make the allocation of authority part of the system design rather than an assumption hidden behind technical components.
Nor does it mean that every inconvenience is a system issue. Local defects still exist. Skills still matter. A poorly designed module can fail because it is poorly designed. A team can lack needed expertise. The production-system perspective becomes necessary when the outcome depends on relationships across boundaries: when local competence cannot resolve a queue, when a policy conflicts with the available environment, when an asset has no viable owner, or when operational evidence cannot reach the people who set demand.
This is why the factory is not equivalent to a platform. A platform can provide exceptional reusable capability and a coherent working environment. It may carry several functions in the reference model. But unless it also owns or connects the outcome flow, accountable actors, controls, external interfaces, and evidence loop, it remains one powerful part of the factory. Where an organization already uses platform for that complete object, the substantive test matters more than the label.
Seven functions, one production system
The reference model has seven functions because each answers a different question about completeness. Removing any one leaves a predictable blind spot. The functions are not arranged as sequential phases. They overlap, recur, and change one another.
1. Flow: what outcome is moving?
A Production Line is the end-to-end, outcome-oriented flow through which a defined class of demand becomes operated software capability and evidence. It gives the factory a reason to exist. Without flow, the model becomes an inventory of teams and tools with no way to judge whether their relationships produce an outcome.
The Production Line begins with bounded demand, not necessarily a settled specification. It may include discovery, parallel work, feedback, stop decisions, and rework. It may route different risk classes differently. Its endpoint is not code completion or deployment but operated capability plus evidence of use, behavior, risk, cost, and work. Chapter 7 will define the Work Order that authorizes demand, and Chapter 9 will examine flow under variability. Here the essential point is only that the system needs an end-to-end path.
If flow is omitted, local activity becomes the proxy for progress. A build service can report success while authorization waits elsewhere. A testing group can maximize utilization while high-risk work accumulates. A supplier can meet a contractual milestone while the institution remains unable to operate the capability. The Production Line makes those delays and disconnections part of one object.
2. Capability: what coherent work can be reused?
A Work Center is a reusable production capability serving one or more Production Lines. It performs a coherent class of work through defined interfaces, policy, skills, assets, ownership, and service expectations. Architecture enablement, verification, identity integration, release assurance, or observability engineering may qualify, but the name of a team does not.
Modularity research provides a useful constraint. Parnas argued that decomposition should hide design decisions likely to change, rather than merely mirror a processing sequence. Later analysis of complex software designs showed that coupling structure affects how change propagates. These sources concern software modules, not organizational design, but their interface logic transfers cautiously: a reusable capability is valuable when consumers can depend on a stable responsibility without absorbing every internal decision.
Without capability, each Production Line must recreate scarce work or negotiate through informal relationships. With poorly designed capability, the Work Center becomes a central queue, hides policy, or transfers its complexity to consumers. Chapter 8 owns the topology choices. Chapter 5 only requires that reusable work be visible as capability with an interface and an owner.
3. Context: where can the work be done safely?
A Work Room is a governed working environment configured for a defined class of work. It contains the access, tools, data, policy, assets, isolation, and evidence capture needed to perform that work safely and effectively. It may be virtual or physical. It is not synonymous with an integrated development environment, repository, chat space, cloud account, or office.
Context matters because identical work performed under different access, data, hardware, regulatory, or operational conditions is not identical in risk. Flight software must run against constrained hardware and mission interfaces. A bank may need a near-production environment to expose interactions that smaller tests miss. A public identity implementation must connect local policy, languages, infrastructure, data governance, and external providers. The environment is part of the production design, not neutral background.
Without context, policy becomes detached from the place where work occurs. People improvise access, copy sensitive data, simulate the wrong conditions, or discover integration constraints late. An overbuilt Work Room can be equally harmful if it imposes cost and delay unrelated to risk. Completeness requires an explicit context; proportionality determines how elaborate it should be.
4. Assets: what prior capability can compound?
A Factory Asset is a governed and reusable resource whose lifecycle is managed to reduce future effort, risk, or variability or to increase quality and learning. Code, platforms, patterns, policies, test capabilities, data, models, documentation, and operational knowledge may qualify. Reusability alone is insufficient. Ownership, provenance, fitness, consumers, support, improvement, and retirement must be known.
Factory Assets differ from Work Centers. An asset is a reusable resource; a Work Center is a capability that performs a coherent class of work. A verification service may be a Work Center. Its test harness, policy library, evidence schema, and operating knowledge may be Factory Assets. The distinction matters because resources can outlive teams, move across capabilities, and accumulate hidden maintenance obligations.
Without assets, each Production Line pays repeatedly for knowledge already created. With unmanaged assets, reuse becomes dependency without accountability. Chapter 10 will examine those economics. In the reference model, assets are necessary because a factory that cannot preserve and responsibly reuse prior capability cannot compound.
5. Actors: who can act, and who remains accountable?
The factory includes people and Digital Workers. A Digital Worker is a bounded software-based actor that performs or coordinates production work under an assigned identity, explicit authority, observable behavior, evidence obligations, escalation conditions, and accountable human ownership. Deterministic automation may qualify when it meets those conditions; an unowned script does not.
Human actors include demand owners, engineers, operators, security and assurance specialists, users, managers, supplier personnel, and other stakeholders whose decisions materially shape production. They do not all sit inside the same hierarchy. Their roles must nevertheless make authority, obligation, and recourse visible at the interfaces where their actions affect the outcome.
Without actors, diagrams turn responsibility into arrows between boxes. A deployment “happens,” a risk is “accepted,” or an exception is “approved” without a named authority. Automation can deepen this ambiguity by executing more work while obscuring who assigned its permissions or who must intervene. Chapter 11 will address bounded delegation. Here the completeness rule is that every consequential action has an identity, authority boundary, evidence obligation, escalation path, and accountable owner.
6. Controls: what constrains, authorizes, or stops work?
Controls express obligations and risk decisions inside production. They can include policy, technical constraints, separation of duties, access rules, acceptance criteria, verification, monitoring, decision rights, and authorized exceptions. A control may be automated, human, or combined. It may allow, correct, stop, or require escalation.
Controls are not necessarily sequential gates. Treating them as a final approval layer recreates the separation between production and assurance. In a coherent factory, a control is located where relevant evidence exists and where correction is economically and operationally possible. Independent challenge remains necessary for consequential risk, but it should be connected to the Production Line rather than hidden in an external queue.
Without controls, speed can move risk without authority. With excessive or opaque controls, the factory creates delay, workarounds, and false assurance. Chapters 13 and 14 will examine quality and Quality Gates in depth. The Chapter 5 requirement is narrower: the system must show what constrains action, who owns the rule, what evidence is evaluated, and what happens when the rule cannot be satisfied.
7. Evidence: what returns from work and operation?
Evidence connects action to consequence. It includes what the factory knows about demand, work state, decisions, changes, assets, controls, operation, user outcomes, failure, cost, burden, and uncertainty. Evidence may be quantitative or qualitative. It must retain enough provenance and context to support a named decision.
Evidence is not a dashboard and does not guarantee learning. A metric can be ignored, misread, or gamed. An incident record can document harm without changing policy. A completed control can prove only that a specified check occurred. Learning requires responsible interpretation and the authority to change later demand, capability, assets, context, controls, or work.
Without evidence, the factory cannot distinguish intended operation from actual consequence. It cannot reconstruct decisions, challenge assumptions, or improve deliberately. Chapters 12, 15, and 16 will address intelligence, Traceability, and measurement. At this stage the completeness rule is simply that production and operation return evidence to the system that can act on it.
Relationships are the architecture
A list of seven functions is not yet a system. The architecture lies in the relationships among them.
Bounded demand enters a Production Line. Human and Digital Worker actors perform or coordinate work through Work Rooms, using and improving Factory Assets. The line invokes Work Centers through explicit interfaces rather than by assuming shared capability is freely available. Controls constrain, authorize, stop, or redirect action. Operated software capability produces evidence. That evidence returns to actors who can change demand, capability, assets, environments, controls, or the line itself.
Figure F05.1 — The Software Factory exists in the governed relationships among seven required functions. A bounded Production Line connects demand to operated capability and evidence. Work Centers provide reusable capability; Work Rooms provide governed context; Factory Assets preserve reusable resources; people and Digital Workers act under explicit authority; controls constrain and authorize; evidence returns from production and operation. Suppliers, regulators, users, open-source communities, and adjacent operations connect through named boundary interfaces.
Source: Author synthesis constrained by A05, A30, A02, A16, S02, S03, S19, B14, B17, R13/R14, R21/R22, and R23/R24. Structured specification: figures/F05.1-spec.md.
The figure should not be read as a fixed lifecycle. Work may loop, branch, wait, stop, or return to discovery. Nor is it an organization chart. A team may participate in several Work Centers; a supplier may operate a Work Room; a Factory Asset may be maintained by an open-source community; a control owner may sit outside the line. The model is concerned with responsibility and relationship, not reporting boxes.
This relationship view produces a practical completeness test. Ask seven questions of any proposed factory boundary:
- Flow: Which recurring demand becomes which operated capability and evidence?
- Capability: Which coherent classes of work are reused through explicit interfaces?
- Context: Where can each class of work be performed with the right access, data, tools, isolation, and evidence capture?
- Assets: Which reusable resources have known ownership, fitness, provenance, consumers, and lifecycle?
- Actors: Who can act, within what authority, and who remains accountable?
- Controls: What constrains, authorizes, stops, corrects, or permits exception?
- Evidence: What returns from work and operation, and who can change the system because of it?
If an answer is absent, the model contains a blind spot. If an answer merely names a tool or department, the relationship remains unresolved. If every answer is visible under another vocabulary, the proposed system may already be complete without adopting the book's labels.
Three contexts, one completeness test
The reference model must survive changes in scale, sourcing, technical architecture, and institutional purpose. Three public records test explanatory coverage without implying that any organization adopted Software Manufacturing.
Regulated banking: the boundary crosses the supplier contract
The DBS record shows a factory boundary that cannot stop at internal engineering. Management described microservices supported by third-party systems, customer journeys spanning applications and infrastructure, near-production testing, real-time monitoring, operational command, playbooks, change controls, service-provider ownership, technology-risk oversight, board governance, and regulator action.
Mapped to the reference model, customer journeys provide the outcome-flow lens. Engineering, testing, resilience, and incident capabilities behave as Work Center candidates. The near-production test environment and operating command context behave as Work Room candidates. Diagnostic capability, playbooks, patterns, and platform services can be treated as potential Factory Assets only where ownership and lifecycle are known. Employees, providers, board members, and regulators occupy different actor and authority positions. Change, risk, and regulatory conditions provide controls. Monitoring and disruption records provide evidence.
The case does not show that these elements formed one coherent factory. Public evidence cannot establish that. It shows why omitting any class would make the production problem harder to explain. A supplier contract did not move the dependency outside the system of interest; it created a boundary interface that required ownership, evidence, and recourse.
Embedded flight software: reusable core, mission-specific context
NASA's core Flight System provides a different test. NASA documents a mission-independent, platform-independent flight-software environment with a reusable core Flight Executive, operating-system and platform abstraction layers, applications, libraries, interfaces, ground tools, and services for software messaging, time, events, tables, files, commands, telemetry, and integrity checks. Earlier program material stated goals around formalized reuse, cross-organization collaboration, sustaining engineering, common standards and tools, and scale from small instruments to systems of systems.
This architecture makes capability, assets, and interfaces unusually visible. Reusable runtime services and abstractions can support multiple mission Production Lines. Mission applications provide context-specific behavior. Hardware, real-time operating systems, ground interfaces, simulators, and mission configurations shape Work Rooms. Engineers, integrators, operators, and automated services remain actors. Command paths, checksum functions, testing, mission assurance, and runtime constraints provide controls and evidence.
But cFS is not itself the complete Software Factory. Its public catalog does not describe every mission's demand decisions, workforce, assurance authority, supplier relationships, operational learning, or asset lifecycle. The case therefore prevents two errors at once: reusable technical architecture is essential capability, and even an excellent reference architecture remains a subsystem of the wider socio-technical production system.
Public digital infrastructure: ownership without total control
MOSIP tests a distributed public-purpose setting. The World Bank described it as a modular open-source approach to foundational identity, designed for country variation, integration with existing infrastructure, open standards, privacy and security, interchangeable modules, external biometric systems, and a community of commercial partners and developers. MOSIP's own documentation describes API-first lifecycle capabilities, configurability, country ownership, integration surfaces, and open contribution.
Here the boundary can cross a public institution, an open-source program, local implementers, technology partners, external service providers, registrars, operators, and the communities affected by identity policy. Modules and interfaces may be Factory Assets. Shared integration or verification capability may behave as Work Centers. Local data, infrastructure, language, access, and policy conditions shape Work Rooms. Public law, privacy, security, inclusion duties, and assurance create controls. Registration, authentication, operation, grievances, and failure produce evidence.
The records do not establish equitable outcomes, effective national implementation, or lower cost. They establish a more limited architectural lesson: open source and modularity distribute ownership; they do not eliminate it. A country may own the public outcome without controlling every contributor or supplier. A governable factory boundary must therefore specify which interfaces it can change, which assets it adopts, what evidence it can obtain, and where recourse lies when an external dependency fails.
Across the three cases, the organizational forms differ sharply. One has a bank and regulator, one has flight software across mission and hardware constraints, and one combines public authority with open-source and partner ecosystems. The seven functions remain intelligible because they describe responsibilities, not a template. Their implementation should differ because the risks, institutions, resources, and operating contexts differ.
What executives should be able to see
Executives do not need to design every Work Center or inspect every interface. They do need to know whether the factory boundary supports accountable decisions.
First, they should be able to name the outcome boundary. “Technology,” “digital,” and “the platform” are usually too vague. A useful boundary identifies a recurring class of demand, an operated capability, the affected stakeholders, and the evidence that returns.
Second, they should see dependency without converting every dependency into central ownership. Which capabilities are shared? Which environments are controlled locally? Which assets come from suppliers or open-source communities? Where can the institution set policy, and where can it only choose whether and how to depend?
Third, they should see authority. Every consequential action should have a named actor, an authority limit, an evidence obligation, and an escalation or recourse path. This is especially important for outsourced work and automation. Delegation can change who performs work; it cannot make accountability disappear.
Fourth, they should distinguish completeness from complexity. A small factory can satisfy all seven functions with modest mechanisms. A large enterprise can own many platforms and teams while leaving several functions incoherent. The model does not reward component count. It asks whether necessary relationships are visible and governable.
Fifth, they should test whether evidence can change the system. If operational data cannot alter priorities, if incident findings cannot change assets or controls, or if supplier evidence cannot be obtained, the return loop is broken. Reporting volume does not repair it.
Finally, leaders should know when the model is unnecessary. A bounded technical improvement may be managed fully through software engineering, SRE, platform engineering, security, or another established discipline. The Software Factory perspective earns attention only when relationships across those capabilities determine the outcome.
What to remember
A Software Factory is a bounded socio-technical production system, not a building, department, platform, pipeline, sourcing model, or collection of tools.
Its minimum complete model makes seven functions visible: outcome flow, reusable capability, governed context, reusable assets, accountable actors, controls, and evidence.
The functions become a system through their relationships. Production Lines invoke Work Centers; actors work through Work Rooms using Factory Assets; controls constrain and authorize; evidence returns from work and operation; external parties connect through explicit interfaces.
Completeness does not prescribe scale or structure. One actor may perform several functions, and one function may span teams, suppliers, communities, or systems. The test is whether responsibility, authority, interfaces, and evidence remain visible.
The model is author synthesis. The cases show that it remains explanatory across regulated, embedded, open-source, public-purpose, and supplier-dependent settings. They do not prove adoption, causality, or performance.
The Software Factory is not its machinery. It is the governed relationship between demand, capability, context, actors, controls, assets, and evidence.
From complete system to coordinated action
The reference model answers what must be connected. It does not explain how those connections operate from one decision to the next.
Production Lines, Work Centers, Work Rooms, Factory Assets, actors, controls, and evidence do not coordinate themselves. Policies must be made executable. Decision rights must remain clear across boundaries. Work state must move without becoming a centralized command queue. Exceptions must be governed. Evidence must reach people and systems able to act. Learning must change the conditions of later work.
That creates the next question, but does not authorize its answer here: What coordinating layer allows autonomous actors and technical systems to operate as one factory without centralizing every decision? Chapter 6 will define the Factory Operating System only after its own evidence gate is passed.