SMDigital Book
Chapter 02The system that was almost finished
© Bhumaha Solutions Private LimitedAuthor: B. Thirumoorthy
02

Part I · Reframing the Production Problem

The Fragmented Delivery System

Why do more people, tools, and activity so often fail to create faster, more reliable outcomes?

23 minute read5,107 wordsPublished · Edition 1.0

The system that was almost finished

Figure 2.1 — Trilogy produced functioning components without delivering the intended case-management capability. The public record identifies multiple contributing conditions; the visual does not assign a single cause.

By the beginning of 2005, the Federal Bureau of Investigation had spent years modernizing the technology used to support investigations. The program was called Trilogy. Its name reflected its structure: one component would replace network infrastructure, another would provide new computers and common office software, and a third would modernize the applications through which agents managed and shared case information.

The first two components eventually operated. Networks had been upgraded. New hardware and software had been installed. The work was visible across FBI offices. Yet the component that was supposed to transform investigative case management—the Virtual Case File, or VCF—did not become operational.

The modernization had not lacked activity. After the attacks of 11 September 2001, the FBI's mission and information needs changed sharply. Congress provided an additional $78 million to accelerate the infrastructure work. Requirements expanded. Contractors developed software. Managers revised commitments. Teams installed equipment across hundreds of locations. By March 2004, however, the three Trilogy components had collectively accumulated about $120 million in cost overruns and at least 21 months of schedule delay against the commitments made in January 2002.

The acceleration produced a revealing result. The infrastructure deployment that had originally been due in May 2004 was pulled forward to July 2002. After repeated delays, it was completed in April 2004—one month before the original date, but almost two years after the accelerated target. The extra effort delivered functioning infrastructure, but it did not deliver the end-to-end investigative capability for which the modernization existed.

The application work had been moving beneath an unstable surface. The FBI's requirements evolved slowly and remained incomplete. The program lacked a firm design baseline. Responsibilities crossed FBI units, management positions, contractors, and oversight bodies. The Department of Justice's inspector general identified weaknesses in requirements, architecture, investment management, management continuity, and contractor oversight. The mission change was real, the technology was difficult, and some needed skills were scarce. No single defect explains what happened.

In March 2005, the FBI terminated the three-year, $170 million VCF effort and began again with a replacement program called Sentinel.

The tempting explanation is that a software project failed. That is true, but incomplete. Trilogy also presents a harder problem: how can an organization contain competent people, substantial funding, functioning technology, extensive management attention, and completed components—and still fail to create the capability that justified the work?

The answer cannot be found by looking only at the effort of any individual team. It lies in the delivery system formed by all of them.

The central question of this chapter is:

Why do more people, tools, and activity so often fail to create faster, more reliable outcomes?

The bounded answer is that fragmented demand, queues, handoffs, overload, dependencies, rework, and conflicting incentives can make delay an emergent property of the system. Local actors may work hard and local components may improve while the end-to-end outcome remains slow and unpredictable.

Fragmentation is the condition in which every part can report progress while the whole loses the ability to move.

This is not a claim that every delay is systemic. Sometimes an organization genuinely lacks a critical skill, enough capacity at a constraint, adequate technology, or an external decision it cannot control. Nor is it an argument against specialists, independent review, suppliers, or distributed teams. Each can be necessary. The claim is narrower: before leaders conclude that disappointing delivery proves insufficient individual effort, they must examine the conditions through which that effort becomes—or fails to become—usable capability.

Busy is not the same as moving

Organizations usually see software work through administrative windows. A portfolio office sees projects. A finance function sees budgets and cost centers. Product leaders see roadmaps. Engineering managers see teams and backlogs. Risk functions see approvals and exceptions. Suppliers see contracts and deliverables. Each view is legitimate, but none necessarily reveals the complete path from an institutional need to an operating result.

This fragmentation becomes consequential when every part is managed as though its local output were the outcome. Analysis produces a specification. A team produces code. Testing produces findings. Security produces an assessment. Operations accepts a release. A supplier completes a milestone. Each group can meet its target while the work between them waits, changes, returns, or loses priority.

The organization then experiences a paradox: activity rises while confidence falls.

More work is authorized to demonstrate urgency. More specialists are added to address risk. More status meetings are created to coordinate the specialists. More tools are purchased to make the work visible. More exceptions are issued to recover dates endangered by the additional work. None of these actions is irrational in isolation. Together, they can create a system in which people spend increasing effort managing the consequences of work already in motion.

Executives usually encounter the result as a forecast problem. Dates move. Estimates widen. Teams report being fully occupied, yet important outcomes remain stuck. Individual leaders can explain every delay: a dependency, a late decision, an environment, an approval, a production incident, a missing expert, a supplier, a change in scope. Each explanation may be accurate. The mistake is to treat the explanations as unrelated events.

When the same kinds of delay recur across work, their recurrence is evidence. A dependency that blocks one initiative may be incidental. A shared service that blocks eight initiatives is part of the production system. A late approval may reflect an exceptional case. A control function with a permanent queue is a designed capacity constraint, whether or not anyone intended it to be one. A priority change may be necessary. Continual reprioritization is a demand condition that changes the economics of every item already under way.

The distinction matters because the response follows the diagnosis. If delay is attributed to weak effort, leaders demand more effort. If it is attributed to a poor tool, they buy another tool. If it is attributed to a team, they reorganize the team. These interventions can improve a real local problem. But when the outcome emerges from relationships among demand, capacity, dependencies, controls, and feedback, local improvement can leave the governing mechanism untouched.

The delivery system is not the organization chart. It is the set of paths through which work must actually travel, including the paths by which decisions, information, evidence, and corrections return. It includes formal process and informal rescue. It includes the work visible in plans and the unplanned work created by incidents, clarification, rework, and coordination. It includes the people doing the work and the conditions that make their work depend on others.

As illustrated in Figure 2.2, that system is often invisible precisely because each participant sees only a credible portion of it.

Figure 2.2 — Administrative views can each be accurate while the connected path from need to operating result remains obscured. Visibility loss is a system condition, not proof that any one function is unnecessary.

What the evidence establishes

The system explanation rests on several evidence paths. Queueing research relates inventory, completion, time, variability, and capacity. Software-engineering studies show how dependencies become coordination demands. Contemporary surveys add evidence about unstable priorities, while public audits show how the mechanisms can combine in a consequential program. None is sufficient alone.

Work accumulates in time, not only in lists

In 1961, John Little proved the queueing relationship now known as Little's Law. In a stable long-run system, the average number of items in the system equals the average rate at which items arrive or complete multiplied by the average time each item spends there: L = λW.

The formula imposes discipline on management language. Work in progress, throughput, and elapsed time are not independent. If a stable system completes ten comparable items a month while carrying twenty, the corresponding average time in the system is two months. If inventory rises while completion does not, average elapsed time must rise.

This is an accounting relationship, not a management recipe. Software portfolios contain work of different size, risk, urgency, and uncertainty; priorities move, and items are redefined or abandoned. Little's Law cannot identify what to stop or prove that cancelling one project will accelerate another. It does show that an organization cannot indefinitely increase active inventory, preserve throughput, and expect average time to stay fixed.

A second 1961 result explains why full calendars can be a warning rather than proof of efficiency. John Kingman's heavy-traffic analysis showed how waiting becomes highly sensitive to variability as utilization approaches capacity. The exact model is a single-server approximation, not a literal representation of an enterprise delivery network. Its directional lesson is nonetheless important. Variability requires room. When a capability is planned at or near complete occupancy, even ordinary variation in arrival or service time produces disproportionate waiting.

Software work contains variation by design: investigation reveals facts, tests expose assumptions, policy needs interpretation, and production events interrupt plans. When every unit of capacity is allocated in advance, the system loses the room needed to absorb that variation.

The result is not that people become idle. The result is that work waits for people who are not.

Dependencies turn individual work into system work

Queueing explains why inventory and occupancy matter. It does not explain why one software change can involve so many people. For that, the unit of analysis must expand from tasks to dependencies.

Software components depend on one another, as do the people changing them. A small interface modification may need knowledge from a consuming team, security assurance, a policy decision, operating test data, and platform support. The code change may be small. The coordination path is not.

Marcelo Cataldo and colleagues developed a way to infer these “coordination requirements” from relationships between technical work and task assignments. Their research made a crucial point: coordination need is not determined only by reporting lines or meeting schedules. It is produced by dependencies in the work.

Two later empirical results show why that distinction matters.

James Herbsleb and Audris Mockus studied modification requests in two geographically distributed telecommunications software organizations. In both, work items involving more than one site took about two and a half times as long, on average, as similar single-site items. Distributed items involved more people, and the number of people involved was strongly related to elapsed time. Survey evidence also showed less frequent communication across sites.

The result is bounded to two early-2000s telecommunications settings, where geography was entangled with organizational history, expertise, communication, and work allocation. It does not establish a timeless “remote work penalty,” and 2.5 is not a forecast for another organization. It shows that when work crosses a boundary, the required interactions can become part of its elapsed time.

Cataldo and Herbsleb later examined two large projects in different companies. They compared the coordination implied by evolving work dependencies with the coordination that actually occurred. Gaps between the two were associated with more software failures; higher alignment was associated with improved development productivity.

This is observational evidence, not proof that adding meetings will improve software. Coordination itself consumes capacity. The finding is that work has an interaction structure, whether management recognizes it or not. When formal arrangements do not match it, individuals compensate through waiting, searching, negotiation, assumption, or rework.

This is how a capable person can be productive inside an unproductive system. The person completes what is locally available, but the outcome depends on interactions that no one participant owns.

Demand instability changes the work already under way

Fragmentation begins before code. It begins when an organization commits more demand than its production system can absorb and then attempts to manage the conflict by repeatedly changing priority.

DORA's 2024 cross-organizational survey associated unstable priorities with lower productivity and higher burnout. Its cross-sectional, self-selected design does not establish causality or a universal effect size. It corroborates a practical mechanism: a priority change does not merely move a line on a plan; it alters work already in motion.

People reconstruct context, renegotiate dependencies, preserve unfinished work, revise tests, and explain the consequences. The displaced item remains in the system after it leaves the executive agenda: it retains commitments, partially changed assets, unresolved decisions, and stakeholder expectations. If it returns, some knowledge has decayed.

Priority change can be rational; a security incident or legal deadline may require it. The system problem arises when reprioritization becomes the default means of reconciling demand with finite capacity. Urgency is then allocated by interruption: work that attracts attention moves, while displaced delay becomes less visible.

The resulting delay is commonly blamed on execution even though it was created partly by the pattern of authorization.

How fragmentation produces delay

The mechanisms can now be joined.

Demand enters through multiple channels: strategic programs, product commitments, operational problems, regulatory change, security findings, maintenance, supplier obligations, and local requests. Each channel has a credible claim. Because the requests are governed separately, the organization can approve more total work than any single decision maker can see.

The approved work moves into teams and specialist functions. It does not move as a single stream. It branches across architecture, data, security, legal, procurement, testing, infrastructure, operations, and business decisions. Some steps proceed in parallel; others cannot begin until preceding information or evidence exists. Every branch creates a chance to wait.

High local utilization lengthens those waits. A team may have capacity to write the change but not to obtain a test environment. The environment group may be busy because several programs reserved overlapping windows. A security review may wait behind other work because the scarce specialists are planned at full occupancy. A policy owner may answer late because software decisions are a small part of that person's role.

While work waits, the world changes. A dependency releases a new version. A requirement is clarified. A production incident consumes the people holding critical knowledge. A date approaches and the item is expedited. The expedite does not create capacity; it moves the item ahead of something else. That displaced work becomes older, riskier, and more likely to require reconstruction when attention returns.

Delayed feedback increases the cost of error. An incorrect assumption found within hours is information. The same assumption found after several teams have built against it is rework. Rework consumes capacity that had been promised to new demand. New demand therefore waits longer, creating pressure for more expediting.

This sequence forms a reinforcing pattern:

  1. Fragmented authorities release more work than the connected system can complete.
  2. High occupancy creates queues at teams, decisions, environments, and controls.
  3. Dependencies turn those queues into handoffs and coordination paths.
  4. Waiting delays feedback, allowing assumptions and conditions to diverge.
  5. Late discovery creates rework, while urgent items trigger priority changes.
  6. Rework, switching, and expediting reduce effective capacity for planned completion.
  7. Lower effective capacity lengthens queues, which prompts still more work and escalation.

Figure 2.3 represents this pattern. It is a causal hypothesis assembled from theory, empirical associations, and one longitudinal case. It is not a measured universal model. The solid relationships express established queueing mechanisms; dashed relationships express context-bound empirical associations; the case overlay shows where the Trilogy record exhibits the pattern without claiming that fragmentation was its sole cause.

Figure 2.3 — Fragmented demand can reinforce overload and delay. Multiple demand channels increase active work; queues and dependencies delay feedback; rework and expediting consume effective capacity and feed further waiting. The map is diagnostic author synthesis, not a universal effect-size model.

Source: Author synthesis based on A03, A04, A18, A23, A24, R02, R16, R17, R18, R19, and CS01.

The figure explains why more activity can reduce outcome yield. Work is not lost only when someone makes a mistake. It is also consumed in the spaces between locally rational actions: waiting for a decision, reconstructing context, reconciling incompatible changes, preparing evidence for a late review, or accelerating one item by delaying another.

The distinction between touch time and elapsed time helps, even when an organization lacks precise data. Touch time is the time in which someone is actively advancing an item. Elapsed time includes both touch and waiting. A change may require only a few days of direct effort while spending weeks in queues and dependencies. Asking an individual to code faster attacks only the visible portion. It may even make the item arrive earlier at the next queue.

This is not an argument that touch time never matters. Slow or poor-quality work exists. Skills matter. Tools matter. Automation matters. The analytical requirement is to estimate their place in the whole. If active effort is most of the elapsed time, improving execution may improve the outcome. If waiting, clarification, and rework dominate, local speed has a ceiling.

That ceiling belongs to the system.

The fragmentation flywheel

Leaders need a memorable way to recognize the pattern without pretending to diagnose an organization from a diagram. The fragmentation flywheel has five linked conditions: divided demand, saturated capacity, dependency distance, delayed truth, and compensating activity. It is descriptive, not a maturity model.

1. Divided demand

Demand is divided when different authorities can start or materially change work without sharing a view of the connected system's commitments. The issue is not that demand has multiple sources; a healthy institution must respond to many kinds of need. The issue is that authorization is locally complete and systemically incomplete.

A strategic program can be fully funded without reserving the scarce data expertise it will need. A product promise can be approved without accounting for a shared service already committed elsewhere. A control remediation can be declared urgent without identifying what existing work it displaces. The commitments become real at different moments and meet only in execution.

The diagnostic question is: Where can work be authorized, and where is its demand on shared capability made visible? A list of projects is not enough if it omits operational, assurance, maintenance, and unplanned work.

2. Saturated capacity

Capacity is saturated when a team, specialist, environment, or decision maker is planned so tightly that normal variation produces a queue. Saturation often appears efficient because every resource is assigned. The hidden cost is the absence of recovery space.

Not every fully occupied period is pathological. During a bounded emergency, concentrated effort may be necessary. A rare specialist may be economically scarce. The system condition appears when chronic occupancy is treated as the desired operating state while leaders continue to expect predictable response to variation.

The diagnostic question is: Which capabilities have a permanent queue, and what variation are they expected to absorb? Utilization without waiting time is an incomplete measure.

3. Dependency distance

Dependency distance is the effort required to cross from one necessary contribution to another. Distance can be geographical, organizational, contractual, technical, temporal, or cognitive. Two people sitting together can be far apart if neither has authority to resolve the issue. Teams on different continents can be close if interfaces, context, and decision paths are strong.

Specialization creates real value. Independent assurance protects institutions and users. Suppliers provide scale and expertise. The goal is not to eliminate boundaries. It is to understand the interaction burden created when work crosses them.

The diagnostic question is: For a typical consequential change, how many ownership, decision, evidence, and environment boundaries must be crossed—and who is responsible for the crossings? Counting organizational handoffs alone misses technical and cognitive distance.

4. Delayed truth

Truth is delayed when the system discovers important facts only after substantial commitment. The fact may be that a requirement was misunderstood, an interface behaves differently, a control is unmet, a user cannot complete the task, or an operational condition invalidates the design.

Software work always contains discovery. Delayed truth is not the existence of surprise; it is a feedback structure that makes surprise unnecessarily expensive. Long queues and handoffs separate action from consequence. Each step can be locally correct based on the information available and globally wrong once the parts meet.

The diagnostic question is: How long passes between making a consequential assumption and obtaining evidence that can disconfirm it? A process can have many reviews and still provide late feedback.

5. Compensating activity

Compensating activity is the work created to manage the flywheel's effects without changing its conditions. Status escalation, manual reconciliation, heroic integration, repeated replanning, priority meetings, exception handling, and emergency coordination can all be necessary. Their growth is also a signal.

This work is dangerous to interpret because it often looks like leadership. People who recover a date, resolve a crisis, or bridge a broken interface deserve credit. But an organization can become dependent on rescue. The rescuers acquire scarce contextual knowledge, which draws them into more work, which makes them a new constraint.

The diagnostic question is: How much management and expert attention is spent advancing outcomes, and how much is spent reconstructing, expediting, and reconciling work already authorized? The answer should not be used to blame coordinators. Their activity reveals the system they are compensating for.

As illustrated in Figure 2.4, the five conditions reinforce one another. Divided demand saturates capacity. Saturation increases dependency waiting. Waiting delays truth. Delayed truth produces compensating activity. Compensating activity consumes the capacity needed to complete the original work, making divided demand harder to reconcile.

The framework's value is not that every organization contains all five conditions at equal strength. It is that it prevents a local explanation from ending the inquiry too early.

The Fragmentation Flywheel

Divided demand → saturated capacity → dependency distance → delayed truth → compensating activity → reduced effective capacity → longer queues.

Every turn makes the system busier while making the whole harder to see.

Figure 2.4 — Five conditions can reinforce one another: divided demand saturates capacity, dependency distance delays truth, and compensating activity consumes the capacity needed to complete the original work. The framework is diagnostic and need not appear with equal strength in every organization.

Approved case study

Trilogy through the system lens

Return to the FBI case. The public record does not allow a precise reconstruction of every queue or handoff, and it should not be forced into the framework. It does, however, allow the framework to be tested against a documented sequence.

Divided demand was visible in the combination of mission change, infrastructure modernization, application replacement, accelerated commitments, congressional funding, FBI management, and contractor delivery. After 11 September 2001, the FBI had legitimate reasons to change priorities and expand information-sharing requirements. The problem was not that the mission changed. It was that a program already spanning three interdependent components had to absorb the change while its management foundations were incomplete.

Saturated or mismatched capacity appeared in management continuity and skill gaps as well as contractor dependence. The audit record identifies vacancies, turnover, inadequate oversight, and immature investment-management practice. These facts support an important counterargument: some of Trilogy's difficulty genuinely involved missing capability. A system explanation must include that scarcity, not explain it away.

Dependency distance existed between the functioning infrastructure components and the unfinished application capability. Hardware and networks were necessary, but their completion could not substitute for stable case-management requirements, usable software, data, operating practice, and acceptance. The program could therefore report substantial deployment while the institutional outcome remained absent.

Delayed truth was visible in the absence of final requirements and a firm baseline. Requirements growth is not inherently failure; learning should change a design. But without a controlled way to translate learning into scope, architecture, contract, test, and decision, new truth arrives as disruption. The cost of an unresolved requirement increases after software and dependent work have been built around it.

Compensating activity appeared in acceleration. The additional $78 million and revised schedule responded to a genuine national urgency. The infrastructure eventually improved and remained useful. Yet the accelerated date was missed by almost two years, and completion occurred only a month before the original target. The application component still failed. This does not prove that acceleration caused the failure. It shows that local urgency and component progress did not overcome the end-to-end production conditions.

Sentinel makes the case more valuable because it prevents a simple failure story. The FBI did eventually replace the old case-management system. After the initial Sentinel plan also encountered cost and schedule problems, the FBI changed its management and delivery approach in 2010, assumed more direct responsibility, and used a more iterative method. Sentinel became available to all users on 1 July 2012.

That outcome should not be romanticized. The inspector general reported a project figure of $441 million, while noting that it excluded two years of operations-and-maintenance costs and FBI personnel. Requirements had been added, modified, transferred, and removed. Of seventeen completion tasks reviewed, the inspector general found eight adequately supported, four partially supported, and could not verify three. The report also questioned parts of the framework used to measure progress.

Figure 2.5 — The longitudinal record shows component progress, VCF termination, and later Sentinel deployment after wider changes in management and delivery. It does not isolate one method as the cause of failure or recovery.

The responsible conclusion is not “agile saved Sentinel.” The evidence does not isolate method from leadership, scope, capability, time, accumulated learning, or changed oversight. The conclusion is that the institution eventually changed more than the code. It altered how it managed the work, brought decision and delivery responsibility closer, revised scope, and continued until a usable system was deployed—while the public record preserved qualifications about cost and completion.

The longitudinal case therefore tests the chapter's proposition in both directions. Fragmented activity can fail to produce an outcome. Changing the production conditions can coincide with progress, but no single practice deserves all the credit. The whole system remains the unit of explanation.

Executive takeaway

What executives should ask

An executive should not respond to this chapter by personally managing queues or drawing dependency maps for teams. The executive responsibility is to keep organizational commitments from becoming invisible production overload and to insist that system evidence accompany local reports.

Five questions change the quality of the conversation.

What outcome is moving, not merely what work is active? A portfolio can report percentage complete while the usable capability remains untestable. Ask for the smallest end-to-end evidence that the intended outcome is becoming real.

Where is work waiting, and what authority or capability does it need? Waiting at an approval, environment, supplier, or shared service is part of delivery time. It should not disappear because it sits outside a team's reporting boundary.

Which commitments compete for the same scarce capabilities? If every initiative is green within its own plan while all depend on the same constrained people or systems, the portfolio view is false. Shared demand must be visible before execution resolves it through delay.

What did the latest priority change displace? Urgent work can be necessary. Its cost includes the work interrupted, the knowledge that will decay, and the downstream commitments that must now change. Naming displacement turns reprioritization from a symbolic act into a decision.

Which explanations would disprove the system hypothesis? A genuine shortage, a poor technical choice, weak performance, an unavoidable external dependency, or a singular event may dominate. Leaders should ask for competing explanations. “The system” must not become a phrase that dissolves accountability.

These questions do not prescribe a universal operating model. They establish a burden of proof. If leaders want faster, more reliable outcomes, evidence about individual busyness is insufficient. They need evidence about the connected path through which work becomes an operating result.

Executive takeaway

What to remember

Five conclusions anchor the chapter.

  1. Activity is a local fact; flow is a system property. Completed tasks and occupied teams do not establish that an end-to-end outcome is advancing.

  2. Inventory, throughput, and time are connected. An organization cannot increase active work indefinitely while completion capacity remains fixed and expect elapsed time to remain stable.

  3. Dependencies create coordination work. When technical and organizational dependencies do not align with actual decision and communication paths, waiting and rework can emerge without any individual choosing them.

  4. Urgency reallocates delay; it does not create capacity. An expedite may be correct, but its displacement cost belongs in the decision.

  5. System explanations require boundaries and counterevidence. Skills, staffing, tools, and individual performance can be real constraints. The task is to locate them in the whole, not to replace one simplistic diagnosis with another.

Figure 2.6 — Chapter 2 shifts attention from local busyness to the connected conditions through which work becomes an operating result. The system explanation remains a hypothesis to test against genuine skill, capacity, technology, and external constraints.

The memory sentence is:

A delivery system can be busy everywhere and still move nowhere important.

Continue the argument

From diagnosis to a disciplined lens

Projects can organize important work, but managing each project harder cannot restore visibility that the delivery system itself has lost. The unresolved question is what operating principles can make the whole—its flow, constraints, quality, and feedback—governable without suppressing discovery and judgment.

Manufacturing offers a disciplined language for that question, along with a dangerous temptation to treat software as an assembly line. Chapter 3 asks what can transfer, what must be adapted, and where the analogy must stop.

Figure 2.7 — Chapter 2 resolves why local activity can coexist with system delay. Chapter 3 tests manufacturing as a disciplined lens while preserving software's discovery burden and reliance on judgment.

End of Chapter 2
A delivery system can be busy everywhere and still move nowhere important.
Return to contents