Research · Working paper · Version of 13 July 2026

Structural Waste in Digital Operations

A Lean Theory of How the Weakest Layer Limits Capability

Mohammad Rashid Azarang Esfandiari1 · Mohammad Reza Azarang Esfandiari, Ph.D.2

  1. 1 Independent Researcher, Mentu, San Pedro Garza García, Nuevo León, Mexico. ORCID 0009-0008-5528-4246
  2. 2 Professor Emeritus, Departamento de Ingeniería Industrial y de Sistemas, Tecnológico de Monterrey, Campus Monterrey
Cite this paper
@misc{azarang2026structural,
  title  = {Structural Waste in Digital Operations: A Lean Theory of How the Weakest Layer Limits Capability},
  author = {Azarang Esfandiari, Mohammad Rashid and Azarang Esfandiari, Mohammad Reza},
  year   = {2026},
  note   = {Working paper, version of 13 July 2026},
  url    = {https://rashidazarang.com/research/structural-waste-in-digital-operations}
}

Azarang Esfandiari, M. R., & Azarang Esfandiari, M. R. (2026). Structural Waste in Digital Operations: A Lean Theory of How the Weakest Layer Limits Capability [Working paper, version of 13 July 2026]. https://rashidazarang.com/research/structural-waste-in-digital-operations

Working paper, version of 13 July 2026. This is the full manuscript prepared for journal submission, not a summary. It supersedes the abridged version published here in May 2026, which predated the formal weakest-layer theorem, the seeded simulation, the corrected pilot, and the two pre-registered studies now reported in section 6 — including a decided null. Not yet submitted; not peer reviewed.

Abstract

Lean production made waste visible and eliminable, yet its migration to information systems has addressed service delivery, not the architecture that determines operational capability. We advance a theory of structural waste: operational inefficiency from architectural misalignment between system components, distinct from technical debt and process waste. By disciplined analogy with the Toyota Production System’s relational structure we derive a five-layer modal architecture (Data, Logic, Interface, Orchestration, Feedback), map the seven wastes to digital equivalents, and specify three sub-dimensions: coordination overhead, semantic drift, dependency concentration. The framework yields two instruments and three propositions: separation lowers waste; flow precedes automation; capability is bounded by the least mature layer. The third, the theory’s central claim, moves from analogy to proof to a decided test: proved as a weakest-layer band with estimable compensation budget, shown by simulation to be separable from additive models, and probed by a pre-registered 500-repository study whose frozen criterion was not met, a verdict on the proxy, not the cross-layer claim, reported in full. The null uncovered a fourth proposition: concentration trades flow efficiency for fragility, the knowledge analogue of just-in-time. Its signature replicated in two further registered samples; its tail prediction remains uncertified. Instrument validation lies ahead.

Keywords structural waste; Lean production; digital operations; modal layers; weakest-layer constraint; Toyota Production System; information systems theory

Introduction

The Toyota Production System (TPS) transformed manufacturing by making waste visible and systematically eliminable (Ohno 1988; Womack et al. 1990). Firms adopting Lean report substantial productivity gains, contingent in magnitude on implementation fidelity and organizational context (Shah and Ward 2003; Netland 2016). As organizational work becomes inseparable from information systems, a paradox is visible: firms equipped with sophisticated digital tools frequently experience more operational disorder than those running simpler systems (Westerman et al. 2014; Kane et al. 2019).

In a typical mid-size organization, analysts spend much of each week reconciling data that different systems report inconsistently; critical operations depend on individuals who “know where everything lives”; each new tool adds coordination burden without adding capability. The vignette is illustrative, not evidence; it names the quantities Section 4 is designed to observe. The inefficiencies persist despite heavy transformation investment and individually sound systems, pointing to a misalignment prevailing frameworks do not name.

Current treatments of Lean in digital settings are insufficient: Lean IT targets IT-service delivery rather than the architectural foundations that determine what operations are capable of (Bell and Orzen 2010); DevOps accelerates delivery but is largely silent on how business logic and data architecture shape operations (Kim et al. 2016); and the technical-debt metaphor captures expedient implementation, not the misalignment between business operations and system design (Kruchten et al. 2012). A system can carry low technical debt and still impose high operational friction, because the friction originates in how components are organized relative to one another, not in how any one is built.

This gap matters because digital architecture increasingly determines organizational capability. When business rules are scattered across systems, when data models fragment without reconciliation, and when processes depend on individual knowledge rather than systematic design, organizations incur what we term structural waste: inefficiency arising from architectural misalignment, not poor execution or poor implementation.1

The paper addresses four questions:

  1. How can the principles of the Toyota Production System be systematically translated to the architecture of organizational information systems?

  2. What forms of waste exist in digital operations that are distinct from both physical waste and technical debt?

  3. What measurement instruments can render structural inefficiency in digital systems observable and assessable?

  4. Is operational capability bounded by the architecture’s least mature layer, and can that claim be made precise enough to prove, and to test?

We make four contributions. First, we conceptualize structural waste as a distinct form of operational inefficiency, explaining why technically excellent systems still produce operational disorder. Second, we develop a five-layer modal architecture that transfers TPS principles to digital operations while respecting the properties that make information systems unlike physical production. Third, we propose measurement instruments, with a validation path, that make architectural misalignment observable. Fourth, we take the framework’s most distinctive claim from analogy to proof to a decided test: we prove, under explicit assumptions, that operational capability is bounded by the least mature layer (Theorem 1), with corollaries that make the compensation budget estimable and derive a second falsifiable signature; we show by simulation that this non-compensatory structure is statistically distinguishable from the additive specifications that predominate in empirical work (Dul 2016; Richter et al. 2020); and we report a pre-registered necessary-condition study of the claim’s most accessible operationalization, dependency concentration on 500 public repositories, frozen before collection; its registered criterion was not met, and we report that null in full. Exploring what it left behind uncovered a fourth proposition: dependency concentration trades flow efficiency for fragility, the knowledge analogue of just-in-time. Its cross-sectional signature replicated in two further pre-registered samples disjoint from the discovery sample, and its tail prediction, twice tested short of its registered event threshold, is handed to the field program with a precise design.

The paper is a work of theory that holds itself to an empirical standard. Its primary aim, in Whetten’s (1989) terms, is to supply the what, how, and why of a new construct and its relationships; the propositions we derive are offered for subsequent examination. But it does not stop at conceptual development: Section 6 proves the weakest-layer proposition’s bound, establishes by simulation that its signature is recoverable, and subjects its most accessible proxy to a pre-registered study with a frozen falsification criterion. Throughout we are explicit about which claims are conceptual, which are proved under stated assumptions, which have met data and with what result, and which still await the field program.

Theoretical Background and Approach

Theoretical approach

This paper builds theory rather than testing it. Conceptual theory papers earn their contribution by clarifying a construct, integrating literatures, or transferring an explanatory structure across domains, and are held to internal consistency, construct clarity, and the plausibility and testability of the relationships they propose (Jaakkola 2020; Corley and Gioia 2011; Murungi and Hirschheim 2022). Our development strategy is theory adaptation by disciplined analogy: we transfer the relational structure of the Toyota Production System to the domain of organizational information systems.

Analogical reasoning is a recognized engine of theoretical progress, but only when the transfer is structural, not superficial (Ketokivi et al. 2017): a surface analogy borrows vocabulary (labeling a backlog a “kanban”) without carrying over the relations that made the source concept explanatory; a structural analogy transfers that system of relations. We discipline the transfer with three requirements. First, relational correspondence: each digital construct must preserve the causal role its manufacturing counterpart plays, not merely its name. Second, domain fidelity: the transfer must respect the properties that distinguish information from material flow (invisibility, costless replication, embedded logic, and network interdependence; Section 2.3) and must fail gracefully where those properties break the analogy. Third, boundary specification: the scope within which the transferred structure is claimed to hold is stated explicitly (Section 3.4). Where an analogy holds only at the surface, we mark it as illustrative, not load-bearing.

This positions the framework alongside design-science and construct-development traditions in information systems research, which likewise proceed by articulating an artifact or construct and the reasoning that warrants it before subjecting it to empirical evaluation (Hevner et al. 2004; MacKenzie et al. 2011). The measurement instruments of Section 4 are the point at which the theory becomes testable; their validation is future work.

Toyota Production System principles

The Toyota Production System admits two canonical stratifications. The classic “TPS house” raises just-in-time and jidoka as its pillars on a foundation of standardized, stable work; the Toyota Way tradition instead names continuous improvement (kaizen) and respect for people as the pillars realized through the interlocking mechanisms beneath them (Liker 2004). The translation that follows deliberately cross-cuts both stratifications: it selects the five mechanisms with distinct operational modes: standardized work, visual management, flow, pull, and jidoka. A mode of operating, not a position in either hierarchy, is what can correspond to an architectural layer. In the layered architecture of Section 3.2 these occupy the four operating layers: standardized work at the Data layer, jidoka at the Logic layer, visual management at the Interface layer, and flow and pull, which operate as a coupled pair where jidoka stands alone, at the Orchestration layer. The fifth layer, Feedback, instantiates the kaizen practice itself, closing the loop the mechanisms serve. That the correspondents sit at different strata of the TPS hierarchies is a consequence of the selection criterion, not an oversight.

Standardized work establishes the current best-known method as the baseline against which abnormality shows (Spear and Bowen 1999). Visual management renders system state observable without investigation: andon, kanban, 5S (Galsworth 1997). Flow organizes work to move continuously, so that when it stops the cause cannot hide in buffers (Rother and Shook 1999). Pull triggers work from downstream demand, suppressing overproduction, the most fundamental muda (Hopp and Spearman 2004). Jidoka (autonomation with a human touch) embeds quality at the source, so defects do not propagate (Shingo 1986). Respect for people enters through the treatment of know-how, taken up at the Feedback layer (Section 3.2).

Information-systems theory and the character of digital operations

Information-processing theories hold that organizational design must match information-processing requirements: as uncertainty rises, an organization must reduce its need to process information or expand its capacity (Galbraith 1974; Tushman and Nadler 1978). The lens is central here because architectural misalignment is a source of unnecessary information-processing load: coordination that exists only because components were organized without regard to how information must move among them.

Digital operations differ from physical production in four ways that any faithful translation of Lean must accommodate. Invisibility: information flows have no physical presence; where work-in-process accumulates as visible pressure, data queues remain hidden inside databases (Davenport and Short 1990), so the framework must make state deliberately visible instead of assuming it as a by-product of physical form. Costless replication: digital assets copy without material cost, but low-cost replication breeds versioning divergence and synchronization overhead as copies evolve independently (Kallinikos et al. 2013). Embedded logic: business rules reside in code, configuration, and human memory rather than in physical constraints, which makes standardization difficult (Yoo et al. 2010). Network interdependence: digital systems form dense interdependencies in which change cascades in ways a linear manufacturing flow does not (Barabási 2016).

Existing approaches to digital waste

Prior research has documented digital inefficiency without reading it through a Lean lens: “technochange” failures when implementation ignores organizational design (Markus 2004), enterprise-system misfits (Strong and Volkoff 2010), and system benefits contingent on organizational complementarities, not technical features (Seddon et al. 2010). Lean information management has applied waste thinking to information flows (Hicks 2007), and a long tradition in data management has argued that the meaning of core entities is neither self-evident nor stable across contexts (Kent 2012), a theme we develop as semantic drift.

The technical-debt metaphor (Cunningham 1992) captures expedient implementation choices that demand future rework, and the literature has extended it well beyond code: architectural technical debt names sub-optimal architectural decisions within a system (coupling violations, eroded module boundaries), and practitioners rank architecture as the largest source of debt (Kruchten et al. 2012; Besker et al. 2018; Ernst et al. 2015; Li et al. 2015). Structural waste is nevertheless distinct from technical debt in all its forms, on two axes. Temporally, technical debt is a liability of past design decisions, quantified as the future rework (the metaphor’s principal and interest) required to correct them; structural waste is the present, recurring operational friction of running work through a given arrangement, which persists whether or not any rework is ever undertaken. Perspectivally, technical debt is assessed from the standpoint of the software artifact and the developers who must change it; structural waste from the standpoint of the business operation that runs across artifacts. The two vary independently: an ensemble of debt-free systems can still force operations to reconcile three definitions of “customer,” while a shortcut-laden monolith, being one system with one data model, can impose little cross-boundary friction. Section 3.1 formalizes this discriminant.

Enterprise-architecture research comes closest in spirit. Ross et al. (2006) argue that a firm’s foundation for execution rests on deliberate choices about process integration and data standardization, the same architectural levers whose failure our construct measures. But theirs is a prescriptive, firm-level design typology, whereas structural waste is a descriptive, measurable property of the running ensemble: a firm may have chosen the right operating model and still exhibit divergent definitions and duplicated rules. Closest of all is the recent enterprise-architecture debt stream, which names the gap between an architecture’s actual and its hypothetical optimal state and proposes methods to elicit and assess it (Hacks et al. 2019; Jung et al. 2021; Daoudi et al. 2023). Structural waste shares the premise that architectural deficiency is a first-class, assessable object, but differs in reference point and in what is measured. EA debt is defined against an optimum: a debt exists to the extent the architecture departs from an ideal or prescribed target, and its assessment is correspondingly expert- and target-relative. Structural waste is defined against operational consequence: the recurring friction a given arrangement imposes on work, measured by the coordination overhead, semantic drift, and dependency concentration it produces, with no optimum posited. The two are complementary, and the outcome anchoring is what makes structural waste directly instrumentable without first fixing an optimal architecture; the same stream has since shown such debt actionable at scale (Haki et al. 2023). Structural waste is also distinct from the enterprise-system misfit that Strong and Volkoff (2010) theorize: a misfit is a relational deficiency between one packaged system and the adopting organization, resolved by adapting one to the other. Structural waste is a property of the ensemble: it can be high where every individual system fits its local users well, because the friction arises between well-fitting systems (three locally adequate, mutually divergent definitions of “customer”) and persists even when no particular misfit can be named. Evidence that operational outcomes hinge on the architectural alignment of processes with digital infrastructure, rather than on any single system’s adequacy, points to this ensemble-level property (Bygstad and Øvrelid 2020). Relatedly, IS research has shown that architectural malleability and innovation stand in productive tension with the debt that accumulated architecture imposes (Rolland and Lyytinen 2021), and that platform and infrastructure renewal is inherently paradoxical (Wimelius et al. 2020); structural waste supplies the standing measure such renewal decisions would act on. Similarly, Lean software development translated the seven wastes to the software-development process (Poppendieck and Poppendieck 2003); those are process wastes of building software, at the execution level of Table 8, whereas structural waste concerns the architecture of the systems that development leaves behind. DevOps and Lean IT (Kim et al. 2016; Bell and Orzen 2010) optimize the delivery pipeline and IT services without questioning whether the architecture delivered matches operational need. The construct we develop next occupies the space these literatures leave open.

Theory Development

Structural waste as a construct

We propose that digital operations exhibit a form of waste that existing frameworks do not capture. Following guidance on construct clarity (Suddaby 2010), we define it, name its dimensions, distinguish it from adjacent constructs, and state its scope conditions.

Definition 1 (Structural waste). Operational inefficiency arising from architectural misalignment between the components of an organizational information system; it arises from how components are organized relative to one another, not from how any component is built or executed. Structural waste is manifested in three sub-dimensions: coordination overhead, semantic drift, and dependency concentration.

The unit of analysis is the organizational information system, the ensemble of coordinated systems, data, and rules through which an organization conducts its operations; not the firm, and not the individual application. Structural waste is a property of that ensemble’s architecture; firm-level conditions such as culture and strategy enter the theory only as moderators of how readily the architecture can be improved (Section 8.3), and application-level properties such as code quality are the province of technical debt, not of this construct.

The three sub-dimensions are conceptually distinct and jointly constitutive: the construct is present when any is elevated, and a system may exhibit one without the others: tolerated drift produces silent error, not coordination effort. Distinctness is a claim about loci and remedies, not statistical independence.

The effort required to maintain consistency across architecturally separated components: when related functions live in different systems without an integrating contract, human effort becomes the integration layer, scaling with the number of boundaries to reconcile.

Ungoverned divergence in the meaning of core concepts across system boundaries: when “customer” denotes subtly different things in different systems, every cross-system operation incurs a translation cost and aggregate figures silently mix incommensurable definitions. The qualifier is constitutive: bounded contexts with explicit, declared mappings are architectural discipline, not drift (Evans 2003); semantic drift is divergence without a declared mapping: meanings that were supposed to agree, or whose disagreement no contract records.

Operational reliance on specific individuals rather than systematic capability: processes embedded in human memory, where they are neither inspectable nor transferable.

Why these three, and not others.

The three sub-dimensions follow from the two conditions an operational architecture must satisfy to convert components into capability. First, the components must be coordinated: related functions must be brought into consistency at acceptable cost. Coordination theory identifies this managing of dependencies as the load an architecture either bears structurally or offloads onto people (Malone and Crowston 1994). Second, the meaning and know-how the operation runs on must be retained in the architecture itself rather than in individual heads, which is the province of organizational memory (Walsh and Ungson 1991). Each sub-dimension is a failure of one of these conditions: coordination overhead is a failure of coordination, so that dependencies are reconciled by human effort rather than by structure; semantic drift is a failure of retained meaning, so that shared definitions decay across boundaries; and dependency concentration is a failure of retained know-how, so that operational capability is stored in people rather than embedded in the system. The count follows from one individuation criterion applied to both branches: a failure is a distinct sub-dimension when it is held by a different mechanism and repaired by a different remedy. Coordination failure surfaces in several forms (duplicated rules, redundant interfaces, push sequencing; Table 9), but all share one mechanism (dependencies reconciled by human effort, not by structure) and one remedy family (explicit integration contracts), so they constitute one sub-dimension; that the surface forms covary as a single dimension is a testable expectation of the validation program, not an assumption. Retention failure splits under the same criterion: shared meaning and operational know-how are held by different mechanisms (declared mappings versus externalized procedure) and repaired by different remedies (semantic governance versus knowledge externalization), so retention contributes two sub-dimensions where coordination contributes one. Organizational-memory theory also individuates memory by function, as acquisition, storage, and retrieval (Walsh and Ungson 1991), and one may ask why retrieval cost is not a third retention failure; it is not, because failed retrieval manifests either as coordination effort or as concentration (only a person can retrieve it) and is carried by the existing sub-dimensions. The same criterion yields one further, economic distinction: coordination overhead and semantic drift are flow wastes, costs paid every operating period and visible in any cross-section, whereas dependency concentration is a contingent liability: a stock of risk priced only when the person it depends on is absent or gone, which may even carry a flow premium in undisrupted operation, since a single knowledge-holder needs no coordination to act. The sub-dimensions are not of the same temporal shape, and Section 6 shows the distinction has empirical teeth. Neither are they independent: divergent meaning and duplicated rules are what generate reconciliation effort, so the expected empirical texture is correlated sub-dimensions with distinct causes and remedies, examined, not assumed, by the validation program. Continuous improvement, the fourth concern an architecture must serve, is not itself a waste but the corrective loop through which the other three are surfaced and reduced; it is carried by the Feedback layer.

Discriminant positioning.

Structural waste is neither technical debt nor process waste, and Table 8 (Appendix C) locates it against both on seven attributes: locus, nature, cause, signature, consequence, remedy, and example. Technical debt, including its architectural form, is a liability of past design decisions within a system, payable as future rework (Kruchten et al. 2012; Besker et al. 2018). Process waste is a property of execution: the delays, hand-offs, and rework that value-stream analysis targets (Rother and Shook 1999). Structural waste is a property of architecture: it can coexist with clean code and efficient local processes, because its cause is the arrangement of components. The discriminant makes the construct non-redundant: different causes, different observable signatures, different remedies; a diagnosis that collapses them prescribes the wrong intervention.

The five-layer modal architecture

To make structural waste addressable, we decompose digital operations into modal layers, following the separation-of-concerns principle (Parnas 1972) and Lean’s emphasis on visual clarity. We call the layers modal in the grammatical sense of mode: each names a distinct mode of operation, not a technical tier of a software stack: knowing (Data), reasoning (Logic), presenting (Interface), sequencing (Orchestration), and improving (Feedback).

Definition 2 (Modal layer). An architectural domain with defined boundaries, transformation rules, and interfaces, independently modifiable except through explicit contracts. Two layers interact only through a declared interface; neither may reach into the other’s internal state.

Relation to layered and enterprise architecture.

The modal layers are not the tiers of a layered software architecture (presentation–business–data) nor the viewpoints of an enterprise-architecture reference model such as Zachman or TOGAF: those stratify a system by technical implementation concern, whereas the modal layers stratify operations by mode of operational work, each defined by the TPS principle it instantiates and the sub-dimension it contains (Table 9). A conventional three-tier application, cleanly built, can still exhibit high structural waste, because a technically well-separated data tier does nothing to prevent the same entity from carrying divergent meaning across systems. The modal decomposition is an operational lens layered onto whatever technical architecture exists, not a competitor to it.

Five layers span the operational concerns the framework targets (Figure 1): four instantiate TPS principles from Section 2.2, the fifth the kaizen pillar those principles serve. The decomposition does not claim to partition every concern: security and identity, observability, and the memory concern developed below cut across all five layers.

Establishes the factual baseline. The correspondence is pitched at standardized work’s role as the authoritative baseline against which deviation shows: one authoritative source for each entity, against which inconsistency becomes visible. Standardized work’s method content (sequence, timing, work-in-process) maps not here but to the Orchestration and Logic layers, a partial transfer we flag rather than paper over. It is the layer at which semantic drift is contained.

Externalizes business rules from implementation code. As jidoka builds quality in by detecting abnormality at its source, the Logic Layer makes business rules explicit, versioned, and enforceable at the point of decision, so that invalid states are caught where they arise, not discovered downstream as reconciliation and manual verification, the digital form of the inspection waste jidoka eliminates. Jidoka itself decomposes in the transfer, and we state the split: its detect-and-block role (the machine stops the defective unit at its station) lives here, in enforcement at the point of decision, while its halt-and-escalate role (the andon pull that summons attention and fixes the root cause) is a learning act that lives at the Feedback layer. One chain unifies the layer’s two appearances in the tables that follow: a rule defined once can enforce; rules duplicated across systems diverge; and divergent rules cannot stop what they no longer agree is a defect: logic duplication and unenforced validity are the same failure seen from two sides. It is the layer at which logic duplication is contained.

Manages the points at which humans and systems interact. As visual management renders system state observable without investigation (Galsworth 1997), the Interface Layer presents exactly the information a decision requires, at the point of use. It is the layer at which redundant and unused interfaces are contained.

Governs how work is sequenced and triggered. It instantiates flow and pull by moving work on downstream demand, not upstream push. It is the layer at which coordination overhead is contained.

Connects outcomes back to change. It instantiates kaizen and the plan–do–check–act cycle, closing the loop from measurement to improvement. It is the layer at which dependency concentration is contained: closing the improvement loop is the mechanism by which know-how is drawn out of individuals and re-embedded in the architecture, so that operational memory is held by the system rather than by the people who happen to run it (Walsh and Ungson 1991). The aim is not to render people interchangeable but to stop using them as the system’s memory, consistent with the respect-for-people pillar. Memory is not a sixth layer but a concern that runs across the stack, retained at the Data layer and returned at the Feedback layer, and the failure of that retention-and-return is what dependency concentration measures. The TPS analogy runs one level deeper here than a waste mapping: dependency concentration is the knowledge analogue of just-in-time: operating without knowledge inventory is fast, because a single holder needs no coordination to act, and fragile, because the operation has no buffer against that holder’s absence. TPS made JIT survivable not by restoring buffers but by the system built around it: stable, standardized work beneath, and the jidoka discipline of rapid detection and organizational response when the line stops; andon is that discipline’s signaling instrument, not its substance. The architectural counterpart of the detection-and-response half is this layer: a strong Feedback loop is what makes concentrated knowledge survivable, by continuously externalizing it before the shock arrives. Proposition 4 states the prediction; its provenance is empirical (Section 6.4).

The five-layer modal architecture. Each layer carries a core TPS principle (left) and exhibits a characteristic layer-level symptom (right) that rolls up to one of the three sub-dimensions (brackets). Layers are independently modifiable only through explicit interface contracts (double-headed connectors); the Feedback Layer closes a kaizen/PDCA loop back to the Data Layer, the cross-cutting memory path along which meaning is retained and returned.

The construct and the architecture operate at two levels, connected by Table 9 (Appendix C): each layer exhibits a characteristic symptom that rolls up to one of the three sub-dimensions, so the five layer-level symptoms are not a competing decomposition but the surface at which the sub-dimensions become locally observable. Coordination overhead, the most pervasive, surfaces at more than one layer (as duplicated logic, redundant interfaces, and push-based sequencing) because coordination is what an architecture must achieve across every boundary; semantic drift concentrates at the Data layer; dependency concentration is contained at the Feedback layer.

Translating the seven wastes to digital operations

The seven wastes (muda) of the Toyota Production System were defined for material flow. Under the discipline of Section 2.1, each transfers to digital operations by preserving the causal role of the original waste (what it consumes and why it is waste), not its material form. Table 7 (Appendix C) states each translation, its primary layer, and its mechanism; in outline: overproduction becomes reports nobody consumes; waiting, items awaiting manual hand-off; transport, moving data between systems to reconcile them; over-processing, re-entering the same data; inventory, accumulated work-in-process whose problems surface late; motion, navigating multiple tools for one task; and defects, inconsistent values propagating because validity is not enforced at the source. Toyota-Way practice also names an eighth waste, the underused talent of people (Liker 2004), a later addition to Ohno’s canonical seven (Ohno 1988). The mapping to dependency concentration is not direct, and the bridge is load-bearing: the eighth waste is unused creativity, people’s improvement capacity left idle, whereas dependency concentration is over-reliance on people as operational memory. The two connect through a single mechanism: people consumed as the system’s memory are people unavailable as its source of improvement, so dependency concentration is the mechanism by which talent goes unused. We include the row with that bridge stated, not as a direct correspondence.

The inventory translation is marked as partial, because the method of Section 2.1 requires the analogy to fail gracefully where the domains differ. In TPS, inventory is waste twice over: it ties up capital, and it conceals problems, the water level that hides the rocks. The capital mechanism does not survive costless replication (Section 2.3); only the concealment mechanism transfers, which is why the digital equivalent is accumulated work-in-process whose problems surface late (Poppendieck and Poppendieck 2003), and why stale caches and divergent copies, sometimes offered as digital “inventory,” are better understood as a Data-layer source of semantic drift. Marking this boundary is the graceful failure the method promises. The translation’s scope is deliberately muda: TPS’s companion categories mura (unevenness) and muri (overburden) are not translated as waste forms but surface in the framework’s propositions, unevenness in Proposition 2’s concern with unstable flow and overburden in Proposition 3’s concern with the binding layer (Ohno 1988; Liker 2004).

Boundary conditions

Scope conditions are part of a construct’s clarity (Suddaby 2010); stating them prevents the theory from being applied where its mechanism does not operate. The framework addresses organizational information systems that support business operations, and is most relevant when

The second condition is keyed to all three sub-dimensions: a single-system organization can still exhibit semantic drift across modules or dependency concentration in one indispensable person, and remains in scope. Whether architectural modification is economically feasible is deliberately not a scope condition (a system whose remedy is uneconomic still exhibits the waste), but it does bound the interventions of Section 8.3: the descriptive theory applies wherever the mechanism operates, prescriptions only where change pays. The framework does not apply, or applies only weakly, to consumer applications, hardware-constrained embedded systems, or purely computational systems without business logic, settings where the transferred TPS structure loses its correspondence.

Measurement Framework

A construct becomes scientifically useful when it can be measured. This section proposes instruments for structural waste and specifies how they are to be validated. The instruments are proposed: we develop their conceptual foundation here and set out the empirical program required to establish their reliability and validity, which we do not claim to have performed.

Formative and reflective measurement

The instruments differ in measurement structure, and conflating the two is a well-documented source of specification error (Petter et al. 2007; MacKenzie et al. 2011): a reflective construct causes its indicators, which are interchangeable manifestations of a latent variable and should covary; a formative construct is composed by its indicators, which define it, need not covary, and are not interchangeable.

The Structural Waste Assessment Tool (Section 4.4) is reflective: within a layer, the survey items are alternative symptoms of the same underlying waste, expected to covary, and assessed by internal-consistency and factor-analytic criteria. The Layer Separation Index (Section 4.3) is formative: it is an index that aggregates separation across five architecturally distinct layers, which need not covary (an organization may separate its Data Layer yet conflate its Orchestration Layer), and must be validated as a composite, not by internal consistency.

Sub-dimensions of structural waste

We define four measurable quantities that operationalize the three sub-dimensions of Section 3.1. Coordination overhead is captured by two quantities of different epistemic standing: a direct measure of the effort it imposes (coordination cost) and an antecedent-based indirect indicator (logic duplication), a cause–effect pair, not two reflections of one latent, so studies should model logic duplication as a structural predictor of coordination cost, not pooled with it. The mapping is thus 343\to 4, and the asymmetry follows from the sub-dimensions’ types: coordination overhead is the construct’s one effort-typed sub-dimension, so it alone requires both the effort measure and its principal structural antecedent; semantic drift and dependency concentration are defined structurally, so their measures already are their structural form, with no separate unmeasured antecedent to instrument.

Appendix C states each quantity as an operational definition with an explicit formula, its normalization rationale, and its anchor in a neighboring measurement tradition. In brief: coordination cost (CC\mathrm{CC}) is reconciliation person-hours as a share of operational effort; logic duplication (LD\mathrm{LD}) is the mean number of locations implementing each business rule, normalized by the declared measurement scope (Roy et al. 2009); semantic drift (SD\mathrm{SD}) is the share of shared entity fields whose definitions diverge without a declared mapping (Wang and Strong 1996; Rahm and Bernstein 2001); and dependency concentration (DC\mathrm{DC}) is a Herfindahl concentration of person-dependent process steps over the individuals whose joint absence halts them, always reported with its companion coverage statistic DCov\mathrm{DCov} (the share of steps that are person-dependent at all) (Avelino et al. 2016). Concentration captures single-point-of-failure risk and coverage the extent of the retention failure; the two are orthogonal by construction. Every boundary these definitions require (what counts as a rule, a location, a shared field, a person-dependent step, or reconciliation work) is a coding judgment governed by the protocol requirement of Section 4.5.

We deliberately do not define a single scalar “structural waste score”: they carry different units, and by Proposition 3 waste is not compensatory across layers, so the object of measurement is the profile. Where a study requires joint modeling, each quantity is standardized against a defined reference population and the aggregation model is stated (Section 4.1).

Layer Separation Index (LSI)

The Layer Separation Index summarizes how cleanly an organization’s architecture separates the five modal layers: LSI=lwlsl,={Data, Logic, Interface, Orchestration, Feedback},\mathrm{LSI} = \sum_{l \in \mathcal{L}} w_l\, s_l, \qquad \mathcal{L}=\{\text{Data, Logic, Interface, Orchestration, Feedback}\}, where sl[0,1]s_l\in[0,1] is the separation score for layer ll and wlw_l is its weight, with lwl=1\sum_l w_l = 1. Each sls_l is the mean of five layer-specific indicators scored on a 0011 rubric (Appendix A), an equal-weight formative choice made on the same grounds as the layer weights, which are fixed a priori and equal (wl=1/5w_l = 1/5). Equal weighting is not a neutral act (no weighting of a formative index is; (Cenfetelli and Bassellier 2009)) but a content-validity commitment that follows from the theory: the framework gives no layer theoretical priority, and Proposition 3 makes the profile’s minimum the load-bearing quantity, so differential weights would encode a compensatory claim the theory rejects. Because LSI is formative, the weights are part of the construct’s definition and must be reported. The validation program accordingly treats weight estimation as a diagnostic (do estimated weights depart materially from equality?), not as a rescoring device; and any estimation criterion must be independent of outcomes later used to test the propositions, since weights tuned to an outcome make tests against it circular (Cenfetelli and Bassellier 2009).

Two functionals of the profile serve two theoretical roles. The weighted sum summarizes the overall level of separation and is the antecedent measure for Proposition 1. Because Proposition 3 posits a non-compensatory constraint, the sum cannot carry it (a high aggregate reached through four strong layers and one failed layer would misrepresent an architecture pinned at the failure), so its load-bearing quantity is the profile minimum, LSImin=minlsl\mathrm{LSI}_{\min} = \min_l s_l, tested as a necessary-condition hypothesis (Dul 2016; Richter et al. 2020), not as an additive effect. As a formative index, the LSI also requires an identification strategy of its own: the validation program pairs it with a global reflective separation item per layer to permit redundancy analysis, and subjects the indicator set to content validation that the five indicators span each layer’s separation facets (Diamantopoulos and Winklhofer 2001; MacKenzie et al. 2011).

Measurement error and the profile minimum.

Because LSImin\mathrm{LSI}_{\min} is an order statistic of imperfectly measured scores, it inherits a known bias: the minimum of noisy estimates is biased below the true minimum, the more so as error variance grows and layer scores cluster. Left uncorrected, this blurs the ceiling and attenuates the estimated effect (conservative for detection) while also scattering genuinely feasible observations into the theoretically empty region, manufacturing spurious ceiling violations; diagnostically, the “weakest layer” flagged in any one measurement is disproportionately the layer whose score drew a negative error. The validation program therefore requires that layer scores be reliability-corrected (or shrunken) before the minimum is taken, that the minimum’s own reliability be reported alongside it, and that necessary-condition analyses of LSImin\mathrm{LSI}_{\min} carry a sensitivity analysis over assumed error variance. Where profile data permit, per-layer necessary-condition tests (Section 5.1) provide a complementary route that does not rest on a single noisy order statistic, and one matching the theorem’s path scoping (Appendix C).

Structural Waste Assessment Tool (SWAT)

The Structural Waste Assessment Tool is a proposed 25-item self-report instrument, five items per modal layer, each rated on a five-point frequency scale (1=1= never true to 5=5= always true; Appendix B). Items are worded as concrete, observable symptoms (“different systems show different values for the same metric”), with frequency carried by the response scale, not embedded in the stem. Because the instrument is reflective, every item is worded as an experienced symptom that waste would cause; the underlying architectural conditions are the business of the formative LSI, and the tautology screen of the validation program polices the boundary between the two instruments. Within a layer the items are reflective indicators of that layer’s latent waste, so a layer subscale is the sum of its five items and the instrument yields a profile across layers. The factorial structure follows the two-level model of Section 3.2 with one identification caveat stated openly: coordination overhead, indicated by three first-order layer factors (Logic, Interface, Orchestration), is a genuine second-order factor and testable as such; semantic drift and dependency concentration are each carried by a single first-order factor (Data and Feedback), and a second-order factor with one indicator is not identified apart from that indicator. The two singleton sub-dimensions are therefore modeled as their layer factors, and the validation program tests the identified structure against flat five-factor and three-factor alternatives, rather than promising a comparison the model cannot deliver. Two items (10 and 16) are pre-registered as expected cross-loaders. Interpretive cut-points and population benchmarks are deliberately omitted: such norms can only follow from a validated administration to a defined population, and are an output of the validation program rather than an assertion in advance.

A validation program

Because the instruments are proposed, not validated, their development must follow an explicit psychometric protocol before any score is interpreted substantively. We adopt the standard sequence for construct measurement in behavioral and information-systems research (Churchill 1979; MacKenzie et al. 2011), summarized in Figure 4 (Appendix C). The conceptual stages (domain specification, item generation, and content assessment) are begun in this paper; the empirical stages are future work.

  1. Domain specification. Define the construct and its sub-dimensions, delimit scope, and specify each instrument as formative or reflective (Sections 3.14.1).

  2. Item generation. Draft an indicator pool per dimension from the conceptual definition and the layer indicators (Appendices AB).

  3. Content validity. Submit items to an expert panel; retain, revise, or discard by content-validity ratio; include an expert item-sort of SWAT items to layers and sub-dimensions, retaining only high-agreement assignments; and screen for item pairs in which a SWAT item restates an LSI indicator with inverted polarity: shared content between antecedent and criterion instruments would render their association tautological.

  4. Pilot and purification. Administer to a development sample; for the reflective SWAT, examine item–total correlations and run exploratory factor analysis against the identified structure of Section 4.4, comparing it with flat five-factor and three-factor alternatives and treating items 10 and 16 as pre-registered cross-loaders; for the formative LSI, screen indicators for excessive collinearity and confirm each contributes distinct content.

  5. Reliability and structure. On an independent sample, estimate internal consistency (Cronbach’s α\alpha with McDonald’s ω\omega alongside, since near-duplicate items inflate α\alpha without adding content) and confirm structure by confirmatory factor analysis for reflective subscales, testing essential tau-equivalence before summing; for the formative index, assess whether empirically estimated weights depart materially from the definitional equal weights (a diagnostic, not a rescoring; Section 4.3).

  6. Construct validity. Establish convergent and discriminant validity (average variance extracted against inter-construct correlations; for the LSI–SWAT pair, discriminant validity between the instruments themselves) and nomological validity by testing the propositions against criteria external to both instruments: cycle time, change lead time, error and rework rates, change-failure rate. LSI–SWAT association is convergent evidence, not a nomological test. The direct quantities offer further convergent anchors from artifacts rather than perception, as do neighboring-field instruments (clone detection for LD, schema-matching distance for SD, truck-factor estimates for DC).

Three requirements attach to the whole program. Sources and levels. The SWAT is perceptual: multiple informants per organization spanning operational roles, composed to the system level under a stated composition model with aggregation justified by inter-rater agreement (Chan 1998); the LSI is an assessor-scored architectural audit, and the two instruments must not share a source. Method variance. Self-report predictors paired with self-report criteria inflate associations, so the program requires procedural separation (LSI from audit, outcomes from archival records, temporal separation) and statistical checks (Podsakoff et al. 2003). Coding reliability. The quantities of Section 4.2 rest on coding judgments; each requires a coding manual with explicit inclusion rules and worked examples, double-coding of a subsample, and a reported reliability coefficient meeting a stated threshold (Hayes and Krippendorff 2007).

Research design for proposition testing.

Psychometric adequacy does not by itself license the propositions’ causal readings. Proposition 1 is threatened by selection and reverse causality (organizations that separate layers may be better managed and resourced, and low-waste organizations may find separation cheaper), so its test requires panel designs with organization fixed effects, a named confounder set (size, age, industry, IT investment and capability, prior performance), and, where feasible, quasi-experimental leverage around discrete separation initiatives (Antonakis et al. 2010). Proposition 2 is a claim about intervention sequence and is not testable in cross-section: it requires longitudinal observation of automation initiatives or validated retrospective coding of whether flow was established before automation. Proposition 3, as a necessary-condition claim, is tested on the profile minimum with necessary-condition analysis, not additive regression (Section 4.3; (Dul 2016)).

Theoretical Propositions

The framework yields three propositions, each stated in antecedent–mechanism–consequent form with its boundary. They form a connected nomological network (Figure 3): layer separation is the antecedent, structural waste the mediating state, operational outcomes the consequence, with Propositions 2 and 3 specifying a sequencing condition and a constraint.

Proposition 1 (Layer separation). Information-system architectures that maintain explicit separation between modal layers exhibit lower structural waste than conflated architectures, other things equal.

Mechanism. Separation creates bounded contexts that localize the impact of change: a change to business logic does not force interface changes, nor a data-model update cascade through orchestration, so coordination overhead falls because fewer boundaries must be reconciled per change. The proposition specializes modularity theory (Baldwin and Clark 2000; Sanchez and Mahoney 1996) to a distinct objective function: the layers modularize by mode of operational work, not by design evolvability, so the predicted return to modularity is lower operational waste, an association modularity theory, whose dependent variables are flexibility and innovation, does not assert. Boundary. The proposition presumes the scope conditions of Section 3.4; where separation is pursued beyond the point of economic return (Section 8.4), the mechanism weakens as separation itself adds coordination cost. A second boundary follows from the mirroring hypothesis: architectures tend to reproduce the communication structure of the organizations that build and run them (Conway 1968; MacCormack et al. 2012; Colfer and Baldwin 2016), so layer separation pursued without corresponding realignment of ownership and communication boundaries is predicted to erode back toward the organizational mirror. Organizational structure is therefore a co-determinant of the antecedent, not merely a moderator of its effect, and durable separation is a socio-technical intervention. This yields a testable failure mode: track separation trajectories sl(t)s_l(t) after separation initiatives and compare decay slopes with versus without accompanying ownership realignment: the mirroring boundary predicts faster decay in the unaccompanied group.

Proposition 2 (Flow before automation). Digital operations that establish information flow prior to automation achieve greater waste reduction than those that automate without first establishing flow.

Mechanism. Automation amplifies whatever pattern it encodes. Establishing stable, standardized flow first surfaces and removes waste before automation fixes it in place; automating an unstable process locks in its waste and raises the cost of subsequent correction. The sequencing principle is Lean’s own (Toyota introduces technology only in service of a process already stabilized and mastered manually; (Liker 2004)), with two independent anchors: the reengineering injunction against automating an unredesigned process (Hammer 1990), and Industry 4.0 evidence that lean process maturity and digitalization are complements (Buer et al. 2018, 2021; Cifone et al. 2021). One caution: that evidence concerns physical lean maturity, so it enters as support-by-analogy across the same boundary this paper polices elsewhere; the proposition’s direct warrant is the domain-general amplification mechanism. Boundary. The claim concerns the sequence of interventions, not whether to automate; it presumes flow can be established at acceptable cost (Section 3.4). Because the claim is about sequence, it is untestable in cross-section; its natural test is on coded event pairs (for each automation event, whether a flow-stabilization event precedes or follows it), with subsequent waste reduction as the outcome.

Proposition 3 (Weakest-layer constraint). Operational capability is bounded by the least mature modal layer: beyond a limited range of compensation, sophistication in other layers cannot offset the binding layer.

Mechanism. Maturity is defined in the framework’s own vocabulary: a layer’s separation score sls_l (Section 4.3) together with its waste subscale (Section 4.4); a mature layer is separated and low-waste. Information quality degrades as it passes through an immature layer along the operational flow path from Data through Logic to Interface, and downstream layers cannot fully recover what an upstream layer failed to preserve. A sophisticated Interface Layer can compensate for a fragmented Data Layer only by adding reconciliation work, compensation purchased in the currency the construct measures, so partial compensation is possible but bounded, and beyond its economic limit the least mature layer imposes a systemic constraint in the sense of the Theory of Constraints (Goldratt and Cox 1984): improving non-constraining layers yields little until the binding layer is addressed (Figure 2). What this adds beyond a TOC reading: the unit the constraint runs over is the modal layer, defined by the TPS principle it instantiates; the verbal bottleneck becomes a proved weakest-layer band with an estimable compensation budget (Section 5.1); and the constraint is empirically separable from a compensatory account by a necessary-condition test (Section 6.1). Two kinds of coupling must be distinguished for Propositions 1 and 3 to cohere: Proposition 1 concerns change coupling, which separation localizes; Proposition 3 concerns information-quality flow, which no architecture eliminates. Separation reduces the cost of changing the system; it does not exempt running work from the quality of what upstream layers deliver. Boundary. The constraint claim is stated for the dominant operational flow path; in densely networked settings (Section 2.3) constraints may be distributed, and the proposition then applies to the binding constraint of each major flow. The bound is on capability attributable to architecture; strategy, market, and resourcing also bound performance.

A formal statement of the weakest-layer bound

The mechanism above can be stated precisely, which both sharpens the claim and makes it falsifiable. Consider an operational flow path as an ordered sequence of layers l=1,,Ll = 1, \dots, L; the canonical exposition instance is Data \to Logic \to Interface. The path is a modeling choice, the flow under test, not an empirical discovery: the result holds for whichever path is modeled, and in networked settings it applies to each path separately. Let sl[0,1]s_l \in [0,1] be the maturity of layer ll as the mechanism above defines it (separation with waste subscale), and let ql[0,1]q_l \in [0,1] denote the quality of the information delivered out of layer ll, with perfect input q0=1q_0 = 1. Realized operational capability is the quality delivered at the end of the path, YqLY \equiv q_L.

Assumption 1 (Ceiling). No layer emits information of higher quality than its own maturity permits: qlslq_l \le s_l for all ll. A layer cannot compensate beyond its own competence.

Assumption 2 (Bounded compensation). A layer may raise incoming quality toward the target by adding reconciliation work, but only within a finite per-layer budget κl0\kappa_l \ge 0: qlql1+κlq_l \le q_{l-1} + \kappa_l. The budget is a structural parameter of the organization (the reconciliation capacity in labor, computation, and attention that can be devoted to layer ll per operating period), to be estimated, not assumed (Corollary 1); the theorem requires only that it is finite.

Assumption 3 (Efficiency). A layer does not degrade information below the lesser of its own competence and the quality it receives: qlmin(sl,ql1)q_l \ge \min(s_l,\, q_{l-1}). This is the formal content of the construct’s scope condition that structural waste is capability-bounded: layers fail by being unable to do better, not by gratuitously doing worse.

Write K=l=1LκlK = \sum_{l=1}^{L}\kappa_l for the total compensation budget along the path.

Theorem 1 (Weakest-layer band). Under Assumptions 13, realized capability is confined to a finite band anchored at the profile minimum: minlslYminlsl+K.\min_{l} s_l \;\le\; Y \;\le\; \min_{l} s_l \;+\; K . In particular, with no compensation (K=0K=0) capability equals the weakest layer exactly, Y=minlslY = \min_l s_l; and for any fixed budget K<1K < 1, capability cannot reach the level implied by the stronger layers once the weakest layer falls more than KK below them. Both bounds are attainable.

The proof (induction for the lower bound, telescoping the compensation budget for the upper) is given in Appendix C, together with explicit attainment witnesses, the deterministic attainment law whose closed form identifies which layer binds (Remark 1), and a second, independently testable signature: maturity investments raise capability only when they land on the binding layer (Corollary 2). The band, closed form, and witnesses are machine-verified over 10510^5 random profiles, including feasible quality sequences that satisfy the inequalities without following the attainment law (supplementary code).

Corollary 1 (Empty-region prediction; KK is estimable). Under Assumptions 13, every observation in the (minlsl,Y)(\min_l s_l,\, Y) plane lies in the band {(x,y):xyx+K}\{(x, y) : x \le y \le x + K\}: the region above y=x+Ky = x + K is empty, the ceiling structure necessary-condition analysis is built to detect (Dul 2016; Richter et al. 2020). Critically for falsifiability, the vertical offset of the estimated ceiling above the 4545^\circ line estimates the compensation budget KK: not free slack that could absorb any observation, but a measured quantity the theory predicts to be small; a large K̂\hat{K} falsifies the claim that compensation is meaningfully bounded. The band’s lower edge is equally diagnostic: observations materially below the 4545^\circ line indicate gratuitous degradation (violating Assumption 3) or measurement error.

Scope of the formalization.

The theorem is stated for a single forward pass over an ordered flow path. Data, Logic, and Interface occupy path positions; Orchestration governs the compensation budgets κl\kappa_l (coordination capacity is what purchases reconciliation) rather than occupying a position; Feedback is a return path whose effects are cross-period, entering the erosion dynamics of Proposition 1, not this static bound. For any single layer jj the upper bound specializes to Ysj+KY \le s_j + K, so one observable coordinate supplies a partial necessary-condition test of the min-governed prediction. This is the license for the dependency-concentration analyses of Section 6, and equally their limit: a one-coordinate test cannot distinguish the full min-structure from a single necessary condition.

Theorem 1 is what distinguishes the claim from a conventional capability model. An additive account predicts YlwlslY \approx \sum_l w_l s_l, a weighted mean in which a strong layer offsets a weak one; the bound instead makes the minimum the governing quantity up to a finite slack KK. A profile with one weak layer (s=(0.9,0.9,0.35,0.9,0.9)s = (0.9,0.9,0.35,0.9,0.9), say) yields an additive prediction near 0.790.79 but a bound-implied capability of at most 0.35+K0.35+K, a gap detectable in data for realistic small budgets, and the reason Proposition 3 is tested as a necessary-condition hypothesis (Corollary 1) rather than by additive regression: the theorem predicts a ceiling, not a slope.

A fourth proposition, uncovered by the exploratory analysis of Section 6.4, completes the set.

Proposition 4 (Concentration trade-off). Dependency concentration trades flow efficiency against resilience: in undisrupted operation, higher concentration is associated with equal or better average operational responsiveness and with greater dispersion of outcomes; under key-person loss, the operational damage is increasing in prior concentration. The Feedback layer moderates the tail: where improvement loops continuously externalize know-how, the damage of a given concentration is capped.

Mechanism. The proposition follows from the sub-dimension typing of Section 4.2: concentration is a contingent liability, not a flow waste. A single knowledge-holder eliminates the coordination overhead distributed ownership must pay (the flow premium) while embedding the operation’s memory in one person, so its cost is realized at absence and departure events rather than in any undisrupted cross-section: the knowledge analogue of just-in-time (Section 3.2), speed purchased by removing buffers, survivable only under a strong feedback discipline. Boundary. The trade-off binds only where the concentrated knowledge is load-bearing (high DCov\mathrm{DCov}); concentration in peripheral know-how carries neither the premium nor the tail. Form. Unlike Proposition 3, this is not a necessary-condition claim: its signatures are cross-sectional dispersion and an event-study dose-response, tested in Section 6 on fresh data after the former was observed exploratorily.

The weakest-layer constraint: realized capability (solid) is bounded at the least mature layer (here Data) while capability latent in more mature layers (hatched) goes unrealized. A schematic of the mechanism, not measured data.
The research model as a nomological network: layer separation reduces structural waste (P1); flow-before-automation sequencing moderates the separation–waste path (P2); the weakest-layer constraint bounds attainable outcomes (P3). Boxes are constructs; sub-text lists their indicators.

A methodological note: because these propositions constitute the construct’s nomological network, testing them supplies the nomological evidence for the instruments; but since the LSI and SWAT partly enumerate complementary descriptions of the same conditions, their mutual association is convergent rather than nomological evidence, and the propositions are tested against operational criteria external to both instruments (Section 4.5).

Formal Analysis and Preliminary Evidence

The propositions above are conceptual claims offered for test. Proposition 3 is the exception: Corollary 1 converts the theorem into a detectable geometry (an empty region above a ceiling line), examinable without field data on the full construct. This section reports what happened when we looked, in five steps: a simulation establishing that the ceiling geometry is statistically distinguishable from the compensatory specifications that predominate in capability research (Dul 2016; Richter et al. 2020); an underpowered-by-design feasibility pilot; a pre-registered, adequately powered necessary-condition study, frozen before collection, which decides against that operationalization (Section 6.3); exploratory analysis of what the null left behind, which motivates Proposition 4 (Section 6.4); and a second pre-registered study testing the new proposition’s tail prediction on fresh data (Section 6.5). Per Section 5.1’s scope paragraph, the repository analyses are single-coordinate partial probes: one sub-dimension, never the cross-layer aggregation, which requires the field program of Section 4.5.

The non-compensatory signature is recoverable

We simulated 300 organizations, with identical measurement error, under two data-generating processes: a compensatory weighted mean, and a non-compensatory minlsl+c\min_l s_l + c process consistent with Theorem 1. We follow standard Monte Carlo practice for stochastic operational models (Azarang Esfandiari and García Dunna 1996), and asked whether standard tools can tell them apart (full detail in Appendix D). They can. Under the compensatory process the profile mean predicts capability far better than the minimum (R2=0.81R^2=0.81 vs. 0.290.29); under the non-compensatory process the ordering reverses (0.790.79 vs. 0.330.33); AIC selects the aggregate matching the generating process in 100% of 200 replications; and the CE-FDH ceiling effect separates by 8×8\times (d=0.32d=0.32 vs. 0.040.04), a large effect on Dul (2016)’s benchmark. When capability is min-governed, high performance does not occur at low minimum maturity, whatever the other layers do (the empty upper-left region of Figure 5). That detectable geometry is why Proposition 3 is posed as a necessary-condition hypothesis, not an additive-regression one.

A real-data pilot: dependency concentration in public repositories

Dependency concentration (Section 4.2) is the one sub-dimension directly observable in public data. Following the truck-factor tradition (Avelino et al. 2016), a feasibility pilot computed commit-ownership concentration (DC\mathrm{DC}) and maintenance responsiveness (log10-\log_{10} of median issue-resolution latency) for 25 heavily starred repositories, under the theorem’s partial single-coordinate prediction. The pilot is inconclusive, and we report it as such: the ceiling points in the predicted direction (d=0.20d=0.20) but does not exceed its permutation null (p=0.32p=0.32) at n=25n=25, and the design carries defects a decisive test must repair (full report and figure in Appendix D). Its yield is methodological: DC\mathrm{DC} is computable on real artifacts and the necessary-condition test runs end to end. Its marginals calibrate the pre-collection power analysis of the study reported next, which fixes each defect by design.

A pre-registered necessary-condition study of the repository proxy

The decisive test was then run as a pre-registered study: sampling frame, measures, hypotheses, power analysis, and falsification criterion frozen before any outcome data were collected; the frozen plan, its version-control hash, and all code accompany this article, with the full design, screening ledger, sample, and robustness grid in Appendix D. In brief: four languages by five popularity strata (25 repositories per cell, seeded selection, content-list repositories excluded outcome-blind); exposure precedes fully observed outcome windows under a fixed 90-day horizon. The confirmatory hypothesis is the single-coordinate ceiling of Corollary 1: support requires both CE-FDH d0.10d \ge 0.10 and one-sided permutation p<0.05p < 0.05; the power analysis fixed the target at n=400n = 400; and failure at this size was registered in advance as a genuine unsupportive result, not an inconclusive one.

Result: the pre-registered criterion is not met (n=500n = 500; all twenty cells filled). The observed ceiling is large in magnitude (CE-FDH d=0.39d = 0.39, ceiling accuracy 0.970.97) but does not exceed its permutation null (p=0.071p = 0.071; Figure 7), which is wide because diffusion’s skewed marginal produces large chance ceilings, the reason the inference was registered against a permutation null and not fixed benchmarks. Two diagnostics undercut a necessity reading independently of the test: the bottleneck table finds no diffusion level required for any responsiveness level, and the ceiling is fragile to single ceiling-defining repositories (dd moves across [0.15,0.40][0.15, 0.40] under removals). The registered regression contrast sharpens the picture: with controls, the average association runs opposite to the necessity intuition (b=0.60b = -0.60, SE=0.16\mathrm{SE} = 0.16 for diffusion; more diffusely owned repositories resolve issues more slowly), and the robustness grid agrees with the primary cell throughout.

What this does and does not establish. Under the criterion registered before collection, the necessary-condition form of the diffusion–responsiveness link is not supported in this domain, and report it as a genuine unsupportive result for this operationalization (commit-ownership concentration as the Feedback-layer waste proxy, issue latency as the capability proxy, open-source repositories as the domain), not as an inconclusive one. The scope limits were registered in advance: a single-coordinate partial test (Section 5.1) that does not measure layer maturity or test LSImin\mathrm{LSI}_{\min}, with volunteer issue triage a distant proxy for organizational capability. The theory-level claim remains open where the construct lives, but the framework’s first powered probe returned a null, and any subsequent defense of Proposition 3 must reckon with it rather than around it. Three design lessons carry to the field program: power analyses for ceiling statistics must simulate from realistic skewed marginals; single-coordinate proxies invite the confound the negative slope displays; and outcome proxies should sit where triage policy cannot mimic or mask responsiveness.

What the null uncovered (exploratory)

Everything in this subsection is exploratory: computed after the confirmatory result was known and reported to motivate theory, not confirm it; these data are spent for Proposition 4, which is why the next subsection tests it on fresh data. Three patterns stand out (detail in Appendix D): the negative diffusion–responsiveness slope is not an artifact of tiny teams, holding within team-size strata and organization-owned repositories; responsiveness dispersion orders monotonically with concentration: the most concentrated tercile is fastest at the median and widest in spread, the most diffuse slowest and tightest, unchanged under issue-volume matching; and the standing level of concentration carries the signal, not its year-over-year change. A flow waste predicts a uniform penalty; a risk variable predicts this widening. Together these are the fingerprint of the contingent-liability reading of Section 4.2: concentrated ownership is fast in the undisrupted regime because it pays no coordination overhead, produces both the best and worst outcomes because nothing buffers it, and its true cost should be priced not by a cross-section but by what happens when the person leaves. That falsifiable statement is Proposition 4’s tail, registered before any new data were examined.

A second pre-registered test: the departure natural experiment

Proposition 4’s tail prediction was tested twice on fresh data under frozen analysis plans (supplementary PREREGISTRATION_DEPARTURE.md and PREREGISTRATION_DEPARTURE_V2.md; full designs and attempt-level reporting in Appendix D). In outline: a repository enters the risk set when one contributor dominates its baseline commits; a departure occurs when that contributor’s commits cease entirely; the registered inference is a one-sided test of a negative concentration coefficient on the pre-to-post responsiveness change among departure events, with a falsifiable verdict at 80 or more events, and difference-in-differences, cessation, and dispersion-replication analyses as registered secondaries. Both attempts fell short of the 80-event floor, so both return no verdict; the tail prediction is untested, not unsupported, and both are reported in full rather than iterated until a criterion clears.

The first attempt found the design assumption instructive: dominant-contributor departures with rich issue traffic are rare in surviving-repository frames (8 analyzable events among 1,474 screened). Its registered secondaries both pointed the proposition’s way: the dispersion signature replicated out of sample (347 disjoint repositories), and whole-project cessation struck 55%55\% of above-median-concentration departures against 9%9\% below. The pre-registered v2 enriched the harvest without conditioning on outcomes, doubling departures to 46 (30 analyzable events; achieved power 0.160.160.400.40). At the achieved sample the registered dose-response estimate is large and in the predicted direction (βDC=1.84\beta_{\mathrm{DC}} = -1.84, SE 0.690.69, one-sided p=.007p = .007), but an estimate that clears nominal significance at power 0.40.4 is the kind that inflates, and we decline to lean on it. Two registered secondaries cut in opposite directions. Against the proposition, the difference-in-differences interaction is null, and the theory explains the contamination: the dispersion signature implies concentrated repositories sit at extremes that regress toward the mean in any change-score design, departure or not. For the proposition, whole-project cessation (an outcome immune to mean reversion) struck 70%70\% of above-median departures against 35%35\% below (n=46n = 46), and the dispersion signature replicated a third time, on the 280 v2 repositories present in neither earlier sample, robust to issue-volume matching.

Where this leaves Proposition 4. Its cross-sectional half is a thrice-observed regularity of this domain awaiting construct-native confirmation. Its tail half has consistent directional evidence (the dose-response estimate, and a cessation gradient mean reversion cannot produce) but no certified verdict, a named confound (mean reversion in change scores, predicted by the proposition’s own dispersion claim), and a now-precise design prescription: survival and cessation outcomes rather than change scores, larger event harvests, or, best, organizational sites where departures are directly observable.

The evidential ledger

Because this section’s analyses concern one sub-dimension observed through one repository proxy, we close by stating what each of the framework’s claims can and cannot inherit (Table 10, Appendix D). Two rows deserve emphasis: the decided null is a verdict on the proxy and licenses no inference about the cross-layer claim, which no analysis here measures (LSImin\mathrm{LSI}_{\min} requires a five-layer profile, and none exists in these data); and no analysis administers the LSI or SWAT, so the instruments’ validity is untouched in either direction. Conflating the proxy’s null with a test of the cross-layer framework would misread both; the decisive tests are the organizational ones the ledger names.

Implications for Practice

The framework’s constructs imply a diagnostic and improvement sequence, offered as a reasoned application of the theory, not a validated program. Its practical force comes from the theorem: because capability is bounded at the profile minimum and Remark 1 identifies which layer binds, diagnosis becomes a computation, not a judgment call, and improvement anywhere but the binding layer yields no first-order return (Corollary 2). Concretely, an organization carrying three locally adequate but mutually divergent definitions of “customer” will read low on Data semantic drift; if Data is the profile minimum, polishing Interface or Orchestration (however visible the improvements) cannot raise capability until the Data constraint is relieved. Diagnosis maps where business logic resides, where operations depend on specific individuals, and how data flows; administers the SWAT and computes the LSI, as provisional readings, not normed scores (Section 4), and identifies the constraining layer. Improvement proceeds in the order the propositions dictate: relieve the binding layer first (Proposition 3); establish authoritative sources and, for entities that legitimately differ across contexts, governed mappings (Section 3.1); externalize business logic judiciously (a rule repository pursued dogmatically becomes a new silo); and only then automate the flows thus established, with pull-based processing and closed feedback loops (Proposition 2): automation last, so it amplifies order, not misalignment. The costs (refactoring, documentation, integration, change management) are real, with declining marginal returns (Section 8.4); the framework predicts directional benefits, and quantifying them is a task for the empirical program.

Discussion

Theoretical contributions

The framework’s first three contributions can be restated briefly: structural waste as a distinct category whose locus, the arrangement of components, neither process waste nor technical debt occupies; a relation-preserving, not terminological, TPS translation (Section 2.1); and instruments, with a validation path, that make the construct investigable, not merely named. The fourth requires more detail: the weakest-layer constraint is a theorem under explicit assumptions (Theorem 1), its non-compensatory signature is statistically separable from the additive default, and a pre-registered study decided its first single-coordinate probe against it (Section 6.3). We count the null itself as part of the contribution: it establishes that the construct’s central proposition can be operationalized, powered, and decided. What survives the null, what it generated (Sections 6.46.5: the concentration trade-off, replicated cross-sectionally, its tail twice tested and uncertified), and what would decide both propositions in organizational data are stated where they arise.

Relationship to existing theory

The framework connects to three established literatures. In information-processing theory (Galbraith 1974; Tushman and Nadler 1978), structural waste is unnecessary information-processing load created by architectural misalignment. In modular-systems theory (Baldwin and Clark 2000), the modal layers are a modularization optimized for operational clarity, not design evolvability, a distinct objective function. In socio-technical systems theory (Trist and Bamforth 1951), structural waste is a symptom of misalignment between technical architecture and the social system of work; its modern, empirically tested form is the mirroring hypothesis (Conway 1968; Colfer and Baldwin 2016), entering the framework as a boundary condition on Proposition 1: architectural separation endures only where ownership and communication boundaries move with it.

Organizational factors

The construct and its propositions are stated at the architecture level; firm-level conditions enter not as parts of the construct but as moderators of its architecture-to-outcome paths, a cross-level structure to be modeled as such (Klein et al. 1994). Three are named, each on a specific path: continuous-improvement culture moderates separation\rightarrowwaste (Proposition 1); executive support moderates whether separation is attempted at all, because architectural change spans boundaries no single unit controls, which makes structural waste strategic, not purely technical; and digital strategy moderates the antecedent itself, since integration-versus-best-of-breed choices determine how many boundaries need reconciling (Bharadwaj et al. 2013). These are hypothesized moderators, not results. Distinct from them, managerial and IT capability, slack, and organizational size and age are potential confounders and enter the empirical program as measured controls, not as theory.

Limitations and risks

The framework has limitations. Measurement is the most immediate: coordination overhead and semantic drift are not routinely instrumented, so the instruments of Section 4 remain proposals pending validation. Implementation cost is real: layer separation is substantial architectural work whose return is not guaranteed. Change resistance is predictable, because separating layers exposes hidden dependencies and informal workarounds stakeholders may prefer undisturbed. And there is an over-engineering risk: separation pursued dogmatically can add complexity beyond the waste it removes, which is why Proposition 1 is bounded and the framework calls for judicious, not maximal, separation.

Future research

The most pressing line of inquiry is empirical validation: executing the program of Section 4.5. The pre-registered study has already decided the cheapest version of the weakest-layer test, so the next step is not another repository study but the harder, construct-native one: organizational sites with measured five-layer maturity profiles, where LSImin\mathrm{LSI}_{\min} exists, the cross-layer min-structure is testable (Corollary 1), the binding-layer signature of Corollary 2 can be examined in panel form, and the design lessons of Section 6.3 are built in. Its unit of analysis is the operational flow path, not the firm: several instrumented flows per site give within-site replication and the panel structure the corollaries need. Its registered outcomes include outcome dispersion and post-departure recovery, the quantities Proposition 4 makes load-bearing and change scores cannot carry. Beyond validation: contingency research on how industry, regulation, size, and culture moderate the separation–performance relationship; dynamic-capabilities research on whether separation’s architectural flexibility contributes to strategic agility; and technology-evolution research on whether artificial-intelligence and low-code platforms relieve architectural misalignment or, by lowering the cost of adding systems, accelerate it.

Conclusion

This paper developed a conceptual theory of structural waste in digital operations, extending Lean production from the shop floor to the architecture of organizational information systems, and answers the four questions posed at the outset. To RQ1, the five-layer modal architecture, in which each layer carries a core element of Lean: the Data Layer standardized work, the Logic Layer jidoka, the Interface Layer visual management, the Orchestration Layer flow and pull, the Feedback Layer the kaizen pillar itself. The translation is structural rather than terminological. To RQ2, structural waste: architectural misalignment producing coordination overhead, semantic drift, and dependency concentration, arising, unlike technical debt or process waste, from how systems are organized relative to one another. To RQ3, the Layer Separation Index and the Structural Waste Assessment Tool, proposed with an explicit formative/reflective specification and a full validation program; establishing their psychometric properties is the task for subsequent research.

To RQ4, the question the title poses, the paper answers with the claim’s precise form and its record: capability is bounded by the least mature layer within an estimable compensation band as a theorem under stated assumptions; the signature is statistically detectable; and the claim’s most accessible operationalization, put to a pre-registered test, was decided against it. Exploring that null uncovered the concentration trade-off of Proposition 4: concentration as the knowledge analogue of just-in-time, fast and fragile. Its cross-sectional signature has been observed in three samples. The framework is offered together with its null and its two uncertified tail tests, and the construct-native tests are now precisely specified.

The practical implication: organizations that invest heavily in digital transformation frequently accelerate disorder rather than eliminate it, because they automate misalignment instead of first establishing architectural order. Just as TPS required a rethinking of manufacturing rather than the adoption of new machines, operational excellence in digital environments requires a rethinking of information architecture, not the adoption of new tools.

Disclosure statement

The first author is the founder of Mentu (the research and software practice of Esfandiari S.A. de C.V.), whose operational work motivated, but did not fund or constrain, the theory developed here. The authors report no other competing interests.

Funding

No funding was obtained for the reported work.

Data availability statement

The data and code supporting Section 6 are provided as supplementary material: the pilot dataset; the frozen pre-registered analysis plans with their pre-collection power analyses; collection code, screening and exclusion logs, and pseudonymized repository-level datasets (contributor identifiers salted-hashed; per-endpoint retrieval timestamps included, as platform state is not replayable); and the seeded analysis code that reproduces every statistic and figure, including the theorem’s machine verification. No human-subject data were generated or analysed.

Declaration of generative AI use

Generative AI tools were used, under full author direction, for literature retrieval, reference verification, and manuscript preparation; the authors reviewed and take complete responsibility for all content.

Appendix A: Layer Separation Indicators

Each modal layer is assessed on five indicators, each scored on a 0011 scale at the anchor points {0,0.25,0.5,0.75,1}\{0,\,0.25,\,0.5,\,0.75,\,1\}. The separation score for a layer, sls_l, is the mean of its five indicator scores; higher scores denote cleaner separation and greater architectural clarity. The rubrics below give the anchor descriptions. As set out in Section 4.5, the rubric is a proposed content specification pending empirical validation, not a normed instrument.

Data Layer

Data Layer separation indicators and scoring anchors.
Dimension 0.0 0.25 0.5 0.75 1.0
Single Source of Truth Multiple conflicting data sources for same entities Some entities have designated primary sources Most entities have single sources with exceptions Clear primary sources with documented exceptions Every entity has one authoritative source
Data Governance Model No formal data governance Ad-hoc ownership assignments Documented ownership for critical data Comprehensive governance with stewardship Full governance with quality metrics and accountability
Schema Version Control No version control for data structures Manual documentation of changes Some schemas under version control Most schemas versioned with migration paths All schemas versioned with automated migration
API-Based Data Access Direct database access predominant Some data accessed through APIs Critical data has API access Most data accessed via APIs All data access through versioned APIs
Data Quality Monitoring No systematic quality monitoring Manual spot checks Automated monitoring for critical data Comprehensive monitoring with alerts Quality checks defined for every authoritative source, with owners and documented remediation paths

Logic Layer

Logic Layer separation indicators and scoring anchors.
Dimension 0.0 0.25 0.5 0.75 1.0
Business Rules Externalization All logic embedded in application code Some rules documented separately Critical rules in configuration files Most rules in rule engines or services All business logic externalized and versioned
Centralized Logic Repository Logic scattered across systems Some consolidation efforts Shared libraries for common logic Central repository with governance Single source for all business logic
Logic Testability Logic only testable through UI Some unit tests for logic Logic partially testable independently Most logic has independent tests All logic testable without system dependencies
Rule Documentation No documentation of business rules Informal documentation exists Key rules documented Comprehensive documentation with examples Complete documentation with decision trees
Logic Version Control No version control for business rules Manual tracking of changes Some rules under version control Most rules versioned with history All logic versioned with rollback capability

Interface Layer

Interface Layer separation indicators and scoring anchors.
Dimension 0.0 0.25 0.5 0.75 1.0
UI–Logic Independence Business logic embedded in interfaces Some separation attempts Critical logic separated from UI Clear separation with exceptions Complete independence of UI and logic
Multiple Interface Options Single interface per function Limited interface flexibility Some functions have multiple interfaces Most functions accessible multiple ways All functions have API, UI, and CLI options
API-First Design Interfaces built without API consideration APIs added after interface development Some API-first development API-first for new development All interfaces built on underlying APIs
Interface Versioning No interface version management Manual version tracking Some interfaces versioned Systematic versioning with deprecation Full version lifecycle management
Role-Based Access Control No systematic access control Basic authentication only Some role-based controls Comprehensive RBAC implementation Access defined entirely through declared role contracts; no individually granted exceptions

Orchestration Layer

Orchestration Layer separation indicators and scoring anchors.
Dimension 0.0 0.25 0.5 0.75 1.0
Workflow Engine Usage All orchestration hard-coded Some scripted workflows Workflow engine for critical processes Most processes in workflow engine All orchestration through workflow platform
Demand-Triggered Sequencing All work pushed on schedules or manual initiation Some steps triggered by downstream demand Mixed push and demand-triggered sequencing Most work triggered by downstream demand All work triggered by declared demand signals; no hidden schedule- or person-dependent triggers
Process Visibility No visibility into process state Manual status checking Some process monitoring Real-time process dashboards Complete process observability
Orchestration Independence Orchestration tied to applications Some separation efforts Critical processes independent Most orchestration separated Complete orchestration independence
SLA Monitoring No SLA tracking Manual SLA reporting Automated SLA monitoring Real-time SLA dashboards Service expectations declared as explicit inter-layer contracts and monitored against commitments

Feedback Layer

Feedback Layer separation indicators and scoring anchors.
Dimension 0.0 0.25 0.5 0.75 1.0
Automated Monitoring Coverage No automated monitoring Basic system monitoring Key metrics monitored Comprehensive monitoring All critical flows monitored through declared indicators with documented response paths
Feedback Loop Documentation No documented feedback processes Informal feedback practices Some loops documented Most feedback loops documented All loops documented with triggers
Performance Metrics Definition No defined metrics Basic operational metrics KPIs defined for critical processes Comprehensive metrics framework Balanced scorecard with leading indicators
Automated Alerting No automated alerts Basic threshold alerts Smart alerting for critical issues Comprehensive alert framework Alerting defined by documented triggers with named owners, independent of any individual’s vigilance
Improvement Process Maturity No systematic improvement Ad-hoc problem solving Regular review cycles Continuous improvement culture Improvement loop runs on documented triggers and owners; outcomes return as recorded change

Appendix B: Structural Waste Assessment Tool (SWAT)

The Structural Waste Assessment Tool is a proposed 25-item self-report instrument, five items per modal layer. Respondents rate each statement on a five-point scale: 1=1= never true, 2=2= rarely true, 3=3= sometimes true, 4=4= often true, 5=5= always true. As specified in Section 4.4, the items within a layer are intended as reflective indicators of that layer’s latent waste, and the instrument is proposed pending the validation program of Section 4.5.

SWAT items, grouped by modal layer.
Data Layer waste (items 1–5)
1 Different systems show different values for the same metric.
2 Two systems disagree about the meaning of a shared field (for example, what counts as an active customer).
3 We must reconcile the same data across systems that store it differently.
4 Reports require manual verification before they can be used.
5 Data definitions vary across departments.
6 Business rules exist in multiple places across our systems.
7 The same calculation produces different results in different systems.
8 We have to update business logic in multiple locations.
9 Business rules cannot be determined without reading application code.
10 Changing business rules requires modifying user-interface code.
11 We have multiple dashboards showing the same information differently.
12 Completing a routine task requires switching between multiple screens or systems.
13 We produce reports and dashboards that have no known consumer.
14 The same function requires different steps in different interfaces.
15 Interfaces that have been replaced remain in production.
16 Processes fail silently without notification.
17 We cannot track where work items are in our processes.
18 Work stalls between systems until someone moves it along manually.
19 A change to one process requires coordination across multiple teams before it can go live.
20 The same workflow is implemented differently across systems.
21 We measure things that don’t drive decisions.
22 Lessons from past incidents live with specific people rather than in the system.
23 Key processes stop working when specific people are unavailable.
24 Improvements are not measured or tracked.
25 When experienced staff leave, their operational knowledge leaves with them.

Scoring.

Each layer subscale is the sum of its five items (range 552525), and the instrument’s output is the per-layer profile. We deliberately define no single total score: by Proposition 3, waste is not compensatory across layers, so a sum over layers would misrepresent the construct, and the profile identifies the layer most in need of attention. Because the instrument is not yet validated, we do not supply population norms or interpretive cut-points: the mapping from a raw score to a substantive judgment (“low” versus “high” structural waste) requires a validated administration to a defined population and is an output of the validation program in Section 4.5, not an input to it.

Appendix C: Operational definitions and formal supplement

Discriminant, mapping, and translation tables

Translation of the seven wastes (muda) to their digital equivalents. The eighth waste (unused talent) is included as it maps onto the dependency-concentration sub-dimension of structural waste.
Manufacturing waste Digital equivalent Primary layer Mechanism (why it is waste)
Overproduction Reports, extracts, and data nobody consumes Interface Producing information ahead of or beyond demand; the digital analogue of the most fundamental muda, and the one that hides the others. Generated by push sequencing at the Orchestration layer; its visible product accumulates at the Interface, which is where it is observed and measured.
Waiting Idle work items awaiting manual hand-off or approval Orchestration Value-carrying items stall between systems because sequencing is push-based rather than pull-based.
Transport Moving data between systems to reconcile them Data Effort spent relocating information across boundaries that exist only because no authoritative source does. Distinct from coordination overhead proper: transport is the movement the reconciliation requires, not the reconciling judgment itself.
Over-processing Re-entering, re-formatting, and re-validating the same data Logic Work performed because a rule or format is duplicated rather than externalized once. Distinct from motion (operator navigation) and from defects (propagated error): over-processing is redundant processing of correct data.
Inventory Accumulated work-in-process: queued items, partially processed batches, unhandled backlogs Orchestration Work that has entered the system but not yet delivered value accumulates between stages, hiding problems and delaying their discovery. Only this half of inventory’s causal role transfers; see the boundary note in Section 3.3.
Motion Navigating multiple screens and tools to complete one task Interface Operator effort expended because the point-of-use does not present what the decision requires.
Defects Inconsistent values and errors that propagate across systems Logic Errors propagate because validity is not enforced at the source of the decision, the jidoka function the Logic layer carries. Recurrence without learning is the corresponding Feedback-layer failure.
(8) Unused talent Processes held in individual memory rather than the system Feedback Capability concentrated in people who “know how things work”; contained at the Feedback layer, where closing the improvement loop externalizes know-how into the architecture, freeing people from serving as its memory so they can serve as its source of improvement. This is the dependency-concentration sub-dimension of structural waste.
Structural waste distinguished from technical debt and process waste.
Technical debt Process waste Structural waste
Locus Implementation (code and within-system design) Execution (workflow) Architecture (arrangement of systems)
Nature Liability of past expedient decisions Execution inefficiency Present, recurring operational friction
Cause Expedient implementation Workflow friction System fragmentation
Signature Rising change cost, defect risk Delays, queues, rework Reconciliation effort, definitional conflict, single-person dependencies
Consequence Future rework Operational delay Degraded operational capability
Remedy Refactoring Flow / value-stream redesign Layer separation
Example Hard-coded value instead of configuration Manual approval bottleneck Same customer defined three ways in three systems
The two-level structure of structural waste: three sub-dimensions of the construct (Section 3.1) related to the five modal layers at which they become observable. Each layer exhibits one characteristic symptom, which rolls up to the sub-dimension named in the right-hand column.
Modal layer TPS principle Layer-level symptom Sub-dimension
Data Standardized work Divergent authoritative meaning Semantic drift
Logic Jidoka Duplicated business rules Coordination overhead
Interface Visual management Redundant / unused interfaces Coordination overhead
Orchestration Flow and pull Push-based sequencing, hand-off stalls Coordination overhead
Feedback Kaizen (pillar) Know-how held in individuals Dependency concentration

The four direct quantities

Person-hours spent maintaining cross-system consistency, expressed as a share of operational effort: CC=ihiH[0,1],\mathrm{CC} = \frac{\sum_{i} h_i}{H}\in[0,1], where hih_i is weekly reconciliation hours for activity ii and HH is total operational labor hours over the same period and scope. The share form makes the quantity unit-free and comparable across settings; normalizing by “units of work delivered” is deliberately avoided, because such units are not commensurable across organizations. The boundary between reconciliation and value-adding work is a coding judgment governed by the protocol requirement of Section 4.5.

A structural antecedent of coordination overhead, located at the Logic layer: duplication is what forces subsequent reconciliation, so it measures the cause whose effort coordination cost measures directly. Operationally, it is the mean number of locations at which each business rule is implemented, LD=rrR,LDnorm=LD1N1[0,1],\mathrm{LD} = \frac{\sum_{r} \ell_r}{R}, \qquad \mathrm{LD}_{\text{norm}} = \frac{\mathrm{LD}-1}{\,N-1\,}\in[0,1], where r\ell_r is the number of locations implementing rule rr, RR is the number of rules, and NN is the number of systems in the declared measurement scope, the maximum number of locations at which a rule could be implemented. Normalizing by a property of the scope rather than by the maximum observed duplication keeps the quantity invariant across organizations and samples; LD=1\mathrm{LD}=1 denotes a single authoritative location per rule. The measure is the operational analogue of code-clone measurement in software engineering (Roy et al. 2009), with the business rule across heterogeneous systems, rather than the code fragment within a codebase, as the unit.

The share of shared entity fields whose definitions diverge across systems without a declared mapping, SD=|{fields with ungoverned divergent definitions}||{shared fields}|[0,1].\mathrm{SD} = \frac{\lvert\{\text{fields with ungoverned divergent definitions}\}\rvert}{\lvert\{\text{shared fields}\}\rvert}\in[0,1]. Fields are shared when two or more systems assert values for the same entity attribute, and divergent when their definitions differ with no explicit mapping recording the difference (Section 3.1); both are coding judgments governed by the protocol requirement of Section 4.5. The measure adapts the consistency dimension of data quality (Wang and Strong 1996) and the correspondence problem of schema matching (Rahm and Bernstein 2001) to the construct’s governance criterion.

A Herfindahl–Hirschman concentration of person-dependent process steps over the individuals they depend on, DC=isi2,si=dependence attributed to person i|person-dependent steps|,\mathrm{DC} = \sum_{i} s_i^{2}, \qquad s_i = \frac{\text{dependence attributed to person } i}{\lvert\text{person-dependent steps}\rvert}, where a step is person-dependent when its execution halts in the absence of its named knowledge-holders (a step two people can execute is scored system-embedded only if it survives the removal of all of them), and each person-dependent step allocates one unit of dependence, fractionally across the holders whose joint absence halts it. Steps executable from system-embedded procedure alone contribute nothing. Under this attribution isi=1\sum_i s_i = 1 whenever any person-dependent step exists, so DC(0,1]\mathrm{DC}\in(0,1] is a pure concentration: it attains 11 only when all person-dependence rests on one individual, and it is undefined (reported as no-dependence) when coverage is zero. Because concentration alone cannot say how much of the operation relies on individuals at all, DC is always reported with its companion coverage statistic, DCov=|person-dependent steps|/|total steps|\mathrm{DCov} = \lvert\text{person-dependent steps}\rvert / \lvert\text{total steps}\rvert: coverage captures the extent of the retention failure of Section 3.1, DC its concentration (single-point-of-failure risk), and the two are orthogonal by construction: two-person redundancy lowers DC without being mistaken for system-embedded capability, and dispersed dependence shows as high coverage with low concentration. The pair adapts the “truck factor” of empirical software engineering (Avelino et al. 2016) from code ownership to operational process steps.

The scale-development and validation program for the proposed instruments, following Churchill (1979) and MacKenzie et al. (2011). Conceptual stages (blue) are begun in this paper; empirical stages (orange) are specified here as the primary task for subsequent research.

Proof of Theorem 1

Proof. Lower bound (Assumption 3 with q0=1q_0 = 1). By induction, qlminjlsjq_l \ge \min_{j \le l} s_j: first q1min(s1,q0)=s1q_1 \ge \min(s_1, q_0) = s_1; then qlmin(sl,ql1)min(sl,minj<lsj)=minjlsjq_l \ge \min(s_l, q_{l-1}) \ge \min\bigl(s_l, \min_{j<l} s_j\bigr) = \min_{j \le l} s_j. Hence Y=qLminlslY = q_L \ge \min_l s_l. Upper bound (Assumptions 12). Let m=argminlslm = \arg\min_l s_l. By Assumption 1, qmsm=minlslq_m \le s_m = \min_l s_l; summing Assumption 2 from mm to LL yields Yqm+l>mκlminlsl+KY \le q_m + \sum_{l>m}\kappa_l \le \min_l s_l + K. Attainability. With κ0\kappa \equiv 0 the assumptions force ql=min(sl,ql1)q_l = \min(s_l, q_{l-1}), so Y=minlslY = \min_l s_l for any profile. For the upper bound, place the weakest layer at position mm, set κi=0\kappa_i = 0 for imi \le m with the entire budget downstream, and take sjminlsl+Ks_j \ge \min_l s_l + K for all j>mj > m with minlsl+K1\min_l s_l + K \le 1; then Y=minlsl+KY = \min_l s_l + K. ◻

Attainment law and the second signature

Remark 1 (Attainment law, closed form, and the binding layer). The deterministic rule ql=min(sl,ql1+κl)q_l = \min(s_l,\, q_{l-1} + \kappa_l), under which each layer delivers the best quality its ceiling and budget permit, satisfies all three assumptions and yields the closed form Y=min1jL(sj+l>jκl),Y \;=\; \min_{1 \le j \le L}\Bigl(s_j + \textstyle\sum_{l>j}\kappa_l\Bigr), whose minimizer identifies which layer binds: the layer whose maturity plus downstream rescue budget is lowest, reducing to argminlsl\arg\min_l s_l when K=0K = 0. The closed form turns the framework’s diagnostic promise into an explicit computation and supplies the structural signal for the simulation of Section 6.1.

Corollary 2 (Binding-layer marginal returns). Under the attainment law of Remark 1, improvements to layers other than the binding layer yield no capability gain until the constraint shifts: Y/sl>0\partial Y / \partial s_l > 0 only at the minimizing layer j*j^{\ast} of the closed form. This is a second, independently testable signature: in panel data, maturity investments should raise capability only when they land on the binding layer, and sequential improvement programs should exhibit the stepwise gains characteristic of constraint relief rather than the smooth gains an additive model predicts.

Index-set note.

The theorem’s minimum ranges over the ordered path layers; the field quantity LSImin\mathrm{LSI}_{\min} (Section 4.3) ranges over all five. A minimum over the larger set cannot exceed the path minimum, so a ceiling test run on LSImin\mathrm{LSI}_{\min} can only be harder to pass than the theorem requires; an apparent violation may be a legitimate observation under the theorem. The five-layer minimum is therefore the demanding operationalization, and the per-layer tests of Section 5.1 are the exact ones.

Appendix D: Full reports of the simulation, pilot, and pre-registered studies

This appendix carries the complete reports summarized in Section 6: design, sample, inference, and robustness detail for the simulation, the feasibility pilot, the pre-registered necessary-condition study, the exploratory analysis it left behind, and both pre-registered departure studies. Every number is generated by the seeded code in the supplementary material, which also contains the frozen analysis plans and their version-control hashes.

The simulation in full

We simulated N=300N=300 organizations, each with five layer-maturity scores sls_l drawn uniformly on [0.15,0.98][0.15,0.98], and generated realized capability under two data-generating processes: a compensatory process, Y*=lwlslY^{\ast}=\sum_l w_l s_l with equal weights (the weighted mean a conventional additive model assumes), and a non-compensatory process consistent with Theorem 1, Y*=minlsl+cY^{\ast}=\min_l s_l + c with compensation cc drawn uniformly on [0,K][0, K], K=0.10K=0.10. In both cases the analyst observes Y=Y*+εY = Y^{\ast} + \varepsilon with ε𝒩(0,0.05)\varepsilon \sim \mathcal{N}(0, 0.05) interpreted as measurement error on the outcome, so the structural signal itself respects the theorem’s band. All parameters and the random seed (20260708) are fixed in the supplementary code. We then asked which aggregate, the profile mean or the profile minimum, better explains observed capability under each process, and whether necessary-condition analysis recovers the ceiling the theorem predicts.

Formal model comparison separates the processes cleanly: the change in AIC from the mean model to the minimum model is +394+394 under the compensatory process (the mean model is preferred) and 341-341 under the non-compensatory process (the minimum model is preferred), and across 200 independent replications the aggregate matching the data-generating process attains the lower AIC in 100% of replications under both processes. The necessary-condition signature behaves as expected: the CE-FDH ceiling effect size is d=0.04d=0.04 under the compensatory process (effectively absent) and d=0.32d=0.32 under the non-compensatory process (Figure 5).

The non-compensatory signature is statistically distinguishable from a linear-additive model. Simulated capability against the profile minimum minlsl\min_l s_l under the compensatory (left) and non-compensatory (center) data-generating processes, with CE-FDH ceilings: only the non-compensatory process leaves the upper-left region empty (high capability requires a high minimum; d=0.32d = 0.32 vs. 0.040.04). (Right) OLS fit of the mean and minimum aggregates under each process: the better-fitting aggregate identifies the generating process, and AIC selects the correct aggregate in 100% of 200 replications. Generated by the seeded supplementary code.

The feasibility pilot in full

For a convenience sample of 25 heavily starred open-source repositories (86,000–391,000 stars; 7 Python, 10 JavaScript, 8 Go by platform language tag), we computed the Herfindahl index of commit ownership, the same DC\mathrm{DC} quantity the construct defines, here over each repository’s top 100 contributors, the maximum the unauthenticated platform interface returns, which truncates the ownership tail, and, as an outcome, the median resolution latency of recently closed issues, read as software-maintenance responsiveness. A conventional linear model finds essentially no average relationship between ownership diffusion and responsiveness (r=0.08r=-0.08, R2=0.007R^2=0.007, responsiveness as log10-\log_{10} of median latency in days). NCA returns a ceiling in the predicted direction (CE-FDH d=0.20d=0.20 for diffusion; d=0.04d=0.04 and 0.150.15 for the truck-factor and top-share variants), but none exceeds its one-sided permutation null over 10,00010{,}000 permutations (p=0.32p=0.32; Figure 6): at n=25n=25 the null’s own 9595th percentile (0.300.30) lies above the observed effect. Beyond low power, the design weaknesses the pre-registered study repairs are: an extreme-popularity frame; documentation and curated-list repositories whose “issues” are content submissions; per-repository medians resting on as few as 3 closed issues; and a concentration index conflating all-time contribution with current ownership.

Feasibility pilot: dependency concentration on 25 public repositories. (a) Maintenance responsiveness (log10-\log_{10} of median issue-resolution latency) against ownership diffusion (1DC1-\mathrm{DC}), with the empirical CE-FDH ceiling (d=0.20d=0.20). (b) The observed effect against its permutation null (10,00010{,}000 permutations): the observed dd falls below the null’s 9595th percentile, so the pilot detects no significant ceiling. The point estimate is in the direction the weakest-layer proposition predicts. Generated by the seeded supplementary code.

The pre-registered necessary-condition study in full

Design. The frame stratifies four languages (Python, JavaScript, Go, Rust) by five popularity strata (500 to 150,000 stars), 25 repositories per cell, drawn by seeded random selection from screened candidates with pre-committed reserve ordering, with no discretionary replacement. Screening excludes archived repositories, forks, templates, mirrors, and (before any outcome data are seen) documentation and curated-list repositories, by declared name, description, and topic rules. Exposure precedes outcome: dependency concentration DCw\mathrm{DC}_w is the Herfindahl index of non-bot commit-author shares over the window 24 to 12 months before collection, with full contributor pagination (no top-100 truncation); the outcome window covers issues created 15 to 3 months before collection, each followed for a fixed 90-day resolution horizon, so every issue is fully observed and multi-year stale artifacts cannot enter. Maintenance responsiveness is log10-\log_{10} of the median latency; bot-authored issues and stale-bot closures are excluded; eligibility requires at least 50 non-bot commits and 20 eligible issues. Median merge latency of pull requests serves as a secondary outcome, and repository age, popularity, contributor count, ownership type, and language enter as controls in the regression contrast.

Pre-registered inference. Support requires both a CE-FDH effect d0.10d \ge 0.10 and a one-sided permutation p<0.05p < 0.05 over 10,00010{,}000 permutations. The pre-collection power analysis, calibrated to the pilot’s marginals, fixed the target at n=400n = 400 analyzed repositories with a false-positive rate below 3%. A robustness grid (alternative concentration measures, the secondary outcome, contemporaneous windows, and within-stratum analyses) is reported descriptively; the primary cell alone decides the hypothesis.

Sample. The frame returned 4,754 candidate repositories, of which 371 were screened out before any outcome data were seen (231 by the content-list name/description rules, 36 by topic, 84 lacking issue trackers, 20 templates). Walking the seeded orderings, 674 candidates failed eligibility (507 with fewer than 50 non-bot commits in the exposure window, 166 with fewer than 20 eligible issues, one inaccessible), yielding the target n=500n = 500, with all twenty cells filled at 25. Median issue latency was 34 days; 158 repositories had median latency at the 90-day horizon (the responsiveness floor); ownership diffusion spanned 00 to 0.980.98 (median 0.630.63); 366 repositories were organization-owned.

Result detail. The observed ceiling (CE-FDH d=0.39d = 0.39, ceiling accuracy 0.970.97) fails its permutation null at p=0.071p = 0.071 against the registered α=0.05\alpha = 0.05, with the null’s own 95th percentile at d=0.40d = 0.40 (Figure 7). The bottleneck table finds no diffusion level required for any responsiveness level: the 90th percentile of responsiveness is attained in the sample at diffusion near zero, so the empty-corner pattern is not present in the strong form. Removing single ceiling-defining repositories moves dd across [0.15,0.40][0.15, 0.40]. The regression contrast yields b=0.60b = -0.60, SE=0.16\mathrm{SE} = 0.16 for diffusion with controls (R2=0.062R^2 = 0.062). In the robustness grid, the secondary outcome (pull-request merge latency) shows no ceiling under any concentration measure (d0.02d \le 0.02 throughout); no language, popularity, or ownership stratum exceeds its null; and the single nominally significant cell among nine (p=0.038p = 0.038 for the top-share variant) is what multiplicity alone would produce and licenses nothing under the registered decision rule.

The pre-registered necessary-condition study (n=500n = 500). (Left) Maintenance responsiveness against ownership diffusion (1DCw1-\mathrm{DC}_w) with the CE-FDH ceiling: the apparent ceiling zone is large (d=0.39d = 0.39) but rests on few ceiling-defining points; the horizontal band at the bottom is the 90-day resolution horizon. (Center) The observed effect against its permutation null (10,00010{,}000 permutations): p=0.071p = 0.071, below the registered criterion; the null is wide because diffusion’s skewed marginal produces large chance ceilings. (Right) Robustness grid: no ceiling for the pull-request outcome under any concentration measure; the one nominally significant cell of nine is consistent with multiplicity. Generated by the seeded supplementary code.

The exploratory analysis in full

First, the negative average association between diffusion and responsiveness is not an artifact of tiny teams: it holds within every mid-size team stratum (r0.20r \approx -0.20 for 3–10 and 11–50 baseline authors) and within organization-owned repositories (supplementary exploratory_analysis.py). Second, responsiveness dispersion orders monotonically with concentration: the most concentrated tercile has the fastest median responsiveness and the widest spread (median 1.33-1.33, IQR 1.251.25), the most diffuse tercile the slowest median and the tightest spread (1.65-1.65, IQR 0.870.87). The widening is not a small-sample artifact of concentrated repositories having fewer issues: the ordering is unchanged when terciles are compared within issue-volume-matched bands (supplementary audit_p4a.py). Third, the standing level of concentration carries the signal; its year-over-year change is uncorrelated with responsiveness.

The departure studies in full

Design (v1). The frame mirrors the necessary-condition study’s twenty cells but drops the activity-recency screen (survivorship after a departure is an outcome here, not a nuisance) and requires repositories old enough to observe a full event window. A repository enters the risk set when one contributor authored at least 30% of its non-bot commits in the twelve-month baseline window; a departure occurs when that contributor’s commits cease entirely for the following eighteen months. The outcome is the change in maintenance responsiveness from a six-month pre-window to a six-month post-window (issues created in each, fixed 90-day horizons, fully observed); repositories that die after a departure enter at the responsiveness floor under a rule registered in advance. Registered inference: among departure events, ordinary least squares of the responsiveness change on baseline concentration with controls (popularity, age, team size, ownership, language); one-sided test of a negative concentration coefficient at α=0.05\alpha = 0.05, falsifiable verdict at 80 or more events; difference-in-differences against still-active dominant-contributor repositories, a cessation analysis, and a cross-sectional replication of the dispersion signature as registered secondaries.

First attempt. Among 1,474 screened repositories, 551 had a dominant contributor but only 22 experienced a departure (the design assumed a prevalence several times higher), and 8 events survived the outcome-volume floors, below the 80-event threshold, so the frozen rule returns no verdict. The dispersion replication was computed on the 347 dominant repositories not shared with the earlier study’s sample (the overlapping 22% of the 444 eligible were excluded): the most concentrated tercile is again fastest at the median and widest in spread (median 1.08-1.08, IQR 1.571.57), the most diffuse tercile slowest and tightest (1.79-1.79, IQR 0.740.74), monotone in both statistics, unchanged within issue-volume-matched bands, and computed on repositories that played no role in generating the proposition. Among the 22 departures, whole-project cessation (no commits by anyone in the post-window) struck 55%55\% of above-median-concentration repositories against 9%9\% of below-median ones.

Design (v2). Because the primary estimand conditions on a departure occurring, the event harvest can be enriched without biasing the within-event gradient, provided enrichment does not select on the outcome side; conditioning only on repository dormancy would do exactly that (diffusely owned survivors keep committing and would be excluded). The v2 frame therefore screens two staleness strata, dormant-since-baseline and currently active, with a stratum covariate in the registered model, relaxes the dominance and outcome-volume thresholds, and shortens the cessation window, all frozen in a second plan before any v2 data were collected; the hypothesis, test, significance level, and 80-event verdict rule are unchanged.

Second attempt. The enrichment doubled the harvest (1,405 screened, 538 dominant, 46 departures, 30 analyzable events), but the 80-event floor was again unmet (achieved power 0.160.160.400.40 for the calibrated effect sizes), so the frozen rule again returns no verdict. The registered model yields βDC=1.84\beta_{\mathrm{DC}} = -1.84 (SE 0.690.69, one-sided p=.007p = .007, R2=0.42R^2 = 0.42) at the achieved sample. The difference-in-differences interaction is null (departure ×\times concentration =0.03= -0.03, t=0.07t = -0.07): concentration predicted pre-to-post decline among still-active dominant repositories too, the mean-reversion contamination the main text names. Whole-project cessation struck 70%70\% of above-median-concentration departures against 35%35\% of below-median ones (n=46n = 46), and the dispersion signature replicated on the 280 v2 repositories present in neither earlier sample (concentrated tercile: median 1.43-1.43, IQR 1.471.47; diffuse: 1.95-1.95, IQR 0.700.70; monotone in both, unchanged under issue-volume matching).

The evidential ledger table

What this article’s evidence does and does not touch. “Decided” and “replicated” statuses are pre-registered; “untested” claims await the designs named in the right column.
Claim Standing after this article What would decide it
Theorem 1 + corollaries Proved under A1–A3; machine-verified — (analytic); K̂\hat{K} estimable in field NCA
P1 (layer separation) Untested; erosion boundary stated Field profiles; separation trajectories sl(t)s_l(t)
P2 (flow before automation) Untested (sequence claim) Coded event pairs in longitudinal field data
P3, cross-layer (LSImin\mathrm{LSI}_{\min}) Untested; no layer profiles exist in these data Organizational five-layer profiles; reliability-corrected min; NCA per Corollary 1
P3, repository proxy Decided against (pre-registered) — (decided, for this proxy)
P4a (trade-off, cross-section) Replicated ×3\times 3 (disjoint samples; volume-matched) in open source Enterprise replication with measured DCov\mathrm{DCov}
P4b (tail dose-response) Uncertified ×2\times 2 (<80<80 events); direction-consistent Survival and recovery outcomes at sites with observable departures
LSI and SWAT validity Not administered anywhere; untouched either way The program of Section 4.5 (Phase-0 panel materials exist)

References

Antonakis, John, Samuel Bendahan, Philippe Jacquart, and Rafael Lalive. 2010. “On Making Causal Claims: A Review and Recommendations.” The Leadership Quarterly 21 (6): 1086–120. https://doi.org/10.1016/j.leaqua.2010.10.010.
Avelino, Guilherme, Leonardo Passos, Andre Hora, and Marco Tulio Valente. 2016. “A Novel Approach for Estimating Truck Factors.” Proceedings of the IEEE 24th International Conference on Program Comprehension (ICPC), 1–10. https://doi.org/10.1109/ICPC.2016.7503718.
Azarang Esfandiari, Mohammad Reza, and Eduardo García Dunna. 1996. Simulación y Análisis de Modelos Estocásticos. McGraw-Hill Interamericana.
Baldwin, Carliss Y., and Kim B. Clark. 2000. Design Rules: The Power of Modularity. MIT Press.
Barabási, Albert-László. 2016. Network Science. Cambridge University Press.
Bell, Steven C., and Michael A. Orzen. 2010. Lean IT: Enabling and Sustaining Your Lean Transformation. CRC Press.
Besker, Terese, Antonio Martini, and Jan Bosch. 2018. “Managing Architectural Technical Debt: A Unified Model and Systematic Literature Review.” Journal of Systems and Software 135: 1–16. https://doi.org/10.1016/j.jss.2017.09.025.
Bharadwaj, Anandhi, Omar A. El Sawy, Paul A. Pavlou, and N. Venkatraman. 2013. “Digital Business Strategy: Toward a Next Generation of Insights.” MIS Quarterly 37 (2): 471–82. https://doi.org/10.25300/MISQ/2013/37:2.3.
Buer, Sven-Vegard, Marco Semini, Jan Ola Strandhagen, and Fabio Sgarbossa. 2021. “The Complementary Effect of Lean Manufacturing and Digitalisation on Operational Performance.” International Journal of Production Research 59 (7): 1976–92. https://doi.org/10.1080/00207543.2020.1790684.
Buer, Sven-Vegard, Jan Ola Strandhagen, and Felix T. S. Chan. 2018. “The Link Between Industry 4.0 and Lean Manufacturing: Mapping Current Research and Establishing a Research Agenda.” International Journal of Production Research 56 (8): 2924–40. https://doi.org/10.1080/00207543.2018.1442945.
Bygstad, Bendik, and Egil Øvrelid. 2020. “Architectural Alignment of Process Innovation and Digital Infrastructure in a High-Tech Hospital.” European Journal of Information Systems 29 (3): 220–37. https://doi.org/10.1080/0960085X.2020.1728201.
Cenfetelli, Ronald T., and Geneviève Bassellier. 2009. “Interpretation of Formative Measurement in Information Systems Research.” MIS Quarterly 33 (4): 689–707. https://doi.org/10.2307/20650323.
Chan, David. 1998. “Functional Relations Among Constructs in the Same Content Domain at Different Levels of Analysis: A Typology of Composition Models.” Journal of Applied Psychology 83 (2): 234–46. https://doi.org/10.1037/0021-9010.83.2.234.
Churchill, Gilbert A. 1979. “A Paradigm for Developing Better Measures of Marketing Constructs.” Journal of Marketing Research 16 (1): 64–73. https://doi.org/10.1177/002224377901600110.
Cifone, Fabiana Dafne, Kai Hoberg, Matthias Holweg, and Alberto Portioli Staudacher. 2021. ‘Lean 4.0’: How Can Digital Technologies Support Lean Practices?” International Journal of Production Economics 241: 108258. https://doi.org/10.1016/j.ijpe.2021.108258.
Colfer, Lyra J., and Carliss Y. Baldwin. 2016. “The Mirroring Hypothesis: Theory, Evidence, and Exceptions.” Industrial and Corporate Change 25 (5): 709–38. https://doi.org/10.1093/icc/dtw027.
Conway, Melvin E. 1968. “How Do Committees Invent?” Datamation 14 (4): 28–31.
Corley, Kevin G., and Dennis A. Gioia. 2011. “Building Theory about Theory Building: What Constitutes a Theoretical Contribution?” Academy of Management Review 36 (1): 12–32. https://doi.org/10.5465/amr.2009.0486.
Cunningham, Ward. 1992. “The WyCash Portfolio Management System.” Addendum to the Proceedings of OOPSLA ’92, 29–30. https://doi.org/10.1145/157709.157715.
Daoudi, Sara, Malin Larsson, Simon Hacks, and Jürgen Jung. 2023. “Discovering and Assessing Enterprise Architecture Debts.” Complex Systems Informatics and Modeling Quarterly, no. 35: 1–29. https://doi.org/10.7250/csimq.2023-35.01.
Davenport, Thomas H., and James E. Short. 1990. “The New Industrial Engineering: Information Technology and Business Process Redesign.” Sloan Management Review 31 (4): 11–27.
Diamantopoulos, Adamantios, and Heidi M. Winklhofer. 2001. “Index Construction with Formative Indicators: An Alternative to Scale Development.” Journal of Marketing Research 38 (2): 269–77. https://doi.org/10.1509/jmkr.38.2.269.18845.
Dul, Jan. 2016. “Necessary Condition Analysis (NCA): Logic and Methodology of ‘Necessary but Not Sufficient’ Causality.” Organizational Research Methods 19 (1): 10–52. https://doi.org/10.1177/1094428115584005.
Ernst, Neil A., Stephany Bellomo, Ipek Ozkaya, Robert L. Nord, and Ian Gorton. 2015. “Measure It? Manage It? Ignore It? Software Practitioners and Technical Debt.” Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering (ESEC/FSE), 50–60. https://doi.org/10.1145/2786805.2786848.
Evans, Eric. 2003. Domain-Driven Design: Tackling Complexity in the Heart of Software. Addison-Wesley.
Galbraith, Jay R. 1974. “Organization Design: An Information Processing View.” Interfaces 4 (3): 28–36. https://doi.org/10.1287/inte.4.3.28.
Galsworth, Gwendolyn D. 1997. Visual Systems: Harnessing the Power of the Visual Workplace. AMACOM.
Goldratt, Eliyahu M., and Jeff Cox. 1984. The Goal: A Process of Ongoing Improvement. North River Press.
Hacks, Simon, Hendrik Höfert, Johannes Salentin, Yoon Chow Yeong, et al. 2019. “Towards the Definition of Enterprise Architecture Debts.” 2019 IEEE 23rd International Enterprise Distributed Object Computing Workshop (EDOCW). https://doi.org/10.1109/edocw.2019.00016.
Haki, Kazem, Annamina Rieder, Lorena Buchmann, and Alexander W. Schneider. 2023. “Digital Nudging for Technical Debt Management at Credit Suisse.” European Journal of Information Systems 32 (1): 64–80. https://doi.org/10.1080/0960085X.2022.2088413.
Hammer, Michael. 1990. “Reengineering Work: Don’t Automate, Obliterate.” Harvard Business Review 68 (4): 104–12.
Hayes, Andrew F., and Klaus Krippendorff. 2007. “Answering the Call for a Standard Reliability Measure for Coding Data.” Communication Methods and Measures 1 (1): 77–89. https://doi.org/10.1080/19312450709336664.
Hevner, Alan R., Salvatore T. March, Jinsoo Park, and Sudha Ram. 2004. “Design Science in Information Systems Research.” MIS Quarterly 28 (1): 75–105. https://doi.org/10.2307/25148625.
Hicks, B. J. 2007. “Lean Information Management: Understanding and Eliminating Waste.” International Journal of Information Management 27 (4): 233–49. https://doi.org/10.1016/j.ijinfomgt.2006.12.001.
Hopp, Wallace J., and Mark L. Spearman. 2004. “To Pull or Not to Pull: What Is the Question?” Manufacturing & Service Operations Management 6 (2): 133–48. https://doi.org/10.1287/msom.1030.0028.
Jaakkola, Elina. 2020. “Designing Conceptual Articles: Four Approaches.” AMS Review 10 (1): 18–26. https://doi.org/10.1007/s13162-020-00161-0.
Jung, Jurgen, Simon Hacks, Thijmen de Gooijer, Matti Kinnunen, et al. 2021. “Revealing Common Enterprise Architecture Debts: Conceptualization and Critical Reflection on a Workshop Format Industry Experience Report.” 2021 IEEE 25th International Enterprise Distributed Object Computing Workshop (EDOCW). https://doi.org/10.1109/edocw52865.2021.00058.
Kallinikos, Jannis, Aleksi Aaltonen, and Attila Marton. 2013. “The Ambivalent Ontology of Digital Artifacts.” MIS Quarterly 37 (2): 357–70. https://doi.org/10.25300/MISQ/2013/37.2.02.
Kane, Gerald C., Doug Palmer, Anh Nguyen Phillips, David Kiron, and Natasha Buckley. 2019. “Accelerating Digital Innovation Inside and Out.” MIT Sloan Management Review and Deloitte Insights, 1–26.
Kent, William. 2012. Data and Reality: A Timeless Perspective on Perceiving and Managing Information. 3rd ed. Technics Publications.
Ketokivi, Mikko, Saku Mantere, and Joep Cornelissen. 2017. “Reasoning by Analogy and the Progress of Theory.” Academy of Management Review 42 (4): 637–58. https://doi.org/10.5465/amr.2015.0322.
Kim, Gene, Jez Humble, Patrick Debois, and John Willis. 2016. The DevOps Handbook. IT Revolution Press.
Klein, Katherine J., Fred Dansereau, and Rosalie J. Hall. 1994. “Levels Issues in Theory Development, Data Collection, and Analysis.” Academy of Management Review 19 (2): 195–229. https://doi.org/10.2307/258703.
Kruchten, Philippe, Robert L. Nord, and Ipek Ozkaya. 2012. “Technical Debt: From Metaphor to Theory and Practice.” IEEE Software 29 (6): 18–21. https://doi.org/10.1109/MS.2012.167.
Li, Zengyang, Paris Avgeriou, and Peng Liang. 2015. “A Systematic Mapping Study on Technical Debt and Its Management.” Journal of Systems and Software 101: 193–220. https://doi.org/10.1016/j.jss.2014.12.027.
Liker, Jeffrey K. 2004. The Toyota Way: 14 Management Principles from the World’s Greatest Manufacturer. McGraw-Hill.
MacCormack, Alan, Carliss Baldwin, and John Rusnak. 2012. “Exploring the Duality Between Product and Organizational Architectures: A Test of the ‘Mirroring’ Hypothesis.” Research Policy 41 (8): 1309–24. https://doi.org/10.1016/j.respol.2012.04.011.
MacKenzie, Scott B., Philip M. Podsakoff, and Nathan P. Podsakoff. 2011. “Construct Measurement and Validation Procedures in MIS and Behavioral Research: Integrating New and Existing Techniques.” MIS Quarterly 35 (2): 293–334. https://doi.org/10.2307/23044045.
Malone, Thomas W., and Kevin Crowston. 1994. “The Interdisciplinary Study of Coordination.” ACM Computing Surveys 26 (1): 87–119. https://doi.org/10.1145/174666.174668.
Markus, M. Lynne. 2004. “Technochange Management: Using IT to Drive Organizational Change.” Journal of Information Technology 19 (1): 4–20. https://doi.org/10.1057/palgrave.jit.2000002.
Murungi, David, and Rudy Hirschheim. 2022. “Theory Through Argument: Applying Argument Mapping to Facilitate Theory Building.” European Journal of Information Systems 31 (4): 437–62. https://doi.org/10.1080/0960085X.2020.1868952.
Netland, Torbjørn H. 2016. “Critical Success Factors for Implementing Lean Production: The Effect of Contingencies.” International Journal of Production Research 54 (8): 2433–48. https://doi.org/10.1080/00207543.2015.1096976.
Ohno, Taiichi. 1988. Toyota Production System: Beyond Large-Scale Production. Productivity Press.
Parnas, David L. 1972. “On the Criteria to Be Used in Decomposing Systems into Modules.” Communications of the ACM 15 (12): 1053–58. https://doi.org/10.1145/361598.361623.
Petter, Stacie, Detmar Straub, and Arun Rai. 2007. “Specifying Formative Constructs in Information Systems Research.” MIS Quarterly 31 (4): 623–56. https://doi.org/10.2307/25148814.
Podsakoff, Philip M., Scott B. MacKenzie, Jeong-Yeon Lee, and Nathan P. Podsakoff. 2003. “Common Method Biases in Behavioral Research: A Critical Review of the Literature and Recommended Remedies.” Journal of Applied Psychology 88 (5): 879–903. https://doi.org/10.1037/0021-9010.88.5.879.
Poppendieck, Mary, and Tom Poppendieck. 2003. Lean Software Development: An Agile Toolkit. Addison-Wesley.
Rahm, Erhard, and Philip A. Bernstein. 2001. “A Survey of Approaches to Automatic Schema Matching.” The VLDB Journal 10 (4): 334–50. https://doi.org/10.1007/s007780100057.
Richter, Nicole Franziska, Sandra Schubring, Sven Hauff, Christian M. Ringle, and Marko Sarstedt. 2020. “When Predictors of Outcomes Are Necessary: Guidelines for the Combined Use of PLS-SEM and NCA.” Industrial Management & Data Systems 120 (12): 2243–67. https://doi.org/10.1108/IMDS-11-2019-0638.
Rolland, Knut-H., and Kalle Lyytinen. 2021. “Exploring the Tensions Between Management of Architectural Debt and Digital Innovation: The Case of a Financial Organization.” Proceedings of the Annual Hawaii International Conference on System Sciences. https://doi.org/10.24251/hicss.2021.807.
Ross, Jeanne W., Peter Weill, and David C. Robertson. 2006. Enterprise Architecture as Strategy: Creating a Foundation for Business Execution. Harvard Business School Press.
Rother, Mike, and John Shook. 1999. Learning to See: Value Stream Mapping to Create Value and Eliminate Muda. Lean Enterprise Institute.
Roy, Chanchal K., James R. Cordy, and Rainer Koschke. 2009. “Comparison and Evaluation of Code Clone Detection Techniques and Tools: A Qualitative Approach.” Science of Computer Programming 74 (7): 470–95. https://doi.org/10.1016/j.scico.2009.02.007.
Sanchez, Ron, and Joseph T. Mahoney. 1996. “Modularity, Flexibility, and Knowledge Management in Product and Organization Design.” Strategic Management Journal 17 (S2): 63–76. https://doi.org/10.1002/smj.4250171107.
Seddon, Peter B., Cheryl Calvert, and Song Yang. 2010. “A Multi-Project Model of Key Factors Affecting Organizational Benefits from Enterprise Systems.” MIS Quarterly 34 (2): 305–28. https://doi.org/10.2307/20721429.
Shah, Rachna, and Peter T. Ward. 2003. “Lean Manufacturing: Context, Practice Bundles, and Performance.” Journal of Operations Management 21 (2): 129–49. https://doi.org/10.1016/S0272-6963(02)00108-0.
Shingo, Shigeo. 1986. Zero Quality Control: Source Inspection and the Poka-Yoke System. Productivity Press.
Spear, Steven, and H. Kent Bowen. 1999. “Decoding the DNA of the Toyota Production System.” Harvard Business Review 77 (5): 96–106.
Strong, Diane M., and Olga Volkoff. 2010. “Understanding Organization–Enterprise System Fit: A Path to Theorizing the Information Technology Artifact.” MIS Quarterly 34 (4): 731–56. https://doi.org/10.2307/25750703.
Suddaby, Roy. 2010. “Editor’s Comments: Construct Clarity in Theories of Management and Organization.” Academy of Management Review 35 (3): 346–57. https://doi.org/10.5465/amr.2010.51141319.
Trist, Eric L., and Ken W. Bamforth. 1951. “Some Social and Psychological Consequences of the Longwall Method of Coal-Getting.” Human Relations 4 (1): 3–38. https://doi.org/10.1177/001872675100400101.
Tushman, Michael L., and David A. Nadler. 1978. “Information Processing as an Integrating Concept in Organizational Design.” Academy of Management Review 3 (3): 613–24. https://doi.org/10.2307/257550.
Walsh, James P., and Gerardo Rivera Ungson. 1991. “Organizational Memory.” Academy of Management Review 16 (1): 57–91. https://doi.org/10.2307/258607.
Wang, Richard Y., and Diane M. Strong. 1996. “Beyond Accuracy: What Data Quality Means to Data Consumers.” Journal of Management Information Systems 12 (4): 5–33. https://doi.org/10.1080/07421222.1996.11518099.
Westerman, George, Didier Bonnet, and Andrew McAfee. 2014. Leading Digital: Turning Technology into Business Transformation. Harvard Business Review Press.
Whetten, David A. 1989. “What Constitutes a Theoretical Contribution?” Academy of Management Review 14 (4): 490–95. https://doi.org/10.2307/258554.
Wimelius, Henrik, Lars Mathiassen, Jonny Holmström, and Mark Keil. 2020. “A Paradoxical Perspective on Technology Renewal in Digital Transformation.” Information Systems Journal 31 (1): 198–225. https://doi.org/10.1111/isj.12307.
Womack, James P., Daniel T. Jones, and Daniel Roos. 1990. The Machine That Changed the World. Rawson Associates.
Yoo, Youngjin, Ola Henfridsson, and Kalle Lyytinen. 2010. “The New Organizing Logic of Digital Innovation: An Agenda for Information Systems Research.” Information Systems Research 21 (4): 724–35. https://doi.org/10.1287/isre.1100.0322.

  1. We use structural waste in a sense specific to information-system architecture. The phrase appears loosely in lean construction (waste of structural materials) and in health-economics waste taxonomies; neither referent is intended here.↩︎