An open research agenda on what retrospective analysis of organizational failure can and cannot support—asked as a question that would exist whether or not this practice did.
Causal Failure Analysis is the study of what can legitimately be claimed to cause organizational failure, and under what evidentiary conditions such a claim is warranted—a question, not yet a settled answer. Causal Failure Diagnostics is the applied practice built on it: using what the analysis has established, and only what it has established, to read a specific situation. Together, they are the science behind FPI™. The practice applies what this research establishes; the research does not exist to justify the practice.
Most organizational research studies success. The reasons are practical—successful firms are easier to study, more willing to participate, and more pleasant to write about—but the consequence is a literature skewed toward survivorship. Far less work asks a harder, prior question: what, exactly, can be claimed on the basis of documented failure?
Under what conditions, if any, can retrospective analysis of organizational failure support mechanism-level causal claims—given that organizations are open, reflexive systems whose regularities are actively disrupted by the failure event itself, rather than the kind of stable, isolable systems classical causal frameworks were built to describe?
Put more plainly: can a documented set of failure patterns support a claim stronger than "this happened before"—without either overclaiming predictive power it can't support, or collapsing into pure narrative? That is the question this research agenda exists to work on. It does not yet have a complete answer.
Most organizational failure analysis—including early stages of this research—is selected entirely on the dependent variable. It studies firms that failed and infers which conditions caused the failure, without checking whether those same conditions were equally present in firms that did not fail. This is a known problem in case-study methodology (selection bias and equifinality), but it is rarely audited at the level of an individual claim.
The more consequential version of the problem is not that this happens—most retrospective failure research shares this limitation—but that standard practice conflates two things that should be measured separately: how many independent sources document a pattern, and whether any of those sources checked the pattern against comparable cases that did not fail. A pattern corroborated by five sources that all studied only failures is not more causally warranted than a pattern corroborated by one source that checked a genuine comparison case. It is only more repeated.
Applying this standard to our own production database: of the patterns documented to date, only a small fraction have been assessed against comparison cases at all, and two now meet the strongest tier—a documented, source-resolved comparison between a failure case and a matched case that did not fail. This is stated here because it is true, not because it is flattering. A well-sourced but uncontrolled pattern is still useful—but it is epistemically distinct from one that has been checked against a real comparison, and treating the two as equivalent produces false confidence.
Four specific, currently unresolved gaps in this line of research:
When a documented pattern is found to hold in one sector, era, or regulatory regime, what would it actually take to establish—rather than assume—that the same causal mechanism is active in a different context? No general answer exists yet. The current position is that transport requires case-specific re-verification, not extrapolation from surface similarity.
Retrospective single-case evidence can support "this factor contributed to the failure." It cannot, without considerably more structural work, support "this factor was necessary" or "this factor was sufficient." Distinguishing these in practice—auditing existing causal-language claims for overclaiming—is unfinished work.
Do documented failure patterns' practical value decay over time as they become more widely known and acted on—a Goodhart- or Merton-style dynamic, where once a pattern is publicized, the population of future firms exhibiting it may no longer resemble the population it was extracted from? This is a plausible, falsifiable hypothesis, not yet testable against sufficient data.
A live debate in safety science (Hollnagel's Safety-II and FRAM literature) argues that failure may be better explained as ordinary variability in multiple normal functions coupling unexpectedly, rather than a traceable chain of causes. This directly challenges the premise underneath most failure analysis, including this one. The working position here is mechanism-based causal claims, cautiously scoped—but the alternative paradigm has not been refuted, only set aside as less immediately actionable.
Two properties of every documented pattern are tracked separately, and deliberately never combined into a single confidence score: how well-corroborated a pattern is (how many independent sources document it, and how convergent their accounts are), and whether the pattern's causal claim has been checked against comparable cases that did not result in failure.
These are different questions. A pattern can be extremely well-corroborated and still rest entirely on failure-only evidence. Conflating the two produces false confidence. A well-sourced but uncontrolled pattern is treated as exactly that—useful, worth attention, but epistemically distinct from a pattern checked against a real comparison case—stated explicitly rather than letting strong sourcing imply a comparison that was never made.
Both cases in the Working Papers section below used the same formal test to reach that determination, applied consistently regardless of which way it came out.
An attempt to apply the matched-comparison standard described above to a documented case pair drawn from Clayton Christensen's The Innovator's Dilemma: Quantum's successful spin-out reintegration against IBM's PC Division's unsuccessful one, where Christensen's own account specifies the actual distinguishing variable. This was the first pattern in the corpus to meet the strongest evidentiary tier described in the Methodological Commitments above.
A second pattern in the same source, using the same structural logic (organizational autonomy versus embeddedness), independently reached the same evidentiary tier—raising the open question of whether this reflects something general about matched-comparison reasoning, or something specific to how this particular author constructs arguments. See Open Question 1, Transportability, and a planned different-author extension below.
A separate case where the same matched-comparison test was attempted and did not resolve—the comparison case was insufficiently documented to support a conclusion either way. The specific reason is recorded rather than the attempt being discarded, consistent with the standard above: an inconclusive test is itself a data point, not a failure to hide.