An organisation that only explains its large failures will never discover which of its small ones were rehearsals.
Every organisation triages. Attention is the scarcest thing a leadership team has, so the events explained are the ones that hurt; the rest are noted and released. Nobody writes the rule down or defends it in a governance forum. It runs anyway, and does more work than almost any policy on the books.
In one review meeting recorded by J.S. Busby in a 1999 Project Management Journal study of post-project reviews, a participant raised a piece of equipment sited where it was awkward to maintain. It was small, low in value and needed little maintenance, so the problem was dismissed as very minor and the discussion moved on. Busby's judgement is that the dismissal made the error of assuming that minor outcomes reflect minor causes — and that maintainability is, for customers in that industry, sometimes critical. He goes further: the ease of the dismissal suggested the organisation's engineers had too little awareness of the issue, perhaps through a lack of training, of formal process, or of knowledge transfer between them. None of that was explored [FACT CHECK REQUIRED] [SOURCE DETAILS REQUIRED].
He is careful about the general case. Organisations with limited time naturally attend to big outcomes; his objection is not to triage but to unthinking triage, because then one misses big issues simply because, on isolated occasions, they happened not to have big outcomes.
The Strategic Context
The triage rule is defensible on its face. Consequence is observable, immediate and easy to rank; cause is none of those things, and ranking by cause would require first doing the work you are deciding whether to do.
But it rests on an assumption rarely stated: that the size of an outcome carries information about the size of its cause. Often it does. A large loss usually had a large mechanism behind it, and investigating large losses is not a mistake.
The assumption fails in an identifiable class of case — where the mechanism is a defect in a definition, an interface or an assumption, and the consequence is set by circumstance rather than by the defect. Ian Whittingham opens a 2009 column on the clearest available instance: two spacecraft lost within ten weeks in 1999, the initial investigation into the first focused on a mismatch in the units used to control its trajectory — one party working in newtons, a subcontractor in imperial units. That account sits behind a paywall and stops mid-sentence, so nothing is concluded here beyond the shape [FACT CHECK REQUIRED]. The shape is the point: the same mismatch produces nothing on most days and, on one day, a total loss.
This is why the second half of Busby's point matters more than the first: good learning stimulated by a minor adverse outcome, he notes, could pay large dividends later, when the same mechanism fires under worse conditions. An organisation triaging by consequence is not choosing which failures to investigate. It is letting circumstance choose.
What the Triage Rule Actually Selects For
Three things follow, none of them obvious from inside.
The stock of explanations becomes a biased sample. It holds the mechanisms that fired under bad conditions and excludes those that fired under good ones, so any inference about what tends to go wrong is drawn from a sample selected on the outcome variable. Whether evidence came from a representative population is taken up elsewhere [Related article: Volunteers Are Not a Sample]; the prior question is which events entered the evidence base at all.
The near miss is excluded by construction, and it is the cheapest teaching material an enterprise owns — an event whose mechanism completed and whose consequence did not. Under a consequence rule it scores zero.
Nobody is accountable for a rule nobody wrote. No forum exists in which somebody says we have decided not to understand this. There is a meeting in which a small thing is mentioned and the conversation moves on, which feels like nothing happening and is a decision with a cost.
Reframing the Issue
Stop treating investigation as a response and start treating it as an allocation. An enterprise has a finite investigation budget — hours of senior attention, mostly — and allocates it by a rule it has never examined. The strategic question is not should we investigate this failure but what is our standing rule for which events earn an explanation, and does it measure anything other than the size of the consequence?
A standing rule can be written down, debated and defended, and can include classes of event a case-by-case judgement would never select — because the point of those classes is that they look unimportant when they arrive.
It also separates this question from a neighbouring one. Designing indicators so a signal reaches you at all, fast enough to act on, is a different discipline with a different failure mode [Related article: Measuring an Outcome You Cannot Predict]. The concern here is a signal that did arrive, was heard, and was released.
The Level at Which a Finding Is Written Down
A second finding from the same study compounds the first, and it decides whether an investigation is worth its cost.
Busby found diagnoses consistently too concrete — too narrow a view of the process, too literal a reading of what went wrong. His example: a product fails because a designed part's wall thickness was insufficient. One diagnosis is we failed to specify an adequate wall thickness, and its remedy is a line in the design codes. Another is our designers do not understand the extremes of operating duty their products have to meet, and its remedy is exposure to customers using them.
Both are true. Only the second is worth the meeting. The first produces a rule catching this defect and nothing adjacent; the second changes a capability.
His conclusion is that being too concrete produces strictly incremental learning — small revisions to current knowledge rather than replacements of it — and that persistent incremental learning leaves an organisation unable to react to large changes in its environment. He immediately hedges: he had no indication that any firm he studied faced such a change, though most organisations face one eventually. Keep the hedge, because the argument needs no more than it says. Incremental learning is not wrong. It is the only kind available to an enterprise that never asks whether the thing in front of it is an instance of something general.
Six References in Twelve Hours
The third finding closes the loop. Across twelve hours of meetings Busby recorded six references to historical experience, serving three purposes: as evidence for an explanation, to show something had changed by contrast with older events, and to explain why people behaved as they did.
None helped anyone establish whether an event was systemic.
He states the stake plainly: you cannot know whether an event on a single project is unique, frequent or systemic unless you examine other completed projects. Without that, every failure arrives looking like a first occurrence — precisely the reading that justifies a small response.
He is equally careful about the opposite error. History can be consulted badly, as a template for a future that will not resemble it; his resolution is to look at it at the right level of generality, since your next undertaking will differ in every detail and still run through very similar processes.
Note what this asks. It asks whether anyone in the room put the question. Whether your records can answer it, and whether they ever change a subsequent estimate, is a different instrument and belongs elsewhere [Related article: The Estimating Loop Nobody Closes].
Two Enterprises, Two Thresholds
Consider, hypothetically, a metropolitan water authority running a pressure-management programme. A control valve at a suburban node behaves unexpectedly during a routine adjustment; the operator corrects it manually and logs it. No burst, no customers affected, no reportable event. Under a consequence rule the entry closes. Under a threshold that includes a device that behaved other than as its configuration described, someone asks why — and finds a commissioning convention under which several hundred valves were configured from a template last revised for a different generation of hardware. The consequence of the event was zero. The consequence of the mechanism is a number nobody has calculated.
The same shape appears in a quick-service restaurant network, hypothetically, auditing franchise compliance across nine hundred stores. One store fails a temperature check by a small margin, corrects it on the spot, and the finding closes as minor. Nobody asks whether it was a store problem, a training problem or an equipment-specification problem, because the outcome was small and the store was compliant by the afternoon. Four years of audits have produced no finding pitched above the level of the individual store — four years of activity and no organisational knowledge.
In both cases the enterprise did nothing wrong on the day, and in both the rule that governed the day was never written down.
Decision Framework
The investigation threshold. A short standing rule, owned by whoever owns operational risk, naming the classes of event that earn an explanation regardless of consequence:
- an interface defect — anything that went wrong at the boundary between two teams, systems, contracts or disciplines;
- a definitional mismatch — two parties using the same word or unit differently;
- a near miss on a safety, regulatory or financial-control boundary, crossed or not;
- a surprise in a counterparty's behaviour, which is evidence your model of them is wrong;
- any event a competent person present found strange — the only class that catches what the other four miss.
The rule should also fix the level of generality at which findings are written down, because that is where the value is created or lost, and it defaults to the specific if left unstated.
The severity test. Would we have investigated this had the consequence been ten times larger? If yes, what are we actually triaging on, and would we defend it in front of a board?
The generality test. Is this finding recorded as a fact about this part, this design habit, or this class of decision — and was anyone asked to choose?
The recurrence test. When we last explained a failure, did anyone ask whether it had happened before, and was the question asked about the part or the class of decision? The test asks only whether the question is put; what your records can answer is a separate matter.
From Strategy to Execution
Immediately. Take the last twenty entries in whatever log receives minor events — incident register, audit findings, defect list, review actions — and mark those that would have earned an explanation under the five classes above. The number is usually larger than expected, and it is the first honest measure of the current rule.
Over the next two quarters. Write the threshold down, give it an owner, and fund it. An investigation with no budget line competes with delivery work and loses — to a competent manager making a defensible decision. Set the level-of-generality requirement explicitly, and require every finding to name the class of decision it belongs to, not only the artefact it was found in.
Over years. The durable capability is a base rate. An enterprise recording findings at a consistent level of generality for five years can answer has this happened before with a number rather than a recollection, and treat a small event as evidence rather than anecdote. That is also what allows a risk rating to be revisited on evidence rather than the calendar [Related article: Which of Your Risk Responses Changes the Probability?].
One boundary. This article treats failure and says nothing about the other half of the distribution: an unexpectedly good outcome is also an event nobody explained, and the case for treating deviation as two-sided is made elsewhere from a fuller source [Related article: Risk Is Not the Chance That Things Go Wrong].
Signals to Monitor
- The correlation between consequence size and whether an event was investigated. Close to one means you have no threshold, only a ranking.
- The proportion of investigations arising from near misses rather than losses. In most enterprises it is close to zero and nobody has noticed.
- The level of generality of your last twenty findings — how many name a class of decision rather than a component.
- Repeat findings nobody recognises as repeats, the signature of a record too specific to match itself.
- Investigations commissioned only after something expensive happened, which biases the stock of explanations toward disasters and teaches the organisation that enquiry is punishment [Related article: The Review That Cannot Ask Why].
- Whether findings reach anyone outside the initiative that produced them, since a threshold that selects well and disseminates nothing has only relocated the loss [Related article: Written Down Is Not Passed On].
Questions for the Leadership Team
- What is our standing rule for which failures earn an explanation, and where is it written down?
- Of the events we investigated last year, how many had small consequences — and what does that say about the rule we are running?
- Who decides whether a finding is about a component, a habit or a class of decision?
- Could we answer, with a number rather than a recollection, whether our last significant failure had happened before?
- What would a year of investigating every interface defect and definitional mismatch cost, and what would we expect to find?
- Which of last quarter's near misses would we now like to have understood?
Closing Perspective
The equipment in that review was small, cheap and easy to leave alone. Everyone made a reasonable judgement, and it was reasonable because it was about the equipment. What was not small was the awareness gap the dismissal revealed — and nobody looked, because nothing in the way the meeting ran required anyone to.
Triage is unavoidable. Triaging on a variable the enterprise did not choose is not. Consequence size is set by circumstance: whether the valve was on a main or a spur, whether the mismatch was caught in integration or in flight, whether the store failed by one degree or ten. An organisation that lets circumstance govern which of its own mechanisms it understands has outsourced its learning agenda to luck.
The decision is not which failure to investigate next. It is whether the rule that answers that question is one the leadership team would be willing to write down.
About EraNorth Insights
EraNorth Insights publishes practical analysis on strategy, projects, operations, transformation and decision intelligence for professional and organisational use. About EraNorth.
