An organisation that cannot state the condition under which it would stop an initiative has not funded a decision. It has funded a direction of travel.
Most enterprises can describe, at length, what their AI initiatives are meant to achieve. Very few can state the condition under which they would be permitted to stop one.
That asymmetry is not carelessness. It follows from a missing measurement. Deployment questions are argued as though the standard were self-evident — is it accurate enough, is it ready, is it safe — when the only standard that means anything commercially is comparative: is it better than what we do now, and better by enough to justify what it costs to run and to supervise? Without a measured account of how the current human decision performs, every result the technical team produces is arguable, and an arguable result never loses a funding round.
The consequence is a portfolio in which nothing is ever shelved. Initiatives are extended, re-scoped, quietly re-staffed and occasionally rebranded, but the disposition that says not this, not yet, and here is what would change our mind is almost never taken, because nobody owns it and nothing triggers it.
The Strategic Context
There is a rule available that resolves this, and it is more useful than its plainness suggests. Compare the system's performance against the human judgement it would replace, and take one of three dispositions. If it clearly surpasses that judgement, it becomes the primary instrument for the decision. If it is close but not better, it becomes decision support — it advises, a person decides. If it lags materially, it is set aside until either the data or the methods improve.
Read as a technical protocol this is unremarkable. Read as a capital instrument it is considerably more interesting, because the three branches are not three grades of the same outcome. They have different cost structures, different owners, different measures of success and different conditions for reversal. Choosing between them is an allocation decision that belongs to the executive, not a threshold that belongs to the data team.
It is also worth being clear about what this test is not. A stage gate asks whether an initiative has done what it undertook to do — whether the plan was followed, the deliverables produced, the assurance completed. That is a question about conformance, and a well-run initiative can pass it while producing something the business should not deploy. The question here is orthogonal: measured against the person doing this job today, is the machine better, comparable, or worse?
What a Single Accuracy Figure Conceals
The benchmark is usually imagined rather than measured. Ask what the current error rate of the human decision is, in the same units, over the same period, and most organisations cannot say. The comparison then defaults to a recollection of how the best practitioner performs on a good day, which is not the process being replaced.
An average conceals the distribution that matters. A system that is better overall and worse on the rare, expensive cases is not better; it is a redistribution of error from the common and cheap towards the uncommon and costly. Any benchmark worth acting on requires an agreed cost weighting across error types, set by the executive who bears those costs rather than by the team building the system.
A test result is a claim about a past sample, not a forecast. Performance is measured against data drawn from conditions that held when it was collected. The familiar practice of dividing a dataset into training, tuning and held-out portions is a working convention of the discipline rather than a standard, and the proportions people quote carry no authority. What matters is that the held-out portion was genuinely unseen and, wherever possible, drawn from a later period than the data used to build the system — a much harder test, and a far more honest one.
Readiness is confused with authority. Passing a benchmark says a system performs. It does not say anyone has granted it the right to act, within what bounds, with what reversal route — a separate governance question with its own instrument [Related article: What Has Your Enterprise Already Authorised a Machine to Decide?].
Reframing the Issue: Three Dispositions, Three Cost Structures
Primary instrument. The system makes the decision; people handle exceptions and monitor performance. The cost structure is a genuine substitution of labour, offset by three costs that business cases routinely omit: continuous monitoring, incident response, and an exception path that will be under-resourced within a year because exceptions are unglamorous and their volume is unpredictable. The owner is the line executive who owns the decision today — not the technology function, which cannot be accountable for an outcome it does not control.
Decision support. The system proposes; a person decides. This branch looks like the cautious middle and is, per decision, the most expensive of the three. You carry the full cost of the system and the full cost of the human decision, plus a new reconciliation cost every time the two disagree. It removes no labour by design. Which means decision support cannot be justified on efficiency and must be justified on decision quality — better consistency, fewer catastrophic errors, faster resolution of hard cases. Most organisations fund this branch while promising the economics of the first, and are then puzzled when the savings never arrive.
Shelved. Not cancelled, and the difference is material. Cancellation destroys the asset; shelving preserves it against a named constraint at a small carrying cost. What must be preserved is specific: the dataset and its labelling conventions, the evaluation harness, the human benchmark and the record of what was tried. What must accompany it is a reopening trigger expressed as an observable event — a data source that becomes available, a labelling backlog cleared, a change in the underlying methods — never a date, and never a person's continuing enthusiasm.
Building the Benchmark Before You Need It
The benchmark is the instrument. Everything else follows from it, and it has to be built while nobody has an interest in the answer.
The protocol is not complicated. Take a representative sample of real decisions from a defined period. Have experienced people decide them under normal conditions, blind to any system output. Record the outcomes where outcomes are observable, and where they are not, use an agreed adjudication panel and say plainly that this is what you have done. Weight the errors by their cost, using weights agreed in advance. Then run the system on the same sample.
Two things follow that leaders consistently underestimate.
The first is that the benchmark is valuable even if nothing is ever deployed. Measuring how your own experienced people decide the same cases usually reveals variation between them that nobody had quantified — and inconsistency among decision-makers is a finding with immediate operational value, entirely independent of any technology.
The second is that the benchmark must be re-measured. People learn; processes change; the population being decided upon shifts. A system that surpassed the human standard two years ago may not surpass today's, particularly if the earlier comparison prompted the improvement. A benchmark measured once and cited thereafter becomes a claim rather than evidence, and it will be cited long after it has stopped being true.
Where the system sits inside something the enterprise sells rather than inside its own operations, the comparator changes entirely: the relevant standard is what a customer's alternative delivers, not what your staff achieve, and the two rarely coincide [Related article: Is Your AI a Feature You Sell, or a Tool You Use?].
Why Nothing Is Ever Shelved
The disposition rule fails in practice for reasons that have nothing to do with measurement.
Shelving has no owner, or its only owner is the person who loses by it. If the decision rests with the sponsor, the sponsor is being asked to end their own initiative on evidence they commissioned. Very few do, and the ones who do are not rewarded for it.
The constraint is never named, so the shelf becomes a graveyard. Set something aside without recording what specifically has to change, and nobody can ever tell whether the moment to reopen has arrived. The initiative does not stay shelved; it decays.
Sunk cost arrives dressed as commitment. The larger the spend, the stronger the case for continuing, which is exactly backwards. What has been spent is not recoverable by any disposition; only the forward cost is a decision variable.
The alternative use of the capacity is invisible. Shelving frees scarce people. If no competing claim on those people is named at the point of the decision, the comparison is between doing this and doing nothing, and doing this always wins.
The design correction is straightforward and unpopular. The disposition owner is appointed at initiation, before results exist and before anyone's reputation is attached — and it is not the sponsor. Where the constraint is data quality or collection practice rather than method, the natural owner is the person accountable for that data across the business, which is itself an operating-model question rather than a technical one [Related article: Your Data Pipeline Is an Operating-Model Decision, Not an IT Project].
Decision Framework
| Primary instrument | Decision support | Shelved | |
|---|---|---|---|
| Condition | Clearly surpasses the measured benchmark, including on high-cost errors | Comparable, or better on average but not on costly cases | Materially behind, or the benchmark cannot be measured |
| Justified on | Labour substitution and consistency at volume | Decision quality only — it removes no cost | Preserved option value against a named constraint |
| Owner | The line executive accountable for the decision | The function head whose people decide | The owner of the constraint that caused the shelving |
| Must produce | Monitoring, exception handling, a reversal route | Override rates, and examination of disagreements | A written constraint and an observable reopening trigger |
| Reverses when | Performance drifts below a re-measured benchmark | Overrides approach zero or approach half | The trigger fires, or the option is formally written off |
Four evidence requirements sit beneath the table. The benchmark must exist, be dated, and have been measured by someone other than the delivery team. The error cost weighting must be signed by the executive who bears the errors. The held-out evaluation must use data the system has never seen. And the reopening trigger must be written before the shelving decision, not after — a trigger drafted afterwards is written by whoever is least reconciled to the outcome.
From Strategy to Execution
Immediately, take the largest AI initiative in the portfolio and ask for its human benchmark: who measured it, when, on what sample, with what error weighting. If there is none, that is the finding, and it is more useful than any technical review. Ask second what would shelve the initiative, and note whether the answer is a condition or a mood.
Over the next two to three quarters, make disposition owners an appointment made at initiation. Build the benchmark protocol once and reuse it, so that measuring the human baseline stops being a negotiation each time. Give every decision-support deployment a decision-quality target rather than a savings target, and report against it — this single change usually resolves the confusion about why the promised efficiencies did not materialise.
Over years, the objective is a portfolio in which shelving is routine and unremarkable, because it is a disposition rather than a verdict on the people involved. That requires visible cases of an initiative being set aside, reopened when its trigger fired, and delivering — which is the only evidence that will persuade a capable sponsor that the shelf is not a euphemism. It also requires the enterprise to hold benchmarks as durable assets, re-measured on a cycle, owned by the business rather than by whichever team last needed one.
Signals to Monitor
- Initiatives with no stated stopping condition, expressed as a proportion of the AI portfolio. This is the primary measure.
- Benchmarks measured by the delivery team, or benchmarks older than the process they describe.
- Override rates in decision support falling towards zero, which means the human has stopped deciding, or sitting near half, which means the system is adding noise rather than judgement.
- Exception volumes in primary deployments growing faster than the resourcing of the exception path.
- Shelved work with no named constraint, or with a reopening trigger written as a date.
- Re-measured human performance improving, which can reverse a disposition without anything about the system having changed.
- Business cases claiming labour savings from a decision-support deployment, which is a category error visible on the page.
Questions for the Leadership Team
- For our largest AI initiative, what exactly would stop it, who decides, and when were they appointed?
- What is the measured performance of the human decision this would replace — and who measured it?
- Which errors are expensive for us, who agreed those weights, and does our evaluation reflect them?
- Where we have deployed decision support, what decision-quality improvement have we obtained for the cost of running both?
- What is currently on the shelf, what constraint put it there, and what observable event would bring it back?
- Whose scarce capacity would be released by shelving one initiative this quarter, and what would we ask them to do instead?
Closing Perspective
The three dispositions are one decision viewed through three cost structures, and refusing to choose does not avoid the choice. An enterprise that declines to decide has chosen decision support by default — running the system and the people, paying for both, and describing the result as prudence.
That is a defensible position if it is taken deliberately and measured on decision quality. It is an expensive accident if it is inherited because no one was willing to say the word shelved.
The deeper capability is not technical at all. It is the ability to conclude, on evidence, that something the organisation has invested in should stop for now — and to say so in a way that preserves the asset, names the constraint, and leaves a door that a future executive can actually find. Organisations that can do this allocate capital. Organisations that cannot merely spend it, and then explain the pattern afterwards as strategy.
About the author
Kevin Jogin is Founder & Principal Advisor at EraNorth. Meet the Founder.
