Risk and Resilience

Why One Machine Failure Costs More Than a Thousand Human Ones

Automated decisions are judged as policy because they trace back to a specification someone approved. That asymmetry belongs in your deployment threshold.

EraNorth Insights · 14 min read

A human error is an accident. A machine error is a policy — specified, approved and shipped — and it will be read as one.

An automated decision system produces one bad outcome. The response is not sympathy. It is a request for documents. Within days the enterprise is assembling the requirement that defined the behaviour, the data the model learnt from, the threshold somebody selected, the evaluation results, the approval, and the risk-register entry where a specialist raised the possibility and a committee closed it.

Now put the same magnitude of harm in the hands of a person doing the same job. The organisation investigates circumstances. It may retrain, redeploy or dismiss. It states what happened and what will change. It is rarely asked to produce the document in which the outcome was contemplated in advance, because no such document exists.

That difference is the subject. It is usually described as public irrationality — people are said to forgive humans and punish machines — and handled as a communications problem. It is better understood as a structural property of automated decisions, with direct consequences for the threshold at which an enterprise should be willing to deploy one.

The Strategic Context

Enterprises are crossing from advisory to autonomous in ordinary functions: screening, pricing, scheduling, credit assessment, triage, control-system operation. The crossing is rarely a single decision. A system recommends; humans agree with it; the agreement rate climbs; the review step becomes ceremonial; somebody removes it as a productivity gain. The authorisation that mattered was distributed across two years of small approvals, none of which felt like the moment.

The expected-value case for making that crossing is usually genuine. A system that outperforms the human average on a defined task will, over enough decisions, reduce total harm as well as total cost. It deserves to win more often than caution allows.

But the case is made on the mean and the consequence is felt in the tail. The paper in front of the board reports accuracy, throughput and cost per decision. Nothing in it describes what happens the first time the system is wrong in a way that matters — the event that decides whether the programme survives.

The Asymmetry Is Not Simply Irrationality

The received explanation is that people accept human imperfection more readily than machine imperfection. It is intuitively familiar, and it reaches most executive settings without a study, an author or a year attached — an observation worth testing rather than a finding to rely on [FACT CHECK REQUIRED]. Taken alone it is also unhelpful, because it invites leaders to manage the perception rather than the exposure.

Three reasons make the premium rational, none of which appeals to sentiment.

Correlation. A thousand human errors are close to independent draws. One machine error is a single draw executed repeatedly across the whole exposed population, simultaneously. Two portfolios with identical expected loss and different correlation are not equally risky; an insurer would not price them alike. Independence is what makes a loss absorbable; correlation is what makes it structural.

Attribution. Human error distributes across many actors, each carrying part of the responsibility. Machine error concentrates on one legal person — the enterprise that specified the behaviour, approved it and put it into production. There is no dilution available.

Recourse. A person harmed by a human decision can usually find someone to ask. A person harmed by an automated one often cannot obtain an explanation, a reconsideration or a name. Much of what looks like hostility towards machines is the absence of a route to remedy — the one component of the premium the enterprise controls.

Reframing the Issue

When a choice is made in advance, written into a specification and executed identically every time, it stops being an accident. It becomes a rule with an author.

This is the design-decision lens, and it is the mechanism behind the asymmetry. A rule attracts the questions rules attract: who wrote it, what alternatives were considered, what evidence supported the threshold, who approved it, and when it was last reviewed. Every one has an answer already in writing, because engineering produces artefacts as a by-product of doing engineering. The human case leaves no comparable trail; the responses differ not because society is fair to people and unfair to machines, but because in the machine case there is something to read.

Two consequences follow.

The first is that "the model did it" is not available as a defence. A model is not an actor. It is a written record of a management decision, executing repeatedly and at speed. Everything it does was authorised by somebody, at a point in time, under conditions that can be reconstructed — and reconstructing them is the first thing any competent counterparty, regulator or claimant will do.

The second is that judgements a human operator makes in a fraction of a second, unobserved, must now be settled in advance, in writing, for every future instance. That transfer is not optional. Declining to decide is itself a decision that gets encoded, and it is the hardest version to defend, because the record will show the question was reachable and nobody reached it.

One Draw, Repeated: Why Scale Changes the Kind of Risk

Independent errors aggregate into a predictable mean with modest variance around it. A systematic defect does not aggregate at all: it is one event of size N, where N is set by throughput — the same variable the business case was built to maximise. Automation improves the mean and concentrates the tail, and the second effect grows in exact proportion to the success of the first.

Three consequences follow.

Detection lag matters more than error rate. A defect operating undetected for a quarter produces a quarter of identical outcomes with one root cause. An organisation that monitors accuracy monthly, in aggregate and never by subgroup, has chosen a detection lag without discussing it.

Remediation is fleet-wide. The answer to a defective operator is to stand that operator down; the answer to a defective specification is to halt the system across every use, so the remediation cost bears no relationship to the size of the original harm. The cost of stopping is a design parameter, and few enterprises know theirs until they need it.

A deliberately hypothetical illustration: an enterprise runs forty thousand assessments a year through an automated screen, and a defect sits in it undetected for a quarter: ten thousand identically wrong outcomes, one cause, one owner, one remedy — and one disclosure. The same aggregate error rate spread across two hundred human assessors produces a scatter with no common cause, no single remedy and no single story. The arithmetic is identical. Nothing else about the two situations is.

Why the Response Is Discovery Rather Than Condolence

Anyone examining an automated failure works from documents, and the set is predictable: the specification and its change history, data provenance and exclusions, evaluation results including performance by subgroup, the threshold and its justification, escalations raised and how each was closed, and the monitoring plan with evidence that it was operated.

The last item is where most enterprises are exposed, and the exposure is self-inflicted: an organisation that documented a monitoring obligation and did not perform it has manufactured evidence of a control it defined and declined to run. Documentation protects only where practice matches it. Where it diverges, the file becomes the case against the enterprise rather than its defence.

A subtler version: where the internal test cannot detect the defect it purports to test for, the file records an assurance that never existed — reported in good faith, which makes it worse rather than better [Related article: Masking the Attribute Does Not Remove It]. Where harm becomes a claim, the questions move on to who carries it, which parts of the exposure sit inside the corporate structure and which sit outside it [Related article: What Limited Liability Actually Excludes].

The Deployment Threshold Belongs Above Parity

The standard proposition is that a system should be deployed once it outperforms the human benchmark. If the costs on each side were symmetric, that would be correct. They are not. Correlated exposure, concentrated attribution, fleet-wide remediation, evidentiary discoverability and the absence of recourse all sit on the machine side of the ledger, and none appears in a comparison of accuracy rates.

The conclusion is not that enterprises should hesitate. Hesitation has a body count of its own, paid in the harms the better system would have prevented. The conclusion is that the deployment threshold should carry a deliberate margin above human parity, that the size of the margin is a governance decision, and that a board should be able to state what margin it required and why.

A margin left unstated is still set — by whoever wanted the productivity benefit, at the level that made the business case work.

The margin is also reducible, which is the commercially interesting part. Build explanation and recourse and part of the premium disappears, because some of what the public prices is the unavailability of an answer. Shorten detection from quarters to days and more of it goes, because exposure is lag multiplied by throughput. Both are engineering investments that shrink the tail.

One further caution: parity against an unmeasured benchmark is not parity. Many organisations cannot state their human error rate on the same task, on the same definition. Where that is so, the business case compares a measured system with a flattering assumption.

Decision Framework

FactorThe question to answer before deploymentEffect on the margin above parity
Population exposedHow many decisions execute before a defect could plausibly be detected?Larger population, larger margin
Detection lagWhat signal would reveal a systematic defect, who watches it, and how often?Longer lag, larger margin
ReversibilityCan an affected outcome be reversed, and at what cost to whom?Irreversible, larger margin
RecourseCan an affected person obtain an explanation and a reconsideration from a named human?No recourse, larger margin
ConsentDid the affected person choose exposure to an automated decision?No choice, larger margin
Benchmark qualityIs the human comparison measured on the same definition, or assumed?Assumed, larger margin

Three governance tests sit alongside it. The authorisation test: produce the decision that permitted autonomous operation — dated, named, with the threshold it set; if no such document exists, an unowned decision is running in production. The halt test: who can stop the system, how long that takes, what it costs per day, and when the path was last exercised. The explanation test: take one real outcome and explain it to the person it affected, in language they can act on, within a stated time. An enterprise failing the third carries the full scrutiny premium and has bought nothing with it.

From Strategy to Execution

The immediate work is an inventory with teeth: for every system already operating without a human in the loop, produce the authorisation record, the current threshold, the monitoring evidence and the halt path. Gaps here are not documentation problems. They are decisions nobody made, executing anyway, and they can be closed within a quarter.

Over the medium term, two capabilities need funding. The first is the recourse channel, built into the product rather than bolted on as a complaints process, because it is the only lever that reduces the premium instead of absorbing it. The second is a costed view of failure: finance should be able to say what one bad automated outcome costs, including remediation, disclosure, stoppage and rework. A threshold expressed in accuracy while the consequence is expressed in dollars cannot be calibrated by anyone [Related article: What Is Your Quality Failure Costing You — and Can Finance Produce the Number?]. The economics are the ones Philip Crosby set out for quality: prevention is purchased before the event and failure is paid for after it, and only one of those prices is negotiable.

The long-term positioning is the ability to demonstrate, after a bad outcome, that the enterprise decided well with what it knew: specification, evidence, review, monitoring, a record, and a margin it can justify. It is the cheapest insurance available to an automating enterprise, and the only kind that cannot be assembled retrospectively.

Signals to Monitor

Agreement between human reviewers and the system approaching total, with no overrides recorded in a quarter — review has become ceremony and the control is already gone. Monitoring reports whose numbers never move. Aggregate accuracy reported without subgroup performance. A rising share of decisions taken without human involvement and no matching change in the risk register. Supplier accuracy assurances paired with liability caps that would not cover one bad quarter. Insurer questionnaires becoming more specific about automated decision-making.

Questions for the Leadership Team

  1. For each system now deciding without human involvement, can we produce the dated authorisation, the threshold approved, and the executive who owns it?
  2. What margin above human performance did we require — and if none, who chose the level we are running at?
  3. Do we know our human error rate on the same task, measured the same way, or are we comparing a tested system with an assumption?
  4. How long would a systematic defect operate before we detected it, and how many decisions is that?
  5. Can a person affected by one of our decisions obtain an explanation and a reconsideration, and how long does that take?
  6. If we had to halt one of these systems tomorrow, who authorises it, what does it cost per day, and when did we last rehearse it?

Closing Perspective

An enterprise does not get to choose whether its automated decisions are read as policy; a specification is an expression of intent and will be read as one. The only choice is whether they were made as policy — deliberately, with evidence, at a threshold somebody selected and can defend.

The question for the board is therefore not whether the models outperform people; that is the easy half. It is whether the enterprise could produce, without preparation, the decision that authorised deployment, the margin it demanded above human performance, and the reasoning behind that number. An organisation that can do so holds a defensible position even after a bad outcome. An organisation that cannot holds an indefensible one, including in the years when nothing goes wrong at all.


About EraNorth Insights
EraNorth Insights publishes practical analysis on strategy, projects, operations, transformation and decision intelligence for professional and organisational use. About EraNorth.