Business

How Much of Your Pilot's Success Was Bought by Its Sponsor?

A pilot that succeeded because a senior sponsor made it succeed proves nothing about scale. Three clauses separate admissible evidence from an expensive rehearsal.

Kevin Jogin · 27 Aug 2026 · 11 min read

A result produced by the person who needed it is not evidence of anything except that person's determination.

Somewhere in your enterprise a pilot has just succeeded. It cleared its gate, the measured outcome beat the threshold, and the sponsor is asking for scale funding. The case is stronger than most because it rests on a real deployment with real users and a number that was observed rather than modelled.

Before you fund it, put a question to the evidence itself. It has three clauses, and each must be answered separately. Would this result have occurred without the person who wanted it? At the price the business case assumes? More than once?

Most internal pilots fail at least one clause. The difficulty is that a contaminated pilot and a clean one produce the same number, and the contamination lives in conditions nobody recorded — because at the time they were not conditions, they were simply how the work got done.

The Strategic Context

Early-stage ventures have a well-understood version of this problem. A product bought only by the founder's friends and family has not been validated, because those purchases were acts of relationship rather than acts of demand. The transaction happened; the evidence did not. Every experienced investor knows to discount it.

The enterprise analogue attracts far less scrutiny and costs considerably more, because a large organisation has more ways to buy its own result than any founder does. It can instruct a business unit to adopt. It can second its strongest operators to a small trial. It can suspend a procurement rule for one site. It can ask a supplier for accommodation that will be quietly recovered on the next contract. It can absorb the cost centrally and show the receiving function a clean margin. None of this is misconduct. Most of it is ordinary sponsorship, and every item of it is a subsidy the scaled version will not receive.

The enterprise's friends and family are its own reporting lines. A captive internal customer, a partner doing a favour, a unit adopting under instruction — each produces adoption without producing evidence of demand, and each is invisible in a results pack that reports the outcome rather than the conditions.

What a Successful Pilot Does Not Prove

A pilot proves that the thing can work. That is genuinely valuable and it is a smaller claim than most pilot reports make. What it does not establish is whether the thing works when the conditions that produced it are withdrawn, which is precisely what happens at scale.

Consider a hypothetical rail infrastructure operator trialling condition-based maintenance at a single depot. The result is strong: unplanned failures fall, intervention costs drop, the depot manager is an advocate. What the report does not say is that the depot was chosen because its manager was interested, that two of the operator's four experienced reliability engineers were assigned to it for the duration, that a data quality problem was resolved by an analyst working evenings, and that the sponsoring executive personally cleared three procurement exceptions in eleven weeks. At scale there is one reliability engineer per region, no evening analyst, and a procurement process that takes its normal course.

The pilot did not measure the capability. It measured the capability plus the sponsor, and reported the sum. A pilot report that records outcomes without recording conditions is not a weak piece of evidence; it is an unreadable one, because the number cannot be decomposed after the fact.

Reframing the Issue

Stop treating a pilot as a demonstration and start treating it as a controlled test with a stated hypothesis and a stated set of conditions. A demonstration is designed to succeed and its sponsor's involvement is a feature. A test is designed to discriminate, and its sponsor's involvement is a variable that must be declared, bounded and — at some point before the funding decision — removed.

The practical form of that reframing is a withdrawal period. The pilot runs, the result is achieved, and then the extraordinary inputs are taken away while measurement continues. What survives the withdrawal is the evidence. What does not survive was the subsidy, and knowing its size is worth more than the original result.

The Price Nobody Was Charged

The second clause is commercial rather than behavioural, and it is the one most often skipped. A pilot delivered to an internal customer at no charge, or at a transfer price set to encourage participation, tells you that the capability is usable. It tells you nothing about whether anyone will fund it.

This matters because the business case for scale almost always assumes a price: a chargeback rate, a cost-per-unit, a headcount reduction, an internal tariff a division has agreed to pay. If that price was never tested — if the pilot site received the service free, or the central budget absorbed the implementation — then the case's most load-bearing assumption is the one thing the pilot did not examine.

The same logic applies to a hypothetical professional services firm testing a new advisory offering with a long-standing client. The engagement was delivered, the client was satisfied, and a partner with a fifteen-year relationship discounted heavily to get it away. The offering has been proven deliverable and has not been proven saleable, and the firm will not discover the difference until a partner without that relationship attempts it at list price.

There is a further asymmetry the case rarely records. A receiving site that adopted early usually gave something up to do so, and what it gave up is the cell most business cases leave blank [Related article: The Quadrant Your Business Case Never Writes]. When that site is willing and the sponsor is present, the loss is absorbed quietly. At scale it is absorbed by units who volunteered for nothing.

Once Is Not a Process

The third clause is repetition, and it is the cheapest to satisfy and the most frequently waived under schedule pressure.

A single successful cycle demonstrates that a capable team, given attention, can produce the outcome once. A process is something that produces the outcome again, with different people, at a different site, without the original team present. The distinction is the whole of the scaling question, and the enterprise version of it is not a matter of counting transactions but of counting independent repetitions under progressively normal conditions.

The sequence worth insisting on is short. The first repetition uses the same team at a different site — testing whether the result depended on local conditions. The second uses a different team at a different site — testing whether it depended on the people. Somewhere in that sequence the sponsor stops attending, and the effect of that absence is measured rather than assumed. An enterprise that cannot afford three cycles before committing capital to a full rollout has decided that speed is worth more than knowing, which is a legitimate decision and should be recorded as one.

When the Result Disappoints, the Reflex Is Usually Wrong

Assume the scale-up proceeds and underperforms. The reflex response in most enterprises is to add execution rigour: tighter governance, more frequent reporting, a stronger delivery lead, a recovery plan. This is the response the organisation is best equipped to produce, and it is the correct response to exactly one class of failure — the class in which the intent was sound and the delivery was poor.

The other classes route elsewhere, and adding rigour to them makes the enterprise more efficient at doing the wrong thing. Persistent heavy discounting is not a sales-discipline problem; it is the market reporting that it cannot distinguish your offering from the alternatives, which is a positioning failure. Attrition in the roles the change depends on is not a management problem; it usually means the capability was defined after the structure was set, so people were hired against a problem nobody had specified. Scale that produces volume without profit is not an efficiency problem; it means expansion preceded validation, which is a sequencing failure and cannot be corrected downstream.

Each of these has the same surface appearance — the programme is not delivering — and each requires an intervention at a different layer. Diagnosing the layer is the work. Routing every symptom to the delivery layer is what makes underperforming programmes expensive rather than merely disappointing.

Decision Framework

Two instruments, applied at different moments.

The admissibility test, applied before scale funding. For the pilot result on the table, answer in writing: what extraordinary inputs were present and what did they cost; what price was charged and by whom it was borne; how many independent repetitions occurred and under what conditions. A result that cannot answer all three is recorded as promising rather than proven — the point of the test is not to block the investment but to prevent the enterprise from believing it has evidence it does not have.

The routing table, applied when results disappoint.

Symptom at scaleLayer that has usually failedWhat adding execution rigour achieves
Sustained heavy discounting to closePositioning — the market cannot distinguish the offeringFaster, better-governed discounting
Loss of the people the change depends onCapability definition — roles specified after the structureMore rigorous recruitment against the wrong specification
Growth in volume without growth in profitSequencing — expansion preceded validationEfficient expansion of an unprofitable unit
Adoption only where a sponsor is presentEvidence — the pilot measured sponsorshipA larger, better-reported dependency on sponsors
Delivery late against a sound designExecution — the one case the reflex fitsThe intended correction

Where the routing points to positioning, the correction is not a campaign. It is a change in what every function stops and starts doing, which is a substantially larger commitment than most enterprises expect [Related article: When You Chose a Position, Did Every Function Change — or Only Marketing?].

The expansion trigger. Write the condition for scaling before the pilot result exists, state it as a threshold with a named owner, and do not set it as a number this article or anyone else supplies — the right threshold is specific to your economics, your capacity and your alternatives. What matters is that it is written in advance, because a trigger written afterwards is a rationalisation, and a trigger set by the sponsor is not a control. It is equally worth recording why the expansion is wanted, since a scale-up pursued to rescue a struggling core is a different decision from one pursued from strength [Related article: Are You Expanding From Strength, or Gambling on Rescue?].

From Strategy to Execution

Immediately, take the pilot closest to a scale decision and reconstruct its conditions: who attended, what was expedited, what was absorbed centrally, who was seconded, what was not charged. Do this with the pilot team rather than to them, and frame it as protecting their result from being misread. The reconstruction usually takes a day and frequently changes the funding conversation.

Over the next two to three quarters, change the pilot template so that conditions are recorded prospectively alongside outcomes, and add a withdrawal period as a standard stage before any scale gate. Both changes are cheap. Neither survives without a governance body prepared to accept "promising, not proven" as a legitimate gate outcome, which is the actual reform.

Over the longer term, the capability to build is diagnostic routing — the organisational habit of asking which layer failed before deciding what to do about it. This is harder than it sounds, because the delivery layer is the only one most governance forums are constituted to discuss, and routing a failure to positioning or sequencing means telling an executive committee that the problem is upstream of the programme it commissioned.

Signals to Monitor

  • Adoption that tracks the sponsor's calendar. If uptake rises during site visits and falls afterwards, the dependency is measurable.
  • The gap between pilot unit cost and scaled unit cost, decomposed rather than averaged. Unexplained gaps are the subsidy, arriving late.
  • Exceptions granted during the pilot — procurement, resourcing, policy — counted and priced.
  • Whether any repetition has been run without the original team. If not, the process claim is untested however good the result.
  • The proportion of failure diagnoses that route to the delivery layer. If it approaches all of them, the routing discipline does not exist.
  • Discount depth on a scaled offering, read as a positioning signal rather than a sales-performance one.

Questions for the Leadership Team

  1. For the pilot we are about to scale, what extraordinary inputs produced the result, and what will they cost when withdrawn?
  2. What price was charged, to whom, and would that party have paid it without the sponsor's involvement?
  3. How many independent repetitions have we run, and did the sponsor attend all of them?
  4. Was the expansion trigger written before the result existed, and by someone other than the person who wanted the result?
  5. When our last major programme underperformed, which layer did we intervene in — and was that the layer that failed?
  6. What would it take for this organisation to record a pilot as promising but unproven, and has it ever done so?

Closing Perspective

The uncomfortable feature of sponsor-purchased evidence is that it is produced by exactly the behaviour enterprises reward. A determined executive who clears obstacles, secures the best people and personally resolves what the process cannot is doing what good sponsors do, and the enterprise is right to value it. The error is in what happens next: the result is recorded as a property of the capability rather than a property of the sponsorship, and funding is committed on that reading.

The correction is not to weaken sponsorship. It is to separate two things a pilot produces — a demonstration that the capability can work, and an estimate of what it takes to make it work — and to recognise that only the second predicts anything about scale. An enterprise that measures both will occasionally decline to fund something that succeeded, which feels like a loss and is usually the most valuable decision available.

The question survives every methodology, every gate design and every reporting format, and it is worth asking of anything about to receive serious money: if the person who wanted this result were removed tomorrow, what part of it would remain?


About the author
Kevin Jogin is Founder & Principal Advisor at EraNorth. Meet the Founder.