Enterprise Transformation

Volunteers Are Not a Sample

Your pilot proved the system works. It was designed and tested by the people who wanted it, which is why it cannot tell you whether anyone else will use it.

EraNorth Insights · 30 Aug 2026 · 14 min read

A pilot run by the willing can prove a system works. It cannot tell you whether it will be used, because the people who would not use it were never in it.

Two statements about the same programme are both true. The pilot succeeded: sites reported the new system as an improvement, defects were found early and fixed, and the business case survived its second gate on that evidence. The rollout failed: adoption stalled at forty per cent, workarounds spread, and eighteen months later the organisation is running two processes at a cost neither budget carries.

The usual explanation is change management — under-communicated, under-trained, under-sponsored. Sometimes that is right. Often it is a diagnosis chosen because it implies a remedy the organisation knows how to buy.

A simpler and less comfortable explanation is available first. The pilot answered a different question from the one the rollout asked: it established that the system works for people who wanted it, and could establish nothing about people who did not, because none were in it.

That is not a criticism of anyone's execution. But an executive reading a pilot report is being handed evidence about a population, and the population was selected twice, by two mechanisms nobody chose deliberately.

The Strategic Context

Enterprise systems are funded in tranches, and each tranche is released on evidence. That makes the quality of the evidence a capital-allocation question rather than a programme-reporting one.

The evidence that authorises the largest commitments in most transformation programmes comes from pilots and from design workshops. Both are populated by people who were available and interested. The commitment they authorise falls on everybody — including the sites with older equipment, the teams already at capacity, the regions with a different customer mix, and the people whose current workaround is quietly better than the new process for their particular case.

An organisation that understands this can still run a voluntary pilot. It simply reads the result correctly: as a demonstration of feasibility, not as a forecast of adoption. The failure mode is not running the pilot. It is treating its output as a sample when it is a self-selection.

What a Successful Pilot Proves

The habit worth challenging is the conflation of two questions that a pilot report answers as though they were one.

Does the thing work? Can it be configured, integrated, taught and operated; does it survive contact with real data and real volumes; does it break in ways that can be fixed. A pilot with enthusiastic participants answers this well. Enthusiasts are, if anything, harder testers than the indifferent — they push a system further and report defects rather than working around them.

Will it be used? Will people who did not ask for it, who have other pressures, whose current method works well enough, and who bear the cost of the transition without receiving its benefit, adopt it and keep adopting it after the programme team has gone.

The first question is about the artefact. The second is about the organisation. Only the second determines whether the investment pays, and a voluntary pilot is structurally incapable of answering it. Note that this is a loss at the point of collection, not in transmission — what gets filtered out as information travels up through reporting layers is a separate and equally expensive problem [Related article: Whose Knowledge Does Your Governance System Actually Hear?].

Reframing the Issue: Selection Happens Twice

The first selection happens at design, and it is usually invisible because it looks like scheduling.

An unusually candid account of this comes from a trade-journal report of a development bank's replacement of its procurement processes with a web-based platform. The organisation is named because its own source names it. The team leader — a senior contracts officer — needed a cross-functional group drawn from an organisation of ten thousand people. The people whose input the design required sat outside her business unit and travelled frequently. Her account of the remedy is that the team worked with whoever was available — and, in her own words, "with the people whom we knew were the most supportive of the project" (Atkinson, Purchasing, 2005).

That is a competent manager solving a real scheduling problem. It is also a selection rule. The design will fit the people who could attend and who liked the idea.

The second selection happens at trial. A consultant retained on the same programme describes the rollout: the bank "asked for volunteer departments and regions to try the new system out, so glitches could be worked out before it was rolled out worldwide". Again sensible, and again a selection.

Then the two compound. The same team leader reports that "to date, the system is being well-received by users", and that mandatory implementation would follow in about six months. The report adds that early results looked promising — supplier competition improving as registration became easier, compliance simpler to monitor, users appreciating accessible reporting.

Nothing in that account is dishonest, and none of it is presented here as an outcome; the source reports intentions and early impressions in 2005 and evidences neither, and no outcome is claimed [FACT CHECK REQUIRED — outcome not established by the source]. The point is what the favourable early reception is evidence of. It is evidence about volunteers, gathered by the person who selected the design group, in a programme whose participants had already been chosen for support. That is exactly the reception a self-selected population produces, whatever the system is like.

What the Willing Cannot Tell You

A voluntary population differs from the whole population in ways that are specific and predictable, which is what makes the omission correctable.

Configuration. Volunteers tend to be sites with newer equipment, cleaner data and fewer local variations — the sites for which the change looks easy.

Capacity. A team volunteers when it has someone to spare. The teams that decide the adoption rate do not.

Connectivity and geography. Central or well-connected sites volunteer more readily than remote ones, and remote sites are where infrastructure assumptions fail.

Incentive. Volunteers frequently gain from the change. The population includes people who lose from it — an autonomy, a relationship, a piece of discretion — and their behaviour is the adoption curve.

Competence mix. Pilot teams are staffed with the best available people, which conceals how much the design depends on skill.

Each is knowable in advance, which is why a coverage statement is cheap. The absences are not mysterious; they are unrecorded.

The Misreading This Argument Must Not License

There is a version of this argument that becomes an excuse, and it must be refused explicitly.

The absence of a population from your evidence says nothing whatever about that population's motives. A site that declined the pilot is not thereby resistant, obstructive or attached to the old order. It may have been mid-audit, short-staffed, or already running something that works. An objection raised late by a group that was never asked early is information about the design, not proof of politics.

This matters because the adjacent argument — that opposition to a transformation is structurally better organised than support — is easily misapplied in exactly this way. That argument concerns the volume and organisation of opposition, never its informational content, and it never licenses discounting a well-founded objection [Related article: Why the Old Order Fights Harder]. A leader who uses selection bias to dismiss the excluded has taken the wrong half of the lesson.

Nor is this an argument for a stakeholder instrument. Classifying people by power, interest or exposure is a different discipline with its own literature and its own failure modes [Related article: Your Stakeholder Map Measures Attention, Not Exposure]. The coverage question is narrower and more mechanical: who is in the evidence, and who is not.

Two Illustrations

Both are hypothetical.

A state-wide public library consortium builds a shared catalogue. Membership is voluntary, and eleven of forty-two services join the first release. They are, predictably, the eleven with modern library management systems, a systems librarian on staff and a council willing to fund the migration. The release is a success by every measure the programme has. The thirty-one that follow have older systems, no dedicated staff and different cataloguing conventions — and the shared catalogue was designed by people who had never had to accommodate any of that. The consortium is now funding an integration layer it did not budget for, to solve problems its pilot could not have surfaced.

A banner group of independently owned pharmacies rolls out a common dispensing and ordering platform. The design group is drawn from the owners who attend the national conference — engaged, commercially confident, running larger stores with several staff. The platform they help design assumes a dispensary with someone available to manage exceptions. Two-thirds of the group are single-pharmacist stores where the exception queue lands on the person already serving customers. Adoption is not resisted; it is simply impossible during trading hours, and it shows up in the numbers as indifference.

Decision Framework

Write a coverage statement for every pilot and every design group, before the work starts, in half a page.

FieldWhat it records
RepresentedWhich populations are in the group, and how they came to be in it
AbsentWhich are not, named specifically rather than as "other sites"
What the absent do differentlyEquipment, volume, staffing, data quality, incentive
Evidence still requiredWhat would have to be shown before the absence stops mattering

Three rules follow.

Recruit at least one site that did not want it. Not a hostile site — a representative one that would not have volunteered. This costs a negotiation at the start and is the single highest-value change available to most programmes. It is also the one that will be resisted internally, because it makes the pilot harder and its results less flattering.

Separate the two questions in the reporting. A pilot report should state its feasibility findings and its adoption findings in different sections, and should say plainly that the second is not generalisable where the population was self-selected. That sentence costs nothing and prevents a tranche being released on the wrong evidence.

Set the mandatory-implementation test before the voluntary phase ends. Ask what will be different when adoption stops being a choice, and what evidence you will have about that condition. If the answer is none, the programme is planning to discover its real adoption problem after the money is committed.

Two related questions belong elsewhere and are deliberately not answered here. Who is entitled to declare the investment a success, against what criterion and on what date, is its own argument [Related article: Who Is Entitled to Say It Worked?]. And whether evidence that something worked is evidence at all, when it comes from people who wanted it, is a question a risk register raises in the same shape [Related article: Which of Your Risk Responses Changes the Probability?] — as does any instrument used to judge a person [Related article: The Model You Score Is the Model You Get].

From Strategy to Execution

Immediately. For every pilot now running, write the coverage statement retrospectively. It takes an hour, and the list of absent populations is usually the programme's forward risk list.

At the next gate. Require the coverage statement as a condition of releasing the next tranche, and make it the sponsor's document rather than the programme manager's. A programme manager has every incentive to report a successful pilot; a sponsor carrying the benefit has an incentive to know what the pilot did not test.

Over the longer term. Build a standing list of the sites, teams or regions that represent the organisation's hard cases, and use them deliberately. Most enterprises can name theirs in ten minutes and have never written them down.

One caution on generalisation. Assuming a finding transfers from a self-selected group to a whole population is a cousin of assuming a method transfers between contexts it was not built for, and an organisation prone to one is usually prone to the other [Related article: Why "The Principles Apply to Any Project" Is Only Half True].

Signals to Monitor

  • Adoption curves that flatten between the volunteer wave and the mandated wave. The gap between them is the size of the evidence problem.
  • Workarounds appearing in populations that were not represented in design. These are design findings arriving late, not compliance failures.
  • Support tickets concentrated in a configuration nobody piloted.
  • Requests for local exceptions. Each is a statement that the design did not contemplate that site, and the pattern of requests maps the coverage gap precisely.
  • Programme reporting that quotes user satisfaction without stating who the users were. Satisfaction figures from a volunteer cohort are the most misleading number in a transformation pack, because they are both true and irrelevant.
  • Benefit measurement deferred. The bank in the source material intended to build task-duration benchmarks once the work was automated, and had not begun. Deferred measurement keeps favourable early evidence as the only evidence.

Questions for the Leadership Team

  1. For our largest current programme, who was in the design group, and how were they selected?
  2. Which populations are absent from our pilot evidence, and what do they do differently?
  3. What would we have to see before we would believe the pilot result applies to the whole organisation?
  4. Have we asked one reluctant site to participate — and if not, what stopped us?
  5. When adoption becomes mandatory, what evidence will we have about the people for whom it was never voluntary?
  6. Where a site has objected late, did anyone check whether it was ever asked early?

Closing Perspective

Every enterprise programme reaches a moment where a favourable result must be interpreted. The temptation is to treat it as a forecast, because it was expensive to obtain and because the alternative is to explain to a board that the evidence does not yet exist.

The honest reading is narrower and more useful. A voluntary pilot tells you the thing can work, which is worth knowing and is not nothing. It tells you almost nothing about whether it will be used, because the question of use is a question about people who were not there.

The remedy is not more rigour, more governance or a larger pilot. It is one reluctant site, one page listing who is absent, and the discipline to say in a board paper which of the two questions the evidence has actually answered.


About EraNorth Insights
EraNorth Insights publishes practical analysis on strategy, projects, operations, transformation and decision intelligence for professional and organisational use. About EraNorth.