Program Governance

Three Scales, One Word

An enterprise that aggregates risk across delivery units is adding numbers produced by incompatible scales, and the portfolio figure that results is not a quantity.

EraNorth Insights · 14 min read

An enterprise running more than one delivery unit is performing arithmetic on incommensurable units when it aggregates risk, because the words on its scales are not the same numbers in different parts of the same organisation.

A chief executive asks what the enterprise's total risk exposure is. It is a reasonable question — arguably the most reasonable one a chief executive can ask — and an answer arrives within a fortnight: a portfolio heat map, a count of items in each band, an aggregate exposure figure, a note on the two units carrying the most red.

Consider what was done to produce it. Each delivery unit scored its risks against its own probability and impact scales, converted the scores into bands using its own thresholds, and passed the results upward. The portfolio office counted the bands and summed the exposures. No arithmetic error was made anywhere.

The number is still not a quantity. It adds figures produced on scales that do not measure the same thing: severe means one proportion of budget in one unit and another elsewhere; almost certain begins at four different likelihoods depending on the template a unit adopted; and at least one unit's impact weights are not percentages of anything, so they cannot be converted into the others' terms even in principle.

The question is not whether the enterprise should aggregate risk. It should. It is whether it can, and how anyone in the governance chain would find out that it cannot.

The Strategic Context

The discipline is aware of the problem. Its own teaching opens on this scenario: two managers in one organisation face the same exposure, one calls it moderate and monitors it, the other calls it high and escalates, and neither is wrong, because no common standard exists. The diagnosis is exact.

What follows is the interesting part. The same material publishes several worked scales — probability bands, impact bands, band-to-priority mappings — each internally sensible, each offered as an example to adopt, and no two agreeing. Elsewhere it instructs practitioners to calibrate scales to their own project rather than use generic wording: sound advice at project level, and a guarantee of divergence above it.

Those two instructions cannot both be honoured. Scales cannot be calibrated to each project and comparable across projects. Either choice costs something: local calibration buys sharper prioritisation inside a unit and destroys comparability between units, while standardisation buys comparability and blunts local judgement. Most enterprises have never made the choice. They took local calibration by default, because that is what the templates encourage, and then behave in the boardroom as though they had chosen standardisation.

That is the whole failure, and it is not a failure of diligence in any unit. Every unit has done what it was told.

What the Portfolio Heat Map Is Taken to Prove

That the units can be ranked against each other. Not on band evidence alone. A unit showing four red items and one showing a single item may carry identical exposure, or the reverse, depending on where each drew its boundaries. Ranking needs a common denominator; the heat map supplies none.

That a trend in the aggregate means something changed in the world. A unit that refreshes its template, adopts a group standard or is acquired shows a step change in its band profile without any underlying risk altering. Aggregated band counts are as sensitive to document lineage as to reality, and nothing in the reporting distinguishes the two.

That the reserve derived from the aggregate is sized against exposure. A reserve set from summed band scores is set from a figure whose components were computed on different bases. Which fund should carry the result, and who may release it, is a distinct question owned by [Related article: Two Funds, Two Authorities]; the point here is prior to it, and concerns whether the input figure means anything at all.

Reframing the Issue

This is a units problem, not a rigour problem. Nobody is careless. The enterprise is adding metres to acres and calling the total a size.

Consider an integrated pulp and paper operation carrying four capital projects at once: a recovery boiler rebuild, a woodyard upgrade, a converting line automation and a mill-wide control system replacement. The illustration is hypothetical; the mechanism is general. Each defines its top cost impact band as an overrun exceeding a stated proportion of its own budget, and the boiler rebuild's budget is an order of magnitude larger than the converting line's. An identical unplanned cost — the same money, leaving the same account — is negligible on one register and severe on the other.

Both roll into one mill-level view whose aggregate sums numbers computed on those two bases. The distortion runs one way: the largest projects, having the largest denominators, always report the mildest bands for a given absolute loss, so the risk picture under-weights the places where the money is. And the boiler rebuild is the project whose failure stops the mill — the genuinely enterprise-level exposure carries the most forgiving scale.

A multi-site private hospital group shows the same effect in a form harder to see. Each site calibrates its clinical impact scale to its own case mix, which is defensible and locally correct. Group assurance then reads one heat map in which major clinical impact is not one thing, never has been, and no document in the group states what would make it one.

Three Places the Units Come Apart

The likelihood bands do not start in the same place

Across published scales, the top likelihood band — labelled almost certain or very high — begins above four-fifths in one scheme, three-quarters in another, seven-tenths in a third, and in a fourth is not a band at all but a point value with no width. A risk assessed at three-quarters likelihood is therefore top band, second band, or undefined, according to which document its owner was handed.

The lower bands are worse, because they do not always tile. At least one published scale leaves a gap between the ceiling of one band and the floor of the next, and at least one score-to-priority mapping leaves scores in no zone at all. These are not exotic edge cases but ordinary values a competent assessor will produce and then place by judgement — the very thing the scale existed to remove.

What happens to items above the top band is a separate and serious matter, examined in [Related article: Where Do Your Near-Certainties Go?] rather than here. The point here is narrower: the band's floor moves between documents, and nothing above unit level records which floor was used.

The impact bands are anchored to a moving denominator

Every cost impact scale in general use is denominated as a proportion of the project's own budget, and the proportions differ. The top band begins above a fifth of budget in one scheme, a quarter in another, two-fifths in a third. An overrun of roughly a quarter of budget is thus the most severe outcome the scale can express, or a mid-table entry, depending on the template.

Because the denominator is the unit's own budget, the scale is inconsistent between units and within one unit over time. A programme rebaselined upward becomes less risky on its own register overnight, with no risk having changed.

A fourth arrangement drops percentages altogether and assigns geometric weights to impact bands — values roughly doubling at each step, anchored to nothing external. Within a single project this is defensible and arguably superior, because it captures how severity is actually perceived. Across an enterprise it is the hardest case: no exchange rate exists between those weights and a percentage of budget, not even a poor one, because the two are not measuring in the same dimension.

The escalation trigger is denominated in currency and nothing else is

Governance frameworks set escalation thresholds in absolute money and elapsed months: above a stated sum the programme director is engaged, above a larger sum the board. This is sensible, and it is the only place in the apparatus where a real, comparable unit appears.

It is also inconsistent with everything feeding it. Impact is scored as a proportion of budget; escalation is triggered by an absolute figure. The same loss therefore reaches the board from a small unit while sitting unescalated inside a large one, and the board sees escalations weighted towards the enterprise's least consequential activities. Whether escalation should be triggered by a single risk's size at all, and whether the portfolio-offsetting argument justifying that design survives scrutiny, is the subject of [Related article: Escalation Sized by One Risk]; what matters here is that the trigger and the scale feeding it are denominated in different things.

A note on scope. A separate objection to these matrices — that ordinal ranks are multiplied as though cardinal — is taken up in ERANORTH's treatment of tender evaluation, where the same defect appears in scored bids. This argument does not rest on it. Even if every multiplication here were impeccable, the scales would still not be the same instrument, and the aggregate still would not be a quantity.

Decision Framework

The scale reconciliation. Run it once, then annually, from the portfolio office. Four steps, one rule.

Step one — collect what is actually in use. Obtain the probability scale, every impact scale, the band mapping and the escalation thresholds in force in each material delivery unit — in use, not as published. Where a unit cannot produce them within a week, record that: an unstated scale is a finding in its own right.

Step two — tabulate the boundaries in absolute terms. For each unit, write the numeric floor of every likelihood band, then convert each impact boundary from a percentage into money using that unit's own budget. Present the second column as one line across the portfolio. The spread between the highest and lowest absolute value of severe is the number the board needs, and it is usually larger than anyone in the room expects.

Step three — run two probe events. Define two events in absolute terms only: a loss of a stated sum at a stated likelihood, and a rare event with a very large loss. Have every unit score both under its own scale and return the band. The spread of bands returned for one defined event is the enterprise's measure of incommensurability, reportable in a single line.

Step four — decide the denomination, at board level. Three positions are coherent. A single enterprise scale, impact denominated in absolute money, local scales abolished. Local scales retained for prioritisation inside a unit, with all reporting above unit level in absolute money and band language forbidden. Or local scales retained with a published conversion table, which obliges someone to write the conversion down and defend it. The second usually costs the least to adopt and is the most honest. The position not available is the current one: the second in practice and the first in presentation.

The rule. No aggregate risk figure reaches a governing body without stating how many distinct scales it was computed across. If the answer is more than one, the figure is a count of words and must be labelled as one.

From Strategy to Execution

Immediate. Ask two units of very different size to score the same defined loss and report the band each returns. It takes an afternoon and either demonstrates the problem or retires the concern. Either way, the answer belongs in the next governance pack.

Medium term. Move portfolio reporting to absolute money for impact, leaving local scales untouched for prioritisation. This is deliberately the smaller change: no unit abandons a calibration it has good reasons for, and band language disappears from the only level at which it is invalid. In parallel, put the scale in force into the header of every risk register, so the basis travels with the data.

Long term. Treat the choice between local calibration and enterprise comparability as a standing governance decision with an owner and a review date, rather than an artefact of which template a unit inherited. Where the enterprise grows by acquisition, add scale reconciliation to integration: an acquired unit arrives with its own definition of severe and will otherwise keep it indefinitely.

Signals to Monitor

A unit's red count moving sharply after a template refresh or a rebaseline, with no corresponding event. Two units reporting the same absolute exposure in different bands. Escalations arriving disproportionately from the smallest units. A heat map that survived an acquisition or a major budget revision unchanged. Risk scales older than the budgets they are proportional to. Assurance reports comparing units by band. And the plainest indicator: no document in the enterprise stating what its impact scale is denominated in.

Questions for the Leadership Team

  1. How many distinct probability and impact scales are in force across our delivery units, and who could tell us within a week?
  2. In money rather than percentages, what is the lowest and highest absolute loss any of our units classifies in its most severe band?
  3. If we gave every unit the same defined loss at the same likelihood, how many different bands would come back?
  4. When our portfolio reserve was last set, what basis of measurement were the aggregated inputs computed on?
  5. Which of our largest programmes would report a lower severity than a small project for an identical absolute cost, and does the board know?
  6. Who owns the choice between locally calibrated scales and enterprise-comparable ones, and when was it last reviewed rather than inherited?

Closing Perspective

The instrument is not broken. A probability and impact scale calibrated to one project, used to order that project's risks by attention, does its job well, and the ordering is useful even where the numbers are not quantities. The failure is one of promotion. An instrument built for ranking within a unit has been lifted into a role — aggregation across an enterprise, reserve setting, capital allocation, board assurance — it was never built to perform, and nothing in the discipline marks where its validity ends.

That leaves a responsibility, and it sits above the delivery units rather than inside them. No project manager can fix this from a project. The enterprise must decide what its risk aggregate is denominated in, say so in writing, and accept the cost of whichever answer it chooses. Until it does, the number the board receives each quarter is not a measure of exposure. It is a tally of vocabulary, and it will be most reassuring about the units that matter most.


About EraNorth Insights
EraNorth Insights publishes practical analysis on strategy, projects, operations, transformation and decision intelligence for professional and organisational use. About EraNorth.