A weighting is a price the organisation has put on something, set months before anyone can see what the market is offering.
An evaluation panel sits down with four submissions and a scoring matrix. The weightings are concealed while compliance is marked, so that nobody can tune a score toward a favoured bidder. Each criterion is assessed on its own. Only when the marking is complete are the weightings revealed and the totals calculated. The winner emerges from arithmetic, and the paper that follows can show exactly how.
It is the most disciplined room in the organisation, and by the time it sits, the decision it believes it is making has largely been made. The weightings were set in evaluation planning, before a single offer existed. The mandatory criteria were fixed earlier still, and they had already removed submissions nobody in the room will ever read.
None of that is improper. Setting weights before bids arrive is exactly right; doing it afterwards is how evaluations are corrupted. The difficulty is that the discipline of the scoring room creates a strong impression that the judgement happened there, when the judgement happened months earlier, in a meeting with no minutes, attended by whoever was drafting the documents.
Every organisation of any size buys this way. Very few executives have ever read the weighting sheet for a contract they signed.
The Strategic Context
Competitive evaluation is capital allocation with a procedural costume on. The sums are frequently larger than the projects that occasioned them, the commitments run for years, and the resulting arrangement determines what the enterprise can and cannot do in that period.
It is also the point at which an organisation states, in numbers, what it values. Ten points here, four points there, a pass or fail on something else: that is a preference structure, and it will be applied literally by people acting in good faith. If it is wrong, they will still apply it, because a scoring model's virtue is precisely that it does not allow the panel to substitute its own view.
This article is about selecting between suppliers for something already approved. Whether the portfolio should have funded it at all, and whether a portfolio function chooses between initiatives or merely supervises them, is a different question [Related article: Is Your Portfolio Function Selecting, or Supervising?]. So is whether the decision to continue should be taken at the phase boundary at all [Related article: What a Stage Gate Is Actually For].
What the Panel Is Not Deciding
The habit worth challenging is the belief that objectivity lives in the marking. It does not. Marking is where consistency lives. Objectivity, if it exists anywhere in the process, lives in the criteria and their weights — and those are not marked, they are declared.
A standard method makes the structure visible. Sellers are shortlisted against mandatory criteria, and the rule is absolute: failure to meet any one of them eliminates the offer. Surviving offers are scored for compliance on a five-point scale from "meets or exceeds requirements fully" down to "has only minor elements of the requirements". Each criterion carries an importance weight — ten points where absence "would compromise key functionality", descending to two for something the organisation could cope without. Compliance scores are multiplied by weights and summed, and a perfect score is calculated as the sum of the weights times the highest possible compliance [SOURCE DETAILS REQUIRED].
Read that sequence again and notice how much of it is finished before anyone opens an envelope.
Reframing the Issue
A mandatory criterion is not a heavy weight. It is an infinite one. It does not influence the ranking; it determines the population from which a ranking can be drawn. An offer that would have won on every weighted criterion is gone, unread, because it lacked a certification, a reference class or a balance-sheet threshold.
That is often correct — some requirements genuinely are binary. But a pass-or-fail list receives a fraction of the scrutiny that weights receive, because it looks like housekeeping rather than judgement. It is worth asking of each mandatory item whether it is genuinely disqualifying, or merely important and easier to write as a gate.
Accreditation is the common example, and it deserves care. A certification is evidence that an organisation runs a system, which is not the same as evidence that it will perform. Why organisations pursue such credentials, and who the rating is really for, is examined elsewhere [Related article: Who Is Your Maturity Rating For?].
A weight is a statement about consequence. Look again at the language of the importance scale: the top weight is defined by what happens if the feature is absent. That is a risk statement wearing a scoring uniform — which means the weighting sheet is one of the few documents in the enterprise where somebody has priced the consequences of a supplier's shortfalls in advance, and almost nobody reads it that way.
Ordinal Scores, Cardinal Arithmetic
The compliance scale runs from five to one, and the model multiplies those numbers as though they were quantities. They are not. "Meets or exceeds requirements fully" is not five times "has only minor elements of the requirements", and the distance from four to five is not the distance from one to two. They are ranks with numerals attached.
Multiply ranks by weights, sum them, and the output looks like a measurement. Express it against a "perfect score" of one hundred and it looks like a percentage. An offer at 82 and an offer at 76 will be described as six points apart, and the six points will be discussed as though they meant something the underlying scale can support.
This is not an argument against scoring. It is an argument for treating the total as an ordering device rather than a magnitude — and for asking, when two offers finish close together, whether the arithmetic can bear the distinction being drawn from it. A four-point gap on a hundred-point scale built from five-point ranks is not a result; it is a tie the model has broken for you.
Two boundaries are worth marking. This is a claim about one scale inside one instrument, not a general argument that measures change species as they move through an organisation [Related article: From Margin to Milestones]. And it is not a method for testing whether a claim in a submission is achievable — that capability sits with the reader, not the model [Related article: Enough Technical Depth to Test the Answer].
The Safeguard That Protects the Smaller Bias
Concealing the weightings during marking is genuine protection against a real failure: a marker who knows that criterion three is worth ten points can move an offer up by nudging one score. Hiding the weights removes that.
What it does not touch is the larger question of what carries weight at all. And here the source material is instructive by contrast. The worked example in the lecture deck uses three criteria, unnamed, with no guidance on choosing them. The course handout that predates it lists what the criteria should cover: value for money and fitness for purpose, delivery and whole-life operating costs, ongoing support and warranty, probity, and risk identification and management capability [SOURCE DETAILS REQUIRED].
Those last two decide whether an arrangement survives its worst year. In many evaluation sheets they carry no weight at all — not because anyone rejected them, but because the sheet was adapted from the last one.
A scoring model can only ever answer the question it was built to ask: does this offer comply with what we wrote? Whether what we wrote was worth complying with is a different question, and a different instrument answers it [Related article: Quality Assurance Cannot Tell You the Specification Was Wrong].
Quality First, Then Price
The teaching source ends with a rule that most evaluations invert: establish the technical quality of the bid first, and compare prices afterwards, because it is easier to negotiate price than to negotiate quality.
The reasoning is commercial rather than moral. Price is the most movable term in any arrangement and the one a supplier can concede fastest. Quality, capability and capacity are structural. An evaluation that lets price into the room early anchors every subsequent judgement to it, and the organisation spends its remaining leverage on the term it could have moved anyway. What can still be obtained after the price is agreed, and what cannot, is its own argument [Related article: The Terms You Cannot Buy Back].
Two Illustrations
Both are hypothetical.
A state fire and emergency service replaces its computer-aided dispatch system. The weightings were set eighteen months before the tender closed, when the programme's concern was interoperability with neighbouring jurisdictions. Interoperability is a mandatory criterion; degraded-mode operation — what the system does when a data centre is unreachable during a fire event — carries a weight of four, because it was written as a technical feature rather than as a consequence. The winning offer scores well and cannot run standalone at a regional control room. Nobody marked it wrongly.
A commercial laundry and linen services contract is let by a hospital group. Hygiene compliance is mandatory, correctly. The weighted criteria are cost per kilogram, delivery reliability and account management. Surge capacity carries no criterion at all, because the specification described normal volumes, and normal volumes are what the previous contract described too. The failure arrives two winters later, in a week when demand rises by a third and the contract has no term that obliges anyone to meet it.
Decision Framework
Make a weighting record, and make it visible to whoever signs the contract. For each criterion:
| Field | What it captures |
|---|---|
| Decision it stands for | What real-world consequence this criterion is a proxy for |
| Weight and who set it | The number, the person, the date |
| Gate or scale | Mandatory, or weighted — and why it is one and not the other |
| Value of a point | What a one-point difference on this criterion means in operation |
Three tests apply before the tender is issued, not after.
The absence test. For each criterion, state what it costs the organisation if the supplier is weak on it. If the answer is negligible, the weight is too high. If the answer is severe and the weight is four, the sheet has misfiled a consequence as a feature.
The gate test. For each mandatory criterion, ask whether an otherwise outstanding offer that failed it should genuinely be discarded unread. Where the honest answer is no, it is a weighted criterion, not a gate.
The worst-year test. Score the offers against the year you do not expect: the surge, the outage, the industrial dispute, the recall. Most weighting sheets describe steady-state performance, and most contract failures do not occur in steady state.
A separate discipline, worth naming because it is easily conflated: recording who asserted something and when is the basis of a challenge protocol used elsewhere in delivery, and this article borrows the habit rather than the method [Related article: Which of Your Dependencies Are Real?].
From Strategy to Execution
Immediately. For any evaluation now in flight, read the weighting sheet at executive level before the tender closes. It takes twenty minutes and it is the last moment at which the decision can still be changed cheaply.
Within the year. Require a weighting record for every procurement above a materiality threshold, signed by the executive who will own the resulting arrangement rather than by the team running the process. Ownership of the criteria and ownership of the consequences should sit with the same person.
Structurally. Stop inheriting sheets. A weighting model copied from the last contract of similar size encodes the priorities of a different problem, and its greatest defect is invisible: the criteria that are missing leave no trace on the page.
Two adjacent arguments are worth reading beside this one. A priority label is a similar act of declaration and commits the organisation to more than it usually realises [Related article: What Does "Critical" Actually Commit You To?]. And every instrument of this kind was designed for a particular kind of work, which is worth knowing before it is applied to another [Related article: What Kind of Work Were These Instruments Built For?]. What an instrument assumes about a person, rather than about work, raises a related problem [Related article: The Model You Score Is the Model You Get].
Signals to Monitor
- Close finishes decided by weights. Where the top two offers finish within a few points, record whether the ranking would survive a one-point change on any single criterion. If it would not, the model chose, not the panel.
- Criteria with no operational owner. A criterion nobody in operations can explain is a criterion nobody set for a reason.
- Repeated sheets. The same weighting model appearing across contracts with different failure modes.
- Post-award variations clustering on unweighted matters. The things you did not score become the things you negotiate later, at a price.
- Mandatory lists growing. Gates accumulate because adding one is easier than defending a weight. A shortlist of two in a competitive market is usually a gate problem.
Questions for the Leadership Team
- For our three largest current arrangements, who set the weightings, when, and did anyone who now owns the outcome see them?
- Which criteria in our standard evaluation model would a supplier's failure actually hurt us on — and do those carry the highest weights?
- How many offers were eliminated by mandatory criteria in our last four tenders, and would we make the same call knowing what we know now?
- When two bids finish within five points, what do we do — and is that written down anywhere?
- What does our current model say about the year we do not expect?
Closing Perspective
A scoring model does not remove judgement from a decision. It relocates it, from a room full of people who will be accountable for the outcome to a document written months earlier by people who will not.
That relocation is defensible; the alternative is an evaluation nobody can audit. What is not defensible is leaving the document unread by the executive who will live with the arrangement for five years. The weighting sheet is short, and it is the most consequential thing the organisation writes about a supplier before the contract exists.
The decision is made when the weights are set. Everything after that is arithmetic doing what it was told.
About EraNorth Insights
EraNorth Insights publishes practical analysis on strategy, projects, operations, transformation and decision intelligence for professional and organisational use. About EraNorth.
