The most expensive data decisions in an enterprise are made years before anyone proposes a model, by people who were never told they were making them.
Consider a familiar sequence, offered as a hypothetical illustration rather than an account of any particular organisation. An enterprise with five operating units commissions a forecasting capability. The business case is sound, the sponsor is credible, the vendor is competent. Eighteen months later the initiative is quietly reclassified as a foundational data programme, its benefits deferred, its cost doubled.
The failure was not algorithmic. It was that the same commercial event — a customer accepted a variation, a job was reworked, a delivery was made short — had been recorded five different ways across five units over seven years. Not wrongly. Each unit's convention served its own reporting perfectly well. The conventions simply could not be reconciled, and nobody had ever held the authority to require that they be.
That is not a technology failure. It is a decision-rights failure, and it happened long before any model was contemplated. The people who created the condition were behaving rationally inside their mandates. Nobody's mandate included the enterprise.
The Strategic Context
Data is routinely described as an asset, and there is real substance to the claim: businesses whose competitive position rests on what they know about their customers and operations have tended to be valued differently from businesses whose position rests on what they own. But the claim carries a condition that gets dropped in the retelling. Data is an asset when it is joinable and when a decision depends on it. Records that cannot be reconciled across units are not an asset; they are five accounts of five different things, held at cost, carrying custody obligations.
Two related claims deserve correction before they reach a board paper. The first is that data does not depreciate. Behavioural and market data decays quickly, and the useful life of a record is set by how fast the world it describes changes. The second is subtler and more damaging: even where the record persists unchanged, its meaning does not. Systems are replaced, field definitions are revised, thresholds move, a category is split. The archive stays; what it means quietly changes underneath it.
The strategic consequence is that data collection is one of the few operating decisions that cannot be made retrospectively. Capital can be redeployed, people can be redeployed, a supplier can be replaced. A field that was never captured in 2023 cannot be captured in 2026. Whatever an enterprise decides about collection this year sets the boundary of what it can decide about analysis three years from now.
Why This Arrives Labelled as a Technology Problem
Because the symptom appears in a technical function. The unusable dataset is discovered by engineers, so the remedy is assumed to be engineering. Standardisation software can reconcile formats. It cannot reconcile two units that genuinely mean different things by the same word, because that requires someone with authority over both to rule on which meaning holds.
Because centralising storage is mistaken for centralising meaning. Consolidating five collections into one repository yields one repository containing five conventions. The reconciliation work is unchanged; it has only been moved, and it is now less visible because the folder structure suggests it has been done.
Because the pipeline is funded as a project rather than an obligation. Projects finish. Collection standards do not — they need an owner, a change-control path and a budget line for as long as the enterprise runs, because every system change, acquisition and process redesign is an opportunity for the convention to fracture again.
Reframing the Issue
The practical instrument is older and simpler than the technology conversation around it. Before any unit is asked to capture anything, five questions must have named answers.
What is being captured, defined precisely enough that a different unit could reproduce the definition without a conversation.
Why — which decision does this field inform, and whose? A field that informs no decision should not be collected. This is the question most often skipped, and skipping it is how enterprises accumulate large quantities of data they pay to hold, are obliged to protect, and never use.
How — the instrument, the timing, the tolerance, whether the field is mandatory, and what validation applies at the point of entry. Data quality is overwhelmingly determined here, not downstream.
Who — the named role accountable for the definition, and the role empowered to approve a deviation from it.
For how long, and who may change it — retention, and the authority required to alter a definition that other units depend on. A definition change made unilaterally is a change to a shared asset.
Answered properly, these five are not documentation. They are an allocation of authority, and they belong to the executive who owns the operating model, not to the function that stores the result.
Collection Is Paid For Locally and Realised Centrally
The economics of this problem explain most of its persistence.
Disciplined capture is a local cost. It slows a transaction, adds a mandatory field, requires training, and irritates people who already have targets. The benefit is realised somewhere else and later — in a pricing model, a demand forecast, a group-level view of customer profitability. A unit leader who tightens capture bears a certain cost this quarter for an uncertain benefit that will be credited to someone else next year.
This is a benefits-ownership failure in its clearest form, and it is not solved by insisting on compliance. It is solved by changing who pays and who is measured. That means funding the collection standard centrally rather than expecting units to absorb it; naming the accountable executive for each enterprise data domain; and, most usefully, closing the feedback loop. The units that collect almost never see what their capture quality did to a downstream decision. Returning that information — this forecast error, this rework, this failed reconciliation, traced to this convention — converts an abstract instruction into a visible consequence, which is the only thing that reliably changes operating behaviour.
There is a portfolio question here as well, and it usually goes unasked. Every field an enterprise collects has a cost, a custody obligation and an opportunity cost against the fields it is not collecting. The most valuable output of a collection review is frequently a list of what to stop capturing, which then funds the capture of what actually informs decisions.
The Definition Is the Asset; the Record Is Only Storage
If the conventions are agreed, the records become comparable. If they are not, volume makes the problem larger rather than smaller. A larger inconsistent dataset is worse than a smaller consistent one: the inconsistency takes longer to find and more to remediate.
Two temptations tend to appear at this point, both presented as improvements to data quality.
The first is enrichment: adding attributes to make a model more accurate. Some of these additions are ordinary. Others — inferred or recorded characteristics of a person, including sensitive ones — are not data-quality improvements at all. They are decisions about what the enterprise is prepared to know about an individual and, downstream, what it is prepared to act on. Any recommendation to enrich a dataset with sensitive personal attributes should be escalated out of the technical workstream and decided as a commercial and ethical question at executive level [Related article: What You May Know About a Customer, and What You May Price On It].
The second is back-filling by estimation. Where fields are missing, statistical estimation can produce a complete table. It cannot produce information that was never captured. It imports an assumption about the missing population, and whoever chooses the method has made a policy choice about people whose records were incomplete — often not at random, and often for reasons correlated with the very outcome being modelled.
Australian privacy obligations bear directly on all of this. Collection of solicited personal information is constrained to what is reasonably necessary for the entity's functions or activities, and entities carry notification obligations at the point of collection — the requirements in Australian Privacy Principles 3 and 5 are the relevant starting point and require professional verification in every specific case [FACT CHECK REQUIRED]. ERANORTH is not a law firm and this is not legal advice. The commercial point stands independently: a "collect everything, decide later" posture and a purpose-bound collection standard are different operating models, and an enterprise should know which one it is running.
The Missing-Data Choice Is a Business Decision, Not a Statistical One
Three responses are available when a field is incomplete: collect it properly from now on, estimate it, or build only on the features that are complete. Each has a different owner, a different cost and a different failure mode.
| Response | What it really costs | Who should decide | Principal risk |
|---|---|---|---|
| Collect it properly going forward | Process change, training, transaction friction; benefit delayed by the collection period | Operating-unit leadership, with the data domain owner | Delay; the decision that needed the data arrives first |
| Estimate the missing values | Modelling effort; an embedded assumption about the incomplete population | Risk and the accountable business owner jointly, never the modelling team alone | Confident output built on an unstated assumption |
| Restrict the model to complete features | Reduced accuracy, narrower application | The sponsor, explicitly, on the record | Quiet scope reduction presented as a technical constraint |
Deciding which of the three applies is not the modeller's call. Whether the resulting capability is fit to deploy at all, and what should shelve it, is a separate governance question with its own benchmark [Related article: When Is a Model Good Enough to Deploy — and What Shelves It?].
Decision Framework
For each enterprise data domain, place it in one of three states and apply the corresponding rule.
Enterprise-defined. A single definition, a named owner, change control that binds every unit. Models may be funded on it.
Locally defined but mapped. Units differ, but a documented, tested mapping exists and someone owns it. Models may be funded with the mapping cost stated in the business case and the mapping's failure modes disclosed.
Unmapped. No agreed definition, no mapping. No model is funded on this domain until mapping is scoped, costed and owned. This rule prevents the most common failure in the sequence: a capability approved on the assumption that reconciliation is a task rather than a negotiation.
Three tests then govern any proposal to collect something new. Does the field inform a named decision owned by a named person? Could another unit reproduce the definition without a conversation? If we captured it for two years and the intended use never eventuated, would we still be comfortable holding it? A proposal that cannot answer the third is asking the enterprise to accept an obligation in exchange for an option nobody has priced.
From Strategy to Execution
Immediately, produce two lists: the enterprise data domains with their current state and their owner, and the fields currently collected that inform no decision. Both are typically a fortnight of work. The second usually pays for the first.
Over the medium term, build the authority rather than the platform. That means a definitions owner per domain with the standing to bind units, field-definition changes routed through the change-control process that already governs shared systems, and collection standards written into acquisition integration plans — an acquisition is where conventions fracture fastest and where the reconciliation cost is least likely to have been priced in the deal.
For long-term positioning, recognise that this compounds in both directions. An enterprise that can join its own records across units can answer questions its competitors cannot, and can automate decisions its competitors cannot evidence. The question of which decisions it should then hand to a machine, and who authorised that, is a distinct governance matter [Related article: What Has Your Enterprise Already Authorised a Machine to Decide?].
Signals to Monitor
Watch how often analytical work is preceded by a reconciliation exercise; a rising ratio of preparation to analysis is the clearest measure of the problem. Watch for unit-level reporting that no longer agrees with group reporting, which usually indicates a definition has moved without notice. Watch the proportion of optional fields left blank — an optional field is an unanswered "why". Watch for models quietly narrowing their scope during build, which is often missing data being absorbed as a technical constraint rather than escalated as a business decision. Externally, privacy and disclosure expectations continue to move; verify any specific claim about their current state before relying on it [FACT CHECK REQUIRED].
Questions for the Leadership Team
- Who is accountable, by name, for the definition of our five most commercially significant data fields?
- Which of our data domains are enterprise-defined, which are mapped, and which are unmapped — and are we funding any capability on the third category?
- What are we collecting today that informs no decision anyone owns, and what is it costing us to hold and protect?
- When a business unit changes a field definition, who is told, and through what control?
- What did our last major analytical initiative spend on reconciliation before it produced a single answer, and where did that appear in the business case?
- Which decisions we expect to make in three years depend on data we would have to begin capturing now?
Closing Perspective
The distinction that matters is between a problem an enterprise can buy its way out of and one it cannot. Storage, tooling and integration can be purchased at any point. Agreement on what the organisation means by its own terms, and authority to enforce that agreement across units that did not ask for it, cannot be purchased at all. It has to be granted by someone with the standing to grant it.
This is why the pipeline belongs on the operating-model agenda rather than the technology agenda. The choice is not which platform to build. It is whether the enterprise is willing to take a small, distributed, permanent cost in every unit today so that it retains the option to know something in three years — and to accept that if it declines, the option does not remain open. It closes quietly, in a period nobody reports on, and its absence is discovered by someone else.
About the author
Kevin Jogin is Founder & Principal Advisor at EraNorth. Meet the Founder.
