The Unobserved Counterfactual: The Same Measurement Failure, Five Times Over

The Unobserved Counterfactual: The Same Measurement Failure, Five Times Over | HL Hunt
Institutional Outlook

The Unobserved Counterfactual: The Same Measurement Failure, Five Times Over

This desk has now analyzed five separate operational problems — loan workouts, approval cutoffs, fraud screening, collections treatment, and payment mix — and found the same thing underneath each. In every case, a decision forecloses its own alternative, so the alternative's outcome never happens, never gets recorded, and is therefore treated as costing nothing. Three of four possible outcomes are measured and the fourth is unobservable by construction. What follows isn't random error. It's a persistent drift in one direction, in operations run by people responding rationally to reporting that is missing exactly one quantity — and the remedy is identical in all five cases, cheap in all five, and almost never applied.

By the HL Hunt Research Desk · 25 min read · Updated August 2026

The general structure

Take any binary decision with an uncertain outcome. Four things can happen:

ActDon't act
Acting was rightObserved — the benefit shows upObserved — the loss shows up
Acting was wrongNot observedObserved — nothing bad happens

The cell that goes missing is always the same one: you acted, and the outcome that would have occurred otherwise never happened, so its foregone value never appears anywhere.

Two properties make this consequential rather than merely interesting:

It's structural, not a data quality problem. No better system or more careful collection recovers the missing cell. The observation doesn't exist because the event didn't occur, and it didn't occur because you prevented it.

The bias is directional and compounds. An unmeasured cost is treated as zero. A cost treated as zero is under-weighted forever, in every subsequent decision, by every manager reading the reports. The error doesn't average out — it accumulates and self-reinforces, because each period's decisions are made using reporting shaped by the previous period's blindness.

No better system recovers the missing observation. The event didn't happen, and it didn't happen because you prevented it.

The five cases

DecisionWhat's unobservedResulting drift
Enforce or forbear
workout analysis
The borrower who'd have recovered if you'd waitedEnforce too readily
Approve or decline
reject inference
The declined applicant who'd have repaidCutoff too tight
Screen or accept
fraud analysis
The declined customer who was legitimateScreen too hard
Treat or leave
cure analysis
The account that would have cured anywayOver-treat where it doesn't help
Add a method or not
incrementality analysis
The sale that would have happened regardlessOver-credit the addition

These were analyzed separately, across four different functions, over several months, without the connection being the point of any of them. That they share a structure is the finding here.

A sixth case is the partial exception and worth noting because it shows the general remedy already exists in practice. Manual overrides, per our override analysis, are accounts funded below a cutoff — which means they're observations from the missing cell, generated accidentally by human judgment rather than deliberately by design. Most lenders hold a small, biased, free sample of exactly the data they say they can't get, and almost none have looked at it.

Two families, opposite errors

The analytical contribution of this synthesis: the five cases split into two families that produce opposite errors, and conflating them leads to the wrong correction.

Family A — the foreclosed benefit. Acting prevents a good outcome from occurring, so the good outcome is never observed and the action looks free.

  • Enforcing prevents the recovery that waiting would have produced.
  • Declining prevents the repayment that approving would have produced.
  • Screening prevents the sale that accepting would have produced.

The error is being too restrictive, and the visible metric — losses avoided, fraud rate, charge-offs — improves with every increment of restriction. Which is why these operations tighten indefinitely: the metric can never tell them to stop.

Family B — the attributed outcome. Acting gets credit for an outcome that would have occurred anyway, so the action looks effective.

  • Collections treatment gets credit for accounts that would have cured on their own.
  • A new payment method gets credit for sales that would have happened through another.
  • The same applies to marketing, retention offers, and most customer interventions.

The error is doing too much of something that doesn't work, and the visible metric — cure rate, adoption, response rate — looks strong regardless of whether the action contributed anything.

The distinction matters practically. Family A costs you revenue you never see. Family B costs you effort spent on outcomes you'd have had for free. Both are invisible, both are expensive, and the corrections point in opposite directions — loosen in the first case, stop doing things in the second.

The metric can never say stop
Fraud rate improves with every tightening. Cure rate looks strong regardless of treatment. A metric that only moves one way cannot tell you when you have gone too far.

Why one operation has both

The observation that makes this more than taxonomy: a single operation frequently exhibits both errors simultaneously, in the same function, on the same accounts.

Consider a collections operation. Our cure analysis found lift of about 3 points at 15 days past due and 23 points at 60 days. So the operation is:

  • Over-treating early accounts — Family B, spending on accounts that would have cured anyway, and reporting a 91% cure rate that looks excellent.
  • Under-forbearing late accounts — Family A, enforcing where waiting would have recovered more, and reporting charge-offs that look like a market condition.

Both errors, one operation, opposite directions — and neither visible in any report the operation produces. The early-stage numbers look good because they measure a population that didn't need help; the late-stage numbers look bad because they measure a population that was mishandled.

The same coexistence appears in acquisition. A lender can simultaneously decline applicants who'd have performed (Family A, at the cutoff) and pay for marketing to customers who'd have applied anyway (Family B, in acquisition). The two are usually owned by different teams with different metrics, so nobody sees them as one problem.

Which suggests a diagnostic worth running across an organization rather than within a function: list every decision where one branch prevents its own alternative from being observed. The list is usually longer than expected and cuts across ownership boundaries.

The remedy is always the same

Every one of the five cases is solved by the same design: take the opposite decision on a small random sample.

CaseThe randomization
WorkoutsOffer relief to a random subset of accounts you'd enforce on
Approval cutoffApprove a random sample below the cutoff, at reduced exposure
Fraud rulesAccept a random sample of transactions the rules would decline
Collections treatmentWithhold treatment from a random subset at each stage
Payment methodsOffer to a random share of traffic and compare total conversion

What makes randomization work where nothing else does: it creates observations in the missing cell. A random sample of the foreclosed branch is, by construction, representative of the population you were foreclosing on — which no amount of modelling on the observed cells can produce, because the observed cells contain no information about the unobserved one.

Three design points that recur across all five:

Random, not selected. A judgment-selected sample answers a different question — which is precisely the limitation of using overrides as a substitute, as our override analysis notes. Overrides are cheap and biased; randomization is the unbiased version.

Small is sufficient. These are questions about population averages, not individual predictions. A few percent of volume typically suffices, and the sample size needed is driven by the effect you're trying to detect rather than by portfolio size.

Tag permanently. An experimental cohort that isn't identifiable afterward has generated cost and no information — the most common way these programs fail.

What the experiment costs

Work it, because the number is usually the objection and usually smaller than assumed.

A lender declining 40,000 applicants a year wants to know whether the cutoff is right. They approve a random 2% of near-cutoff declines — 800 accounts — at a reduced exposure of $600 each.

  • Total exposure: $480,000
  • If the declined population performs badly, say a 30% loss rate: expected loss $144,000
  • Less interest and fees earned on the 70% that perform, say $85,000
  • Net cost of the experiment: roughly $59,000

Now the value. If the test shows the cutoff is 20 points too tight, and moving it captures 3,000 additional accounts a year at $180 of contribution each, that's $540,000 annually, recurring.

A one-time cost of $59,000 answers a question worth $540,000 a year — and the same arithmetic holds if the answer is the other way, because learning the cutoff is correct ends a recurring debate and justifies the current policy under examination.

The framing that makes it defensible internally: this is not a loss, it is the price of information, and it should be budgeted as research rather than reported as credit performance. An operation that books experimental losses in the same line as ordinary losses has guaranteed the program will be cancelled.

Why it isn't done

If the remedy is this cheap and this general, the near-total absence of it needs explaining. Four reasons, none of them about the analysis being wrong.

It requires deliberately doing what you believe is wrong. Approving applicants your policy declines, or not calling accounts you're paid to call, feels like negligence even when it's research. That discomfort is real and it stops programs before they're proposed.

The cost is visible and the benefit isn't yet. The experiment's losses appear in reporting immediately; the ongoing cost of not knowing appears nowhere. This is the original problem applied to its own solution — the asymmetry that created the blindness also blocks the remedy, which is the most self-reinforcing feature of the whole structure.

Institutional consequence is asymmetric. "We ran an experiment and lost money" is attributable to a person. "We have been leaving money on the table invisibly for six years" is attributable to nobody. A manager choosing between an attributable small loss and an unattributable large one is not behaving irrationally by choosing the second.

Governance friction. Deliberately extending credit outside policy, or withholding treatment, raises questions that need answering in advance — and the answers exist but require work that nobody's job description includes.

Which suggests the barrier is organizational rather than analytical, and the interventions that would help are correspondingly organizational: a research budget line separate from operating losses, pre-approved governance for holdout designs, and ownership assigned to someone whose metric is knowledge rather than performance.

How to spot one you haven't found

The general test, applicable to any decision in any function:

  1. Identify the decision and both branches.
  2. Ask what evidence each branch generates. If one branch produces an outcome record and the other produces nothing, you've found one.
  3. Determine the family. Does acting foreclose a benefit (A) or claim an outcome that would have occurred anyway (B)?
  4. Check the direction of the reported metric. If the headline metric improves monotonically with more of the action, it cannot evaluate the action — this is the fastest tell.
  5. Ask who bears the invisible cost and whether anyone's metric includes it.
  6. Design the randomization and price the experiment.

Step four is the diagnostic worth memorizing. Fraud rate always improves with tighter screening. Cure rate always looks good on early accounts. Charge-off rate always improves with a tighter cutoff. Any metric that moves in only one direction as you do more of something is a metric that cannot tell you when to stop — and every case in this report was identifiable from that property alone.

Candidates worth checking that this desk hasn't analyzed: credit line assignment, account closure decisions, retention offers, dispute resolution thresholds, manual review routing, and marketing spend of every kind. Each has a branch that forecloses its own alternative.

Testable implications

  1. Operations that have never randomized should be miscalibrated in the predicted direction — too restrictive in Family A cases, over-treating in Family B ones. This is the core claim and it's checkable wherever a first holdout is run.
  2. The first holdout in any of these cases should produce a surprising result, because the prior was formed without the missing cell.
  3. Operations with a research budget separate from operating losses should randomize more than those without, holding size constant.
  4. Family A and Family B errors should coexist within single operations, since they arise from the same structure at different decision points.
  5. Override performance should predict holdout results in the same direction but with attenuated magnitude, since overrides are a selected version of the same sample.
  6. The size of the correction should be largest where the decision has been automated longest, because automation removes the human overrides that were accidentally generating observations.

The sixth is the one worth dwelling on. Automating a decision eliminates the discretionary exceptions that were producing the only data about the foreclosed branch — so an operation that automates without simultaneously introducing randomization has made itself permanently blind at exactly the moment it gained the ability to be systematic. That's a specific and underappreciated cost of automation, and it argues that a holdout should be built into any decision system at the point of automation rather than added later, when the historical evidence has already stopped accumulating.

Frequently asked questions

What is an unobserved counterfactual in a business decision?

The outcome that would have occurred had the opposite decision been taken. It never happens, so it never generates data, so it never appears as a cost — and decisions drift toward whichever side is measured.

Why does this produce systematic rather than random error?

The missing information is always on the same side — the cost of having acted when you shouldn't have. That cost never appears while the alternative's costs do, so the drift is consistent and compounds.

What is the remedy?

Randomization. Taking the opposite decision on a small random sample creates observations in the missing cell, and the design is identical across every case.

Why do so few organizations do this?

It requires deliberately doing what you believe is wrong at a visible cost, for information whose value isn't yet proven. The experiment's cost is attributable; the cost of not knowing is attributable to nobody.

Key takeaways

  • Whenever a decision forecloses its own alternative, three of four outcomes are measured and the fourth is unobservable by construction.
  • The missing cell is always the same one — the cost of having acted when you shouldn't have — so the error is directional and compounds.
  • Family A errors make operations too restrictive; Family B errors make them over-treat, and a single operation frequently has both at once.
  • A metric that improves monotonically with more of an action cannot evaluate that action — the fastest way to detect the structure.
  • The remedy is randomization, identical in every case, and typically costs a fraction of the recurring value of the answer.
  • Automating a decision removes the discretionary exceptions that were accidentally generating the only evidence, so holdouts belong in the design.

This report presents an analytical framework and the authors' interpretation; it is not legal, compliance, or financial advice. Worked figures are stylized illustrations. Deliberately extending credit outside policy, withholding collection treatment, or accepting transactions that screening would decline all raise regulatory and contractual considerations that must be addressed before any such program begins — consult qualified counsel and your model risk and compliance functions.