Reject Inference: The Problem Every Credit Model Has and Few Solve

Reject Inference: The Problem Every Credit Model Has and Few Solve | HL Hunt
Payments & AI

Reject Inference: The Problem Every Credit Model Has and Few Solve

A credit model learns from accounts that were funded, because those are the only ones with repayment outcomes. But funding was decided by a prior policy — which means the training sample is selected on precisely the variable being predicted. The consequence is subtle and compounding: your model knows a great deal about the population it already serves, nothing at all about the population just beyond its boundary, and no amount of retraining will ever tell it that the boundary is in the wrong place. Every cycle confirms the previous cycle's judgment. This guide covers why the standard corrections rest on an assumption that is usually false, and why the only genuine solution is to deliberately fund a small number of applicants you would have declined.

By the HL Hunt Research Desk · 16 min read · Updated August 2026

The censored sample

State the structure precisely, because the precision is what makes the implications clear.

You want to estimate P(default | applicant characteristics) across all applicants. What you can estimate is P(default | characteristics, approved). Those are the same quantity only if approval was independent of default risk — and approval was determined by an estimate of default risk. The conditioning is the entire problem.

Three consequences:

  • Below the cutoff you have zero observations. Not few — none. The model is extrapolating into a region where it has never seen an outcome.
  • Near the cutoff, the approved accounts are unrepresentative. The marginal approvals that got funded were those with favorable values on whatever the policy looked at, which means even the observations closest to the boundary are a filtered subset.
  • Model fit statistics are computed on the censored sample, so a model can show excellent discrimination on approved accounts while being badly wrong about the population it never sees. Good validation metrics are not evidence that the cutoff is correct.

This is not a defect of any particular model. It's structural: any lender who declines anyone has a censored sample, and the censoring is more severe the more selective the policy. The most disciplined underwriters have the worst data problem.

Why it compounds

The dynamic is what makes this urgent rather than merely interesting.

  1. Model version one draws a boundary, partly on judgment and partly on limited data.
  2. Approved accounts generate outcomes; declined applicants generate nothing.
  3. Version two trains on those outcomes and learns the approved region well.
  4. Version two draws a boundary informed only by the region version one approved.
  5. Repeat.

Each cycle the model becomes more confident about a region defined by a decision it can no longer evaluate. If version one's boundary was mispositioned — too tight in some segment, drawn on a variable that was proxying for something transient — nothing in the subsequent data will reveal it. The evidence that would falsify the boundary is precisely the evidence the boundary prevents you from collecting.

This connects directly to the swap-set analysis in our champion and challenger guide, and identifies a limit on it: a challenger tested only within the approved population can tell you how to reallocate approvals, but it cannot tell you whether the approval boundary itself should move outward. Challenger testing optimizes inside the frame. Reject inference is about the frame.

It also explains a pattern lenders observe and misdiagnose: approval rates drifting down across model generations without anyone deciding to tighten. Each retraining sharpens discrimination in the observed region and, absent countervailing pressure, tends to pull the cutoff inward. The drift looks like improving risk management. It is partly an artifact of the sample.

Zero observations below the line
Your model isn't estimating risk beyond the cutoff — it's extrapolating. And every retraining makes it more confident about a boundary it has never tested.

What it costs you

The cost is invisible by construction, which is why it goes unaddressed. It has three components.

Forgone good loans. Some share of your declined population would have performed acceptably. You never learn which, so the loss compounds indefinitely. Suppose 5% of your declines would have performed at or above your approved-population average — on 40,000 annual declines that's 2,000 creditworthy applicants refused every year, permanently, because the data that would identify them is the data you decline to generate.

Mispriced approvals near the boundary. Extrapolation error doesn't only affect declines; the accounts just above the cutoff are priced using a curve fitted where the data thins.

Foreclosed segment discovery. The most consequential and least discussed. If an entire population — thin-file applicants, a particular income structure, a channel — is systematically declined, you will never learn that a subset performs well. The thin-file population is the standard example: excluded historically for want of data, and the exclusion is self-perpetuating because it prevents the data from accumulating.

Reframed usefully: reject inference is not a modeling nicety. It is the question of whether your addressable market is larger than you believe, and you currently have no way to find out.

Two connections worth drawing. The censoring interacts badly with the vintage analysis most lenders rely on, because every vintage in the series was selected by the policy in force at the time — so a loss curve built from history describes the population you chose, not the one that applied. And it distorts the alternative-data case: a lender evaluating whether cash flow signals add predictive power can only test that on approved accounts, where the signal has least room to help, since the applicants it would most illuminate were declined.

Why the standard methods fall short

MethodWhat it doesThe problem
Ignore declinesTrain on approvals onlyHonest about its limits; makes no attempt at the population that matters
Assign all declines as badTreat every decline as a defaultGuarantees the model reproduces the existing boundary — the worst option, and common
Parcelling / augmentationAssign inferred outcomes by similarity to approved accountsAssumes approval carried no information beyond recorded variables
ReweightingWeight approvals to resemble the applicant populationSame assumption; can only reweight regions with some observations
Selection modelsModel approval and outcome jointlyNeeds a variable affecting approval but not repayment — genuinely hard to find
External performanceObserve declines' behavior on obligations elsewhereActually generates new information — see below
Random holdout approvalsFund a random sample of declinesThe only method producing unbiased outcome data

The shared weakness of rows three through five deserves stating plainly, because it is frequently glossed over in practice: they assume the approval decision contained no information beyond the variables in your dataset. That assumption is nearly always false. Manual overrides, verification failures, fraud flags, document inconsistencies, and underwriter judgment all influenced approval, and much of it was never recorded in a modelable form. So a declined applicant who looks identical to an approved one on the recorded variables differs on the unrecorded ones — and those unrecorded differences are exactly what the correction cannot see.

Our position: these methods are worth applying and should not be described as solutions. They reduce bias under assumptions you cannot verify. Presenting a model corrected by parcelling as though the selection problem has been handled is the actual error — worse than acknowledging the limitation, because it substitutes false confidence for a known gap.

Row two deserves a specific warning. Assigning all declines as bad is a common default and it is the worst available choice, because it doesn't merely fail to correct the bias — it encodes the existing boundary as ground truth. A model trained this way is guaranteed to reproduce the current cutoff, which is precisely the outcome you were trying to test.

The only real solution

Approve a small random sample of applicants your policy would decline, and observe what happens.

The logic is simple. Selection bias arises because approval correlates with risk. If a subset of approvals is assigned randomly among declines, that subset is uncorrelated with the unobserved factors — which makes it an unbiased sample of the decline population, and the only source of genuine information about it.

What it buys that nothing else does:

  • Actual outcomes below the cutoff, which no statistical technique can manufacture.
  • A test of whether the boundary is correctly placed, rather than an optimization within it.
  • Segment discovery — identification of decline subpopulations that perform acceptably.
  • Calibration of extrapolation error, revealing how wrong the model was about the region it guessed at.
  • A defensible empirical basis for later expanding access, which matters for both commercial and fair lending purposes.

The objection is immediate and legitimate: you are deliberately funding loans expected to perform worse. Yes. That expected loss is the price of the information, and it should be budgeted as a research cost rather than treated as an underwriting failure. A lender who will spend on data vendors, model development, and consultants but will not spend on generating the one data set that would resolve the largest uncertainty in their portfolio has an inconsistent view of what information is worth paying for.

Sizing and pricing the experiment

Make it a budget line rather than an act of faith.

Worked example. A lender declining 40,000 applications annually, average loan $3,000, approved-population loss rate 8%. They fund a random 1% of declines — 400 loans, $1.2 million deployed.

  • If the decline population loses at 25%: expected credit loss $300,000, against $96,000 had they performed like approvals — an incremental cost of roughly $204,000, offset by interest earned on the funded balances.
  • If the decline population loses at 15%: incremental cost roughly $84,000 — and this result would itself be the finding, since a 15% loss rate is priceable and implies the boundary is meaningfully too tight.

Now the return. If the experiment identifies a decline segment representing 5% of declines that performs acceptably, that's 2,000 applicants a year at $3,000 — $6 million of annual originations that were previously invisible. A one-time cost in the low hundreds of thousands against a recurring revenue stream of that size is not a close call.

Design parameters that reduce the cost without destroying the inference:

  • Sample near the boundary preferentially, where the information value per dollar is highest and expected losses are lowest — while retaining some deeper sampling, or you learn nothing about the wider region.
  • Reduce loan size and term for holdout accounts, which caps exposure and shortens time to outcome.
  • Price to the segment where lawful, so the experiment partly funds itself.
  • Run continuously at low volume rather than as one large cohort, producing a rolling unbiased sample and avoiding a single vintage's economic conditions.
  • Stratify the sample so you get observations across the decline distribution rather than concentrated in one region.

Cheaper partial substitutes

Where a holdout isn't feasible, these recover some information at lower cost:

  • Bureau follow-up on declined applicants. Observe, subject to permissible purpose and appropriate controls, how declines performed on obligations they obtained elsewhere. This is genuinely informative — it's real outcome data on real people you declined — with the caveat that a decline who borrowed elsewhere is itself a selected subset, and one that likely skews better than the full decline population.
  • Override outcomes. Manual overrides that funded applicants the model declined are a naturally occurring sample below the cutoff. They are not random — an underwriter overrode for a reason — but they're free, they already exist in your data, and most lenders have never analyzed them as a group. This is the highest-return unexploited analysis in most shops.
  • Policy change natural experiments. When a cutoff moved, the newly approved band is observable. Compare it against the model's prediction for that band.
  • Acquired portfolios underwritten to different standards, which contain accounts your policy would have declined.
  • Partner or channel variation, where different criteria applied to comparable populations.

Ranked honestly: a random holdout is decisively better than all of these, and the override analysis is the best of the free options and should be run first because it costs nothing but a query.

Doing it responsibly

A deliberate approval experiment extends credit to people your model expects to struggle. That deserves care, not just documentation.

  1. Exclude categories where the decline reason isn't about model uncertainty — fraud indicators, identity failures, and ability-to-repay determinations. You are testing a risk boundary, not overriding a protective one.
  2. Limit exposure through smaller amounts and shorter terms, so a borrower who does struggle struggles less.
  3. Apply identical servicing, hardship options, and collections treatment as any other account. Holdout borrowers are customers, not subjects.
  4. Meet all disclosure and adverse action obligations for the applicants not selected, per our notices guide.
  5. Ensure random selection is genuinely random, since any selection rule reintroduces the bias you're trying to eliminate.
  6. Set stopping rules in advance — the loss level at which the experiment halts regardless of how informative it is proving.
  7. Document as a governed model risk activity with defined objective, design, approvals, and monitoring, per our governance framework.
  8. Obtain fair lending review before launch, not after.

The fair lending dimension

Worth being explicit about, because the framing cuts in a direction people don't expect.

The selection problem has a fair lending consequence that is rarely stated: if a group is disproportionately declined, the model has less data about that group, which makes its estimates for them less reliable, which tends to perpetuate the decline rate. The mechanism is statistical rather than intentional, and it is self-reinforcing in exactly the way that produces persistent disparate outcomes without anyone choosing them.

Which means:

  • Reject inference work is fair lending work, not merely a commercial optimization. Reducing extrapolation error where data is thinnest disproportionately affects populations who were declined most.
  • Less discriminatory alternative analysis is weakened by the censoring. Searching for an alternative model with comparable performance and lower disparity is constrained if the data can't evaluate alternatives beyond the current boundary.
  • A documented holdout program is strong evidence of good faith — a lender actively testing whether its boundary excludes creditworthy applicants is doing something concrete about a known limitation rather than asserting the model is fine.

The honest counterweight: an experiment that disproportionately places worse-performing credit with a protected class would be its own problem, however good the intent. This is precisely why the sample must be genuinely random within the decline population and why the composition should be reviewed before launch rather than discovered afterward.

Find out what's beyond your boundary

HL Hunt AI Underwriting supports holdout sampling and override analysis alongside standard decisioning — with cohort tagging that keeps experimental accounts separable from your book, so the information they generate is usable and the losses they carry are measured rather than absorbed invisibly.

Explore HL Hunt AI Underwriting

Frequently asked questions

What is reject inference in credit scoring?

The problem of estimating how declined applicants would have performed when they were never funded. Because approval was determined by risk, the training sample is selected on the outcome being predicted.

Why do credit models get worse near the approval boundary?

Because data thins toward the cutoff and vanishes below it. Each retraining learns confidently about the region already served and nothing about the boundary, so successive models reproduce and narrow it.

Do standard reject inference techniques actually work?

Modestly, on assumptions usually false — they presume the approval decision carried no information beyond recorded variables, when overrides, verification failures, and judgment all influenced it and went unrecorded.

Is it responsible to deliberately approve applicants your model declines?

With design and governance, yes. Use small stratified samples, limit exposure, exclude fraud and ability-to-repay declines, apply identical servicing, and obtain fair lending review before launch.

Key takeaways

  • Credit models train on approved accounts only, so the sample is selected on the outcome being predicted and there are zero observations below the cutoff.
  • The bias compounds — each generation grows more confident about a boundary it cannot test, which quietly pulls approval rates down.
  • Most reject inference methods assume approval contained no unrecorded information, which is nearly always false; assigning all declines as bad is the worst common choice.
  • Randomly approving a small sample of declines is the only method that generates unbiased outcome data, and the expected loss is the price of the information.
  • Analyzing past manual overrides is free, already in your data, and the highest-return unexploited substitute.
  • Because censoring is worst where declines are most concentrated, reject inference work is fair lending work rather than only a commercial optimization.

Start with the data you already have

Before running any experiment, the override population in your existing book is an unexploited sample below your cutoff. HL Hunt AI Underwriting can run in shadow mode across your history to show where the model and your prior decisions disagreed — and how those accounts actually performed.

Get Started with HL Hunt AI Underwriting


This guide is educational and does not constitute legal or compliance advice. Deliberate approval experiments, use of consumer report data on declined applicants, and model risk practices are subject to fair lending, permissible purpose, and model governance requirements; consult qualified counsel and your model risk function before implementing any program described here. Worked figures are stylized illustrations.