Mining Your Overrides: The Free Data Sitting in Every Credit File

Mining Your Overrides: The Free Data Sitting in Every Credit File | HL Hunt
Payments & AI

Mining Your Overrides: The Free Data Sitting in Every Credit File

Our reject inference analysis argued that credit models can never evaluate their own approval boundary, because declined applicants generate no outcomes. It also noted that most lenders hold a partial exception they've never examined: manual overrides — the applicants an underwriter funded despite the policy saying no. Those accounts sit below the cutoff, they have real outcomes, they're already in the portfolio, and analyzing them costs a database query. For most lenders this is the cheapest available evidence about whether their approval boundary is in the right place, and almost nobody has looked.

By the HL Hunt Research Desk · 15 min read · Updated August 2026

Why overrides are the exception

Restate the underlying problem precisely. A credit model learns from funded accounts. Funding was decided by a policy. So the model has outcomes only from the region the policy approved, and no information whatsoever about the region it declined — which is why successive model generations tend to confirm and narrow the boundary their predecessors drew.

Overrides break that, partially. They are funded accounts from below the cutoff, and they have observed outcomes. Every property that makes them awkward — they were chosen, not randomized — is a limitation on interpretation rather than a reason to ignore them.

Three reasons this is worth doing before anything more elaborate:

  • The data already exists. No experiment, no expected loss to budget, no approval to fund declined applicants. The accounts were funded years ago and the outcomes are known.
  • It's fast. A lender with tagged override decisions can run this in a day.
  • It informs whether to run the expensive version. If overrides suggest the boundary is badly placed, that's the argument for a randomized holdout program. If they suggest it's about right, you've saved the cost of finding out.

The prerequisite is that overrides were tagged — recorded as overrides, with the reason and the decision maker. Where they weren't, the accounts are unidentifiable and the information is lost permanently, which is the strongest practical argument for the capture discipline in our cold start analysis.

Low-side and high-side

Two categories that must never be pooled, because they answer different questions.

Low-side overrideHigh-side override
What happenedPolicy said decline; approved anywayPolicy said approve; declined anyway
Typical reasonUnderwriter saw something the model couldn'tA concern the model doesn't capture
Outcome observable?Yes — the account was fundedNo — nothing was funded
Question it answersIs the cutoff too tight?Is the model missing a risk?
Analytical valueHigh — real outcomes below the cutoffLimited directly; useful indirectly

Low-side overrides carry nearly all the value, for the obvious reason that they produce outcomes. High-side overrides produce nothing to observe — the applicant was declined, so you learn no more about them than about any other decline.

High-side overrides are still worth examining for a different purpose: their volume and reasons tell you what your model systematically fails to capture. If underwriters routinely decline model-approved applicants for a consistent reason, that reason is a missing variable — and the model is approving those applicants at scale wherever no underwriter reviews the file. That's a finding about model coverage rather than about the cutoff, and it's frequently more urgent.

Below the cutoff, with outcomes
Low-side overrides are the only funded accounts most lenders have from the population their model declines. The evidence already exists — it just needs a query.

The core analysis

Six steps, in order:

  1. Identify all low-side overrides across a period long enough for outcomes to have matured — several years for most consumer products.
  2. Establish what the policy said at the time, and by how far each applicant fell short. An applicant just below the cutoff and one far below are different observations.
  3. Measure realized performance — delinquency, charge-off, and net loss — at matched months on book.
  4. Build a comparison group of approved accounts from the same vintages, ideally matched on observable characteristics other than the failing criterion.
  5. Compare. How did overrides perform relative to the approved population, and relative to what the model predicted for them?
  6. Segment by override reason, by distance below cutoff, by decision maker, and by vintage.

Worked illustration. A lender examines 840 low-side overrides across three years against 40,000 standard approvals from the same vintages.

GroupAccountsCharge-off rateModel-predicted rate
Standard approvals40,0007.2%7.0%
Low-side overrides, all8409.8%16.4%
— near cutoff5107.9%13.1%
— far below cutoff33012.7%21.5%

Read the last two columns together, because the gap between them is the finding.

Reading the result

Three conclusions from that table, in ascending order of consequence.

The model over-predicted risk for every override group. Predicted 16.4%, realized 9.8%. That gap is the value the underwriter's judgment added — they identified applicants the model scored too harshly.

Near-cutoff overrides performed almost identically to standard approvals — 7.9% against 7.2%. Given that these applicants were declined by policy, that's a substantial finding: the cutoff is excluding a population that performs at roughly the approved rate. At standard pricing that population is profitable, and it's being turned away.

Far-below overrides performed materially worse — 12.7% — but still far better than predicted. Whether that's acceptable depends on pricing, which is a different question from whether it's an error.

What a lender should do with this:

  • Consider moving the cutoff to capture the near-cutoff population, or introduce a pricing tier for it rather than declining.
  • Investigate why the model over-predicts in this region — frequently it's extrapolation error where training data thinned, exactly as expected.
  • Identify which override reasons did the work, covered next.
  • Consider a randomized holdout to test the boundary properly, now that there's evidence justifying the cost.

And the opposite result matters just as much. If overrides had performed worse than predicted, the conclusion would be that the override process adds negative value — underwriters are approving applicants the model correctly declined, and the exception process should be tightened or removed. That's a clean, actionable finding and it's equally likely a priori.

The selection bias question

The objection that stops most people from running this, and it deserves a proper answer rather than dismissal.

The objection: overrides weren't randomly selected. Underwriters chose them for reasons. So their performance reflects that selection, and you can't generalize to the whole declined population.

The response, in three parts:

The bias is directional and known. Underwriters selected applicants they believed would perform. So override performance is an upper bound on how the declined population would perform — not an unbiased estimate. That's a limitation you can reason with rather than a fatal flaw.

One direction of the result is clean. If deliberately selected overrides perform poorly, you've learned something unambiguous: the judgment adds nothing, because even hand-picked exceptions fail. No selection argument rescues that finding.

The other direction is informative but bounded. If overrides perform well, you've learned that a repeatable signal exists — underwriters can identify good applicants the model rejects. You cannot conclude that all similar declined applicants would perform equally. What you've established is that the model has a blind spot and that someone can see into it, which is grounds for finding out what they're seeing.

The framing that resolves it: override analysis is a screening test, not a definitive one. It tells you cheaply whether the expensive investigation is warranted. Refusing to run it because it isn't randomized is refusing free information because it isn't perfect information — and the perfect version costs money the screening test would justify spending.

Analyzing by reason

The most operationally useful cut, because it converts a finding into a change.

Group overrides by the documented reason and compare performance. Reasons that recur and predict well are information your model is missing:

  • Verified income exceeding the application figure — a data quality issue rather than a credit one.
  • Existing relationship with the institution, which is the behavioral evidence our renewal analysis argues outperforms bureau data.
  • An explained derogatory — a documented one-time event rather than a pattern.
  • Substantial liquid assets, which is the buffer our liquidity analysis argues predicts shock absorption better than anything in a bureau file.
  • Thin file rather than bad file — the distinction our thin-file analysis describes, where the model scores absence of data as though it were negative data.
  • Employment stability not captured in the application.
  • Collateral or structure improving the position.

What the distribution tells you:

A reason appearing frequently and predicting well is a missing variable. It should be encoded rather than left to judgment — see below.

A reason appearing frequently and predicting nothing is a rationalization. Underwriters wanted to approve and reached for a justification. Worth naming, since it identifies where the exception process is being used to bypass policy rather than to improve it.

Undocumented or vague reasons are a governance finding regardless of performance. An override without a recorded reason can't be analyzed, can't be defended, and shouldn't have been permitted.

Analyzing by decision maker

Uncomfortable and valuable, and it should be handled as a process finding rather than a performance review.

Override quality varies by person, and the variation is itself informative about what judgment contributes:

  • An underwriter whose overrides consistently perform is applying a signal worth extracting. Interview them — the reasoning is a model input waiting to be encoded.
  • One whose overrides consistently underperform needs coaching or reduced authority, and finding this out from outcomes beats finding it out from a bad vintage.
  • Wide dispersion across underwriters indicates the criteria aren't shared, which is a fair lending exposure as well as an inconsistency — discretion applied unevenly across similar applicants is precisely the pattern outcome testing surfaces.
  • Override rate variation matters separately from quality: one person overriding at 8% and another at 0.5% means applicants are being treated differently based on file assignment.

The constructive framing: the goal is to extract what good underwriters know, not to rank them. An underwriter with a demonstrable edge is producing evidence that a variable exists which the model lacks, and the right response is to find it and encode it — which makes their judgment available to every applicant rather than to the ones whose files they happen to review.

Converting findings into policy

The step that realizes the value, and the one most often skipped.

  1. Encode validated override reasons as policy rules or model variables. A reason that reliably identifies good applicants shouldn't live in an exception process.
  2. Adjust the cutoff where near-cutoff overrides perform at approved rates.
  3. Introduce a pricing tier rather than a decline where performance is acceptable but worse — converting a binary boundary into a priced one, which is generally the better structure.
  4. Tighten the exception process where overrides underperform.
  5. Require documented reasons from a defined list, so future analysis is possible.
  6. Set an override rate target, since both extremes are problems — a very low rate means judgment isn't being applied where it could help, and a very high rate means the policy isn't governing.
  7. Test the change through the challenger design in our deployment guide rather than switching outright.

The principle underneath: the exception process should be a discovery mechanism, not a permanent home for knowledge you've already validated. Every reason that repeatedly proves itself and stays in the exception process is being applied inconsistently, to a fraction of the applicants it fits, depending on who reviewed the file.

Governance and fair lending

Overrides are where discretion enters an otherwise systematic process, which makes them the highest-scrutiny part of a credit operation.

  • Outcome-test overrides by group. Discretion applied unevenly is exactly the pattern fair lending analysis is designed to detect, and an override population skewed on protected characteristics is a serious finding regardless of intent — per our governance framework.
  • High-side overrides deserve particular attention, since declining a model-approved applicant on judgment is discretion operating against the applicant.
  • Document every override with reason, approver, and date.
  • Ensure adverse action reasons remain accurate where an override results in a decline, per our notices guide.
  • Monitor override rates by underwriter and by segment as a standing control rather than a periodic review.
  • Treat consistent override reasons as a model deficiency to be remediated, which is both the better outcome and the better governance posture — a documented process for converting exceptions into policy is far stronger than a permanent exception volume nobody analyzes.

The evidence is already in your book

HL Hunt AI Underwriting tags every decision with policy result, override status, reason, and approver — and can run in shadow mode across your history to compare model predictions against what your overrides actually did.

Explore HL Hunt AI Underwriting

Frequently asked questions

What is a credit override?

A decision departing from what policy prescribed. Low-side approves a declined applicant; high-side declines an approved one. They answer different questions and must be analyzed separately.

Why are overrides valuable for analysis?

Low-side overrides are usually the only funded accounts you have from below your own cutoff — the population your model has no outcome data about. The accounts exist and the outcomes are known.

Doesn't selection bias make override analysis unreliable?

It bounds the conclusion rather than voiding it. Poor override performance is a clean finding that judgment adds nothing; good performance establishes a blind spot exists, without generalizing to all declines.

What should you do when an override reason consistently predicts good performance?

Encode it as policy or a model variable. Left in the exception process it's applied inconsistently, to a fraction of the applicants it fits, depending on who reviews the file.

Key takeaways

  • Low-side overrides are the only funded accounts most lenders hold from below their own cutoff, and analyzing them costs a query.
  • In the worked example, near-cutoff overrides charged off at 7.9% against 7.2% for standard approvals — a declined population performing at approved rates.
  • The model over-predicted risk for every override group, which is the extrapolation error you'd expect where training data thins.
  • Selection bias makes override performance an upper bound, so poor performance is a clean finding and good performance is grounds for a proper test.
  • Override reasons that recur and predict well are missing model variables and should be encoded rather than left to judgment.
  • Overrides are where discretion enters, which makes outcome testing by group a standing requirement rather than a periodic one.

Find your boundary before paying to test it

Get started with HL Hunt AI Underwriting to tag every decision, reason, and override from origination — so the analysis that tells you where your cutoff belongs is available whenever you want to run it.

Get Started with HL Hunt AI Underwriting


This guide is educational and does not constitute legal or compliance advice. Fair lending, adverse action, and model risk obligations apply to override processes and to any policy change derived from them; consult qualified counsel and your model risk function. Worked figures are stylized illustrations.