Buying a Model: What Vendor Due Diligence Should Actually Cover
Buying a Model: What Vendor Due Diligence Should Actually Cover
A vendor presents a credit model with impressive performance statistics, and the natural response is to compare those statistics against alternatives. That's the wrong comparison. Every number a vendor quotes describes how the model performed on the population it was built and tested on — and the only question that determines whether it will work for you is whether your applicants resemble that population. Most buyers never ask, because the statistics look like an answer. Meanwhile the responsibility for every decision the model makes stays with you: fair lending obligations, adverse action accuracy, and model risk management don't transfer with a licence.
What you'll learn
The question that decides it
What did the development sample look like?
A model learns relationships from a set of accounts. Those relationships hold for populations resembling that set and degrade as the population diverges. So the transferability question is entirely about resemblance, and it dominates every performance statistic.
What to establish about the sample:
- Product type. Instalment, revolving, secured, and small-dollar behave differently — a model built on one won't necessarily rank order on another.
- Credit quality distribution. A model built on prime applicants may separate poorly across a subprime range simply because that range was sparse in training.
- Thin-file representation. If your population is thin-file and the sample wasn't, the imputation problems in our missing data analysis will be doing more work than the model.
- Loan size distribution.
- Channel composition, which selects for characteristics that persist and aren't fully captured in a file.
- Geographic and economic period coverage. A model developed entirely in benign conditions hasn't observed the correlated stress our correlation analysis describes.
- Sample size and, critically, default count. Defaults are what a model learns from — a large sample with few defaults is a small sample.
The finding that should follow: a vendor who cannot describe their development sample in this detail is selling something you cannot evaluate. That's a legitimate reason to decline regardless of the statistics, and it's the single most useful screening question available.
And the honest note about statistics generally: a discrimination statistic quoted without the population it was measured on is uninterpretable. Two models can differ substantially on the same measure purely because one was tested on a wider credit range, which is easier to separate.
The outcome definition
The second question, and the one that produces the most surprising misestimates.
A model predicts a specific outcome, defined a specific way. Common definitions vary on:
- Delinquency threshold — 60, 90, or 120 days past due.
- Performance window — 12, 18, or 24 months.
- Treatment of charge-off versus delinquency.
- Whether cures count, which matters enormously given the natural cure rates in our cure analysis.
- Whether the definition is account-level or borrower-level.
- Whether fraud is excluded — and per our attribution analysis, a model trained on a target contaminated with fraud is predicting something other than credit risk.
Why it matters practically: a model predicting 90-day delinquency at 18 months will misestimate your 120-day charge-off rate at 24 months, and the direction and magnitude of the error aren't obvious. The rank ordering may transfer while the level doesn't — which is the same distinction our cold start analysis draws about generic scores.
What to do: get the definition in writing, compare it to yours, and if they differ, recalibrate on your own data rather than using the vendor's implied rates. A vendor's odds-to-score relationship is calibrated to their definition and their population, and using it directly is how loss forecasts miss badly.
Validating on your own data
Non-negotiable, and frequently skipped because the vendor already validated.
A vendor's validation tells you the model works on their data. It tells you nothing about yours.
What to run before deployment:
- Score your own historical accounts — ones with known outcomes, from before the model would have influenced anything.
- Measure rank ordering on your population. Does the score separate your good accounts from your bad ones?
- Check calibration. Does a score band's predicted rate match your realized rate? Expect a gap, and quantify it.
- Test by segment — by channel, product, geography, and credit band. A model performing well overall can perform poorly on a segment that matters to you, and aggregate statistics hide it.
- Compare against your current approach. The question isn't whether the model is good; it's whether it's better than what you have.
- Run fair lending analysis on your population, per our governance framework — a vendor's disparity testing was conducted on their applicants.
- Shadow it live before it decides anything, per our deployment guide.
Step four deserves emphasis. Segment-level performance is where purchased models most commonly disappoint, because a vendor optimizes for aggregate performance across their whole market, and your book may be concentrated in a segment they treated as marginal.
Reason codes and explainability
Where a purchase creates obligations a buyer frequently doesn't anticipate.
You must provide accurate principal reasons for adverse action, per the requirements in our notices guide. That obligation is yours whatever the vendor supplies.
What to establish:
- Does the model produce reason codes, and on what basis?
- Are they accurate for your population? Reason codes derived from the vendor's sample may not describe what drove a decision on yours.
- Can you explain an individual decision if asked to?
- Do reasons trace to actual applicant facts rather than to imputed values — the specific accuracy problem in our missing data analysis.
- Is the variable list disclosed? You cannot assess whether a variable is problematic if you don't know it's in there.
The position worth taking: a vendor who won't disclose the variables under confidentiality is asking you to accept obligations you cannot discharge. You're required to explain decisions you'd be making with a mechanism you're not permitted to understand — and "the vendor considers it proprietary" is not an answer an examiner accepts.
What you can't outsource
| Responsibility | Transfers? |
|---|---|
| Fair lending compliance | No |
| Adverse action accuracy | No |
| Model risk management | No |
| Validation on your population | No |
| Ongoing performance monitoring | No |
| Documentation for examination | Supported, not transferred |
| Model development | Yes |
| Maintenance and updates | Partly |
Only the bottom two genuinely transfer. Everything that generates regulatory exposure stays with you, which reframes what a purchase actually buys: development effort and expertise, not responsibility.
Two implications. The monitoring obligation is yours — the input distribution and performance tracking in our monitoring guide must run on your side, since a vendor monitoring their model across all clients isn't monitoring it on your population. And the override analysis in our override guide applies with more force to a purchased model, because it's the cheapest evidence available about whether the vendor's cutoff suits your applicants.
When building is better
The honest comparison, since buying is the right answer more often than not.
Buy when: you lack sufficient matured defaults, your population resembles a vendor's market, you need to launch quickly, or you lack the modelling capability. That covers most institutions and most new products — and it's why the expert scorecard approach in our cold start analysis pairs with a purchased score rather than replacing it.
Build when:
- Your population differs materially from anything a vendor built for.
- You hold data no vendor has — cash flow, behavioural, or relationship data that predicts on your book specifically.
- You have enough matured defaults, which is a count of defaults rather than accounts.
- Your outcome definition differs and recalibration isn't sufficient.
- The dependency is strategically uncomfortable, since a model you can't leave is a commercial position.
The threshold that decides it in practice: defaults, matured. A book of 60,000 accounts with 400 matured defaults cannot support a reliable model, and a book of 6,000 with 900 can. This is the constraint that makes building unavailable to most institutions regardless of their size or ambition — and the low-loss products where building is hardest are precisely the ones where a purchased model's calibration will be furthest off.
The commercial terms that matter
- Pricing structure — per score, per approval, or subscription, and how it scales if your volume changes.
- What happens at renewal. A model embedded in your decisioning is a weak negotiating position at renewal.
- Exit provisions. Can you retain scored history? Can you operate during a transition? A model you can't leave prices differently the second time.
- Update cadence, and whether you can decline an update that changes behaviour mid-portfolio.
- Support for examination, including whether the vendor will produce documentation and respond to regulator questions.
- Liability allocation, which will not shift the regulatory obligations above but does allocate commercial loss.
- Data rights. What does the vendor do with the data you send? Does it improve a model your competitors use?
The last is worth examining carefully. Sending applicant data to a vendor whose model improves with it may mean funding a capability your competitors also license — which is a reasonable trade if the terms reflect it and a poor one if nobody noticed.
The diligence checklist
- Development sample composition, in detail, including default counts.
- Outcome definition, in writing.
- Performance by segment, not only aggregate.
- Variable list, under confidentiality if necessary.
- Reason code basis.
- Fair lending testing the vendor conducted, and on what population.
- Development documentation sufficient for your model risk function.
- Validation on your own historical data before any commitment.
- Shadow deployment before live decisions.
- Monitoring plan, with your obligations specified.
- Commercial terms, particularly renewal and exit.
- Reference calls with institutions resembling yours.
Items one and eight carry most of the weight. Everything else is confirmatory once you know the sample and have seen the model work on your applicants — and a purchase made without either is a purchase made on a brochure.
Validate before you commit
HL Hunt AI Underwriting runs in shadow mode against your historical accounts before any live decision, with segment-level performance, calibration against your own outcome definition, and reason codes traced to actual applicant facts.
Frequently asked questions
What the development sample looked like. Every performance statistic describes that population, and transferability depends entirely on whether your applicants resemble it.
No. Fair lending, adverse action accuracy, validation, and monitoring all stay with the institution deploying it. A vendor's documentation supports your compliance rather than substituting for it.
When your population differs materially, you hold data no vendor has, or you have enough matured defaults — which is a count of defaults, not accounts.
Methodology, sample composition, outcome definition, variables, segment-level performance, and reason code basis. Refusal to share under confidentiality is itself a finding.
Key takeaways
- Performance statistics describe the vendor's population; ask what the development sample looked like before anything else.
- Outcome definitions vary on threshold, window, cures, and fraud treatment — rank ordering may transfer while calibration doesn't.
- Validate on your own historical accounts and test by segment, where purchased models most commonly disappoint.
- Fair lending, adverse action accuracy, validation, and monitoring don't transfer with a licence — only development effort does.
- A vendor who won't disclose variables under confidentiality is asking you to accept obligations you can't discharge.
- The build threshold is matured defaults rather than accounts, which puts building out of reach for most low-loss books.
Keep the obligations you can't transfer discharged
Get started with HL Hunt AI Underwriting for full development documentation, disclosed variables, per-segment monitoring on your population, and validation support built for examination rather than assembled during one.
This guide is educational and does not constitute legal or compliance advice. Model risk management expectations, third-party risk requirements, fair lending obligations, and adverse action requirements vary by institution type and regulator. Consult qualified counsel and your model risk function before adopting any credit model.