The Cold Start: Underwriting a Product With No Loss History
The Cold Start: Underwriting a Product With No Loss History
Every credit product's first cohort is underwritten without evidence. There are no outcomes, no loss curve, no scorecard fitted to this population and this structure — and yet decisions get made and money goes out. The choice is not between judgment and data. It's between judgment that's written down and judgment that isn't. Undocumented judgment can't be evaluated when outcomes arrive, can't be defended to an examiner, and can't be improved, because nobody recorded what was believed. This guide covers making the initial judgment explicit and testable, choosing benchmarks that don't mislead, structuring early cohorts to produce information rather than just volume, and knowing when you actually have enough data.
What you'll learn
- Judgment is unavoidable
- Building an expert scorecard
- Choosing benchmarks that don't mislead
- What a generic score does and doesn't give you
- Structuring the first cohorts
- Why narrow approval is a trap
- Capture everything
- When you have enough data
- Governance for a judgment model
- Frequently asked questions
Judgment is unavoidable
Worth establishing first, because the alternative framing produces bad practice.
Teams launching a product frequently describe their approach as data-driven because they're using a purchased score, or a vendor model, or rules derived from an adjacent portfolio. Each of those embeds judgment — about which score, which vendor, which portfolio is comparable — and the judgment is doing more work than the data. A team that believes it's being empirical when it's actually making assumptions is worse off than one that knows it's guessing, because the first won't document the reasoning.
So the honest position: the first cohort is priced on a hypothesis. That's fine, and it's the normal condition of launching anything. What matters is:
- The hypothesis is written down in enough detail to be proven wrong.
- Exposure is limited so being wrong is survivable.
- The cohort is structured to produce the information that will replace the hypothesis.
- Nobody claims more confidence than the situation supports.
That last point has a governance dimension. A model risk function encountering a "data-driven" launch with no outcome data will find the gap. Presenting the same launch as a documented expert judgment with a defined validation plan is both more honest and considerably easier to defend.
Building an expert scorecard
The practical instrument: a documented set of characteristics, weights, and reasoning that functions as a model until a real one exists.
- Write the repayment story. How does this product actually get repaid? What has to be true about the borrower for that to happen? This is the step teams skip, and everything else depends on it.
- List the characteristics you believe matter, and for each, state the mechanism — not "income matters" but "income above X because the payment is Y and we believe a payment above Z% of income fails."
- Assign weights, acknowledging they're judgment.
- Set the cutoff against your target loss rate, using the benchmarks below.
- Record predicted performance by band, explicitly. This is the crucial step: write down what you expect each score band to do, so that when outcomes arrive you can compare prediction against reality rather than reconstructing what you thought.
- Note what would falsify each assumption, so you know what you're watching for.
- Date it and version it.
Step five is the one that converts an expert scorecard from a decision rule into a research instrument. A team that records "we expect band C to lose 9%" and observes 22% has learned something specific. A team that recorded nothing has only learned that losses were higher than hoped, which doesn't locate the error.
The characteristics worth considering for most consumer products: capacity relative to the payment, stability of the income source, existing obligations, the liquidity buffer our illiquidity analysis argues predicts shock absorption better than wealth, prior credit behavior where a file exists, and the repayment mechanism itself.
Choosing benchmarks that don't mislead
Borrowed loss rates are the standard starting point and the standard source of large early errors. The errors are predictable, which means they're avoidable.
The wrong basis for comparison: loan size, borrower score band, or industry category. The right basis: repayment mechanism, borrower situation, and collection dynamics.
| Dimension | Why it drives loss more than size does |
|---|---|
| Repayment mechanism | Automatic debit timed to a payroll deposit fails at a fraction of the rate of a monthly statement requiring an action |
| Position in the payment hierarchy | Obligations that threaten transport or housing get paid first, per the triage ordering in our shortfall guidance |
| Whether it reports | A reporting tradeline carries the exercise cost our option analysis describes; a non-reporting one doesn't |
| Term | Longer exposure means more opportunity for circumstances to change |
| Acquisition channel | Selection differs enormously between organic, partner, and paid channels at identical stated criteria |
| Purpose | Distress borrowing and planned purchases perform differently at the same borrower quality |
The single most common benchmarking error: taking loss rates from a secured product and applying them to an unsecured one at the same borrower score band. The borrower is the same; the exercise cost isn't, and the difference is large.
The second most common: taking rates from an established product with a mature, self-selected customer base and applying them to a launch reaching a different population through a different channel. Channel selection alone can move loss rates by multiples at identical stated criteria.
The practical instruction: build a range rather than a point estimate — a base case, a pessimistic case at roughly twice the base, and a plan for what you'd do at the pessimistic case. Launches that assume a point estimate and discover reality outside it have no prepared response, which is when the expensive decisions get made quickly.
What a generic score does and doesn't give you
Purchased scores are the natural fallback and it's worth being precise about what they provide.
What they give you: a rank ordering. A population scored higher will, in general, perform better than one scored lower. That's genuinely valuable and it's the main thing you need at launch.
What they don't give you:
- Calibration for your product. A score band's historical default rate reflects the products in the developer's sample, not yours. The rank ordering transfers; the level does not.
- Coverage of thin-file applicants, which may be much of your target population, per our thin-file analysis.
- Anything about the repayment mechanism, which the table above says drives loss substantially.
- Capacity for this specific payment relative to this borrower's obligations.
- Liquidity position, which no bureau file contains.
The practical synthesis: use the generic score for rank ordering and your own judgment for the cutoff and the level. The common error is trusting the score's implied default rate, which produces forecasts that miss by a wide margin in whichever direction your product differs from the developer's sample. The cash flow inputs in our underwriting guide address several of the gaps directly and are available at launch, which makes them unusually valuable in a cold start.
Structuring the first cohorts
Design the launch to be cheap to be wrong about and fast to learn from.
- Small amounts. Caps the loss per account and lets you fund more accounts within the same budget — which matters because defaults, not dollars, are what you learn from.
- Short terms. Shortens time to outcome. A twelve-month product tells you what you need to know far sooner than a sixty-month one, and early cohorts are for learning rather than earning.
- Volume sufficient for defaults. See the sizing discussion below.
- Staged rollout with defined checkpoints and pre-agreed stopping rules — the circuit-breaker discipline in our monitoring guide.
- Single channel first. Adding channels early makes attribution impossible when performance differs.
- Heavy fraud screening. New products attract testing, and early losses are disproportionately fraud rather than credit — per our fraud guide. Misreading fraud as credit loss leads to tightening credit criteria that weren't the problem.
- Servicing built before launch, since a product that can't collect will produce loss rates reflecting your operations rather than your underwriting.
That last point is underrated. A meaningful share of "the underwriting was wrong" conclusions on new products are actually servicing failures — payments that didn't process, contact information never verified, no reminder before the first due date. Separating the two requires instrumenting the servicing path before the first account funds.
Why narrow approval is a trap
The instinct at launch is to approve only the strongest applicants. It's the right instinct on exposure and the wrong one on learning, and the distinction matters.
A launch approving only a narrow top band produces a cohort that tells you about that band and nothing else. When it performs well, you've learned that good applicants repay — which you assumed. You've learned nothing about where the boundary should actually be, and you've created exactly the censored sample our reject inference analysis describes, at the moment when it's cheapest to avoid.
The alternative: limit exposure through size and term rather than through narrowness of approval. A launch funding $400 loans across a wide band of applicants generates far more information per dollar of loss than one funding $4,000 loans to a narrow band — and it costs less to be wrong.
Concretely:
- Deliberately approve some applicants below your intended cutoff, at reduced size, and tag them.
- Vary limits and terms across the cohort so you learn about structure as well as about borrowers.
- Budget the expected extra loss as a research cost, the same framing our reject inference guide applies.
- Tag every experimental account permanently, or the information is unusable later.
The cheapest time to learn where your boundary belongs is before you've committed to one. A lender three years in with a narrow book faces the same question at far higher cost.
Capture everything
The lowest-cost, highest-return discipline available, and it has a hard deadline.
Data not captured at origination is unavailable forever. You cannot reconstruct an applicant's stated income, their bank balance at application, or which page they abandoned on. Capture is cheap; reconstruction is impossible.
What to capture even without a current use:
- Every application field, including declines and abandonments.
- Full bureau file as of the decision, archived rather than summarized to a score.
- Cash flow data where consented — balances, volatility, and the minimum balance pattern.
- The decision and its reasons, including overrides and who made them.
- Channel, campaign, and device, since selection effects show up here.
- Timing and funnel behavior.
- Every servicing event — payment timing, failed attempts, contacts, and cures. The behavioral data our renewal analysis shows outperforms bureau data starts accumulating here.
- The scorecard version applied to each decision.
The judgment to apply: capture broadly, use narrowly. Capturing a variable creates no obligation to use it, and using variables you haven't validated is its own risk. But the option to use it later requires having it, and that option costs almost nothing to preserve.
When you have enough data
The question everyone asks in terms of accounts, which is the wrong unit.
What matters is the number of defaults, not the number of accounts — defaults are what the model learns from. A book of 10,000 accounts with 40 defaults contains less modeling information than 1,000 accounts with 200 defaults.
Two implications that surprise people:
Low-loss products take much longer to reach modelable data. A product losing 2% needs roughly ten times the volume of one losing 20% to accumulate the same number of defaults. A well-performing product can be starved of the information needed to optimize it — which is an argument for the deliberate variation described above, since it generates defaults where you need them.
The accounts must be mature. Building on a large young book means most future defaults haven't happened, so you'd be modeling early-cycle behavior and calling it credit risk. Compare cohorts at matched months on book, per our loss forecasting guide.
A staged progression that reflects the actual sequence:
- Expert scorecard with documented predictions.
- Early monitoring — operational metrics and fraud, not credit conclusions.
- First calibration once cohorts have partial maturity: adjust the level, not the structure.
- Simple empirical model once defaults accumulate — few variables, robustly estimated, beats a complex model on thin data.
- Refinement as data deepens, tested through the challenger design in our deployment guide.
Skipping to step four too early is the characteristic failure — a model fitted to 60 defaults will find patterns that are noise, and it will find them confidently.
Governance for a judgment model
An expert scorecard is a model for governance purposes and should be treated as one:
- Document the methodology — characteristics, weights, reasoning, and the predictions.
- State the limitations explicitly, including that it rests on judgment rather than fitted data.
- Define the validation plan and when it triggers.
- Ensure reason codes are accurate for adverse action, per our notices guide — an expert scorecard must still produce specific principal reasons.
- Run fair lending review at launch, not after. Judgment-based criteria deserve more scrutiny than fitted ones, because the reasoning is a human's rather than a fitting procedure's, and the analysis in our governance framework applies fully.
- Version everything, so any decision can be reproduced.
- Set review triggers on volume, maturity, and performance deviation.
The framing that makes this easier: an expert scorecard with documented predictions and a validation plan is a stronger governance position than an undocumented process that claims to be data-driven. Examiners respond well to acknowledged uncertainty with a plan and poorly to unsupported confidence.
Launch with structure, not with guesses
HL Hunt AI Underwriting supports documented expert scorecards with versioning, per-decision reason attribution, full application and bureau archiving, and cohort tagging — so the first cohort generates the evidence that replaces the judgment rather than just producing volume.
Frequently asked questions
With structured judgment, documented so it can be tested. Write down the characteristics, weights, reasoning, and — critically — your predicted performance by band, so outcomes can locate the error.
Defaults matter, not accounts. Ten thousand accounts with forty defaults contains less information than one thousand with two hundred. The accounts must also be mature enough for outcomes to have emerged.
As a starting range only, and choose the comparison by repayment mechanism and borrower situation rather than by loan size. Applying secured product rates to an unsecured product at the same score band is the classic error.
Limit exposure through size and term, not through narrowness. Approving only a narrow band creates a censored sample at the moment it's cheapest to avoid.
Key takeaways
- The first cohort is priced on judgment either way — the only choice is whether the judgment is documented well enough to be proven wrong.
- Record predicted performance by band before launch; that's what converts an expert scorecard into a research instrument.
- Choose benchmarks by repayment mechanism and borrower situation, not loan size — and build a range rather than a point estimate.
- Generic scores transfer rank ordering but not calibration, so use them for ordering and your own judgment for the level.
- Limit exposure through small amounts and short terms rather than narrow approval, or you create a censored sample immediately.
- Data not captured at origination is gone forever, and defaults rather than accounts determine when you can build a real model.
Instrument the launch before the first account funds
Get started with HL Hunt AI Underwriting to capture full application, bureau, and cash flow data alongside every decision from day one — so the information you need in month eighteen exists, rather than having to be reconstructed from what was kept.
This guide is educational and does not constitute legal or compliance advice. Model risk management, adverse action, and fair lending obligations apply to judgment-based underwriting criteria as fully as to statistically derived models; consult qualified counsel and your model risk function before launching a credit product.