The Diversification Illusion: Why 50,000 Loans Is Not 50,000 Independent Bets

The Diversification Illusion: Why 50,000 Loans Is Not 50,000 Independent Bets | HL Hunt
Institutional Outlook

The Diversification Illusion: Why 50,000 Loans Is Not 50,000 Independent Bets

A lender with fifty thousand accounts spread across every state describes the book as well diversified, and in one sense it is — no single borrower can hurt it. But diversification only removes the risk that's independent across borrowers, and a large share of consumer default risk isn't. Households share exposure to the labour market, to fuel and energy costs, to rates, to local employment. When those move, many borrowers move together. The arithmetic below shows the consequence: past a few thousand accounts, adding loans stops reducing loss volatility almost entirely. What's left is the correlated component, it doesn't shrink with scale, and it's the part that determines whether a portfolio survives a bad year.

By the HL Hunt Research Desk · 25 min read · Updated August 2026

The core thesis

Decompose a borrower's default risk into two parts:

  • Idiosyncratic — the divorce, the illness, the individual job loss unrelated to anything happening elsewhere.
  • Systematic — the component driven by conditions affecting many households at once.

Diversification averages away the first. It does nothing to the second, ever, at any scale.

The portfolio loss variance for N accounts with individual default variance σ² and pairwise correlation ρ behaves as:

Var(portfolio loss rate) ≈ σ² [ ρ + (1 − ρ)/N ]

Read the bracket. As N grows, the second term vanishes — that's the idiosyncratic risk being averaged out. The first term, ρ, doesn't move at all.

Which yields the central result: portfolio loss volatility converges to a floor set entirely by correlation, and it converges fast. A lender at that floor gains nothing from further scale, and no amount of geographic spread changes it if the correlation runs through factors that operate nationally.

Our claims: that consumer lenders reach the floor at portfolio sizes well below where they think they do; that the correlations they monitor are not the ones doing the damage; and that underwriting improvements reduce the level of losses without reducing their correlation, which means a lender can improve credit quality substantially and remain exactly as exposed to a bad year.

Better underwriting lowers where losses sit. It doesn't lower how much they move together. Those are different problems and only one of them is usually being worked on.

Where diversification stops working

Put numbers on it. Take a book with an 8% expected default rate, so individual default variance σ² = 0.08 × 0.92 = 0.0736. Compute the standard deviation of the portfolio loss rate at various sizes and correlations.

Accountsρ = 0 (independent)ρ = 0.01ρ = 0.03ρ = 0.05
1002.71%2.84%3.09%3.32%
1,0000.86%1.17%1.65%2.00%
5,0000.38%0.90%1.52%1.92%
25,0000.17%0.85%1.49%1.90%
50,0000.12%0.85%1.49%1.90%
500,0000.04%0.84%1.48%1.89%

Three findings, and the third is the one that matters most.

Under independence, scale works beautifully. Loss volatility falls by a factor of more than twenty from 100 to 50,000 accounts. This is the intuition every lender carries.

Under even slight correlation, it stops early. At ρ = 0.03 — a correlation so small it would look like noise between any two borrowers — volatility falls from 3.09% at 100 accounts to 1.52% at 5,000, and then to 1.49% at 50,000. Going from 5,000 to 50,000 accounts, a tenfold increase, reduced loss volatility by about 2%.

The floor is set by correlation and nothing else. At half a million accounts, the ρ = 0.05 book still has 1.89% loss volatility — more than fifteen times the independent case. Scale bought essentially nothing beyond a few thousand accounts.

The operational implication is worth stating bluntly: a lender growing from five thousand to fifty thousand accounts has multiplied its exposure without materially improving the stability of its loss rate. Growth of that kind is a revenue strategy, not a risk strategy, and describing it as diversification is a mistake with capital consequences.

10× the accounts, 2% less volatility
At a correlation of just 0.03, going from 5,000 to 50,000 accounts reduced loss rate volatility from 1.52% to 1.49%. The idiosyncratic risk was already gone; the rest never diversifies.

Why small correlations matter so much

The table shows correlations of 0.01 to 0.05 producing dramatic effects, which is counterintuitive enough to need explaining — and the explanation is the reason this gets underestimated so consistently.

Accounts grow linearly. Pairs grow quadratically. A portfolio of N accounts contains N(N−1)/2 pairs. At 1,000 accounts that's about 500,000 pairs. At 50,000 accounts it's nearly 1.25 billion.

Portfolio variance sums the individual variances and every pairwise covariance. So the correlated contribution is multiplied by a number growing with the square of the portfolio, while the diversifying contribution grows only linearly. Beyond a modest size, the pairwise terms dominate completely, and a correlation invisible between any two borrowers becomes the entire story.

This is why intuitions built from examining individual accounts fail. A credit officer reviewing two files and finding no meaningful connection between them is observing something true and irrelevant — the correlation doesn't have to be visible in any pair to dominate the portfolio.

It also explains a specific institutional blind spot. Risk functions that model borrower-level probability of default with great care frequently treat portfolio loss as the sum of those probabilities, which implicitly assumes independence. The sophistication goes into the marginal distribution while the dependence structure — where the tail risk lives — gets a much rougher treatment or none at all.

The common factors

What actually links household defaults, roughly in order of strength:

  • Employment conditions. The dominant factor. Job loss is the leading cause of consumer default, and it arrives in waves rather than independently — which is precisely the shock our income volatility analysis describes at the household level, aggregated.
  • Local economic conditions. A plant closure defaults its borrowers together regardless of their individual credit quality.
  • Interest rates, which move payments on variable obligations and refinancing options simultaneously.
  • Essential cost shocks. Fuel, energy, food, and — increasingly — the property insurance premiums our insurance analysis documents, which compress every affected household's budget at once.
  • Asset prices, particularly housing, which determine whether stressed borrowers have the extraction option our equity analysis describes.
  • Credit availability itself. The most underappreciated: when credit tightens, borrowers who would have refinanced or borrowed to bridge a gap can't, so the same underlying stress produces more defaults. Tightening is itself a common factor, and it correlates with all the others.
  • Policy changes — benefit expirations, payment resumptions, forbearance program ends — which affect defined populations on defined dates.

The sixth deserves emphasis because it's a feedback loop rather than an exogenous shock. Lenders tightening in response to rising losses cause additional defaults among borrowers who would otherwise have refinanced, which raises losses further. The industry's collective risk response is itself a correlated factor, and no individual lender's underwriting can address it.

The correlations lenders miss

Geographic and employer concentration get monitored. These generally don't, and they're frequently larger.

Vintage. Loans originated in the same period share underwriting standards, competitive conditions, and the economic environment at origination — which is why cohort analysis exists in the first place, per our forecasting guide. A lender that tripled originations in one strong quarter has a concentration in that vintage, and vintage effects are among the strongest observable patterns in consumer credit. Rapid growth is a correlation event, not merely a volume event.

Channel. Acquisition channel selects for borrower characteristics that persist and that aren't fully captured in the file — the effect our cold start analysis notes can move loss rates by multiples at identical stated criteria. A book heavily weighted to one channel is concentrated in whatever that channel selects for, observed or not.

Product structure and repayment mechanism. Loans sharing a repayment path fail together when it fails. A portfolio repaid by automatic debit is exposed to the account-access dynamics in our bank payments guide; one repaid by statement is exposed to different failures. This is a correlation almost nobody measures, and it's mechanical rather than economic.

Industry. Borrowers concentrated in one sector — even across many employers and states — share that sector's cycle. A lender serving a particular occupation is concentrated in it regardless of geographic spread.

Model dependence. The subtle one. If every lender uses similar models on similar data, they approve similar borrowers and decline similar ones. Portfolios that look independent are in fact selected by a shared process, so a model failure is a market-wide correlated event and the increasing convergence of underwriting technology raises it over time.

The pattern across all five: they create concentration inside a book that looks well spread along the dimensions being monitored. A lender reporting on state and employer distribution can be heavily concentrated in vintage, channel, and structure without any report flagging it.

Why better underwriting doesn't help

The conclusion most likely to be resisted, and the one with the clearest logic.

Underwriting improvements shift the expected loss rate down. A book that was going to lose 8% now loses 6%. That's genuinely valuable and it is not a reduction in correlation.

The reason is that better underwriting identifies borrowers with lower idiosyncratic risk — more stable individual circumstances, better individual history. It doesn't identify borrowers with less exposure to the common factors, because those factors affect nearly everyone. Employment shocks hit prime borrowers too; they hit them less often, but the correlation of the hitting is unchanged.

Which produces a specific and important failure mode: a lender that improves credit quality and interprets the resulting lower loss volatility as reduced correlation will be surprised in the next downturn. The volatility fell because the level fell, not because the loss rate became more stable in relative terms.

Worth stating the exception honestly. Underwriting can reduce correlated exposure if it explicitly targets it — screening for concentration in a single employer, requiring the liquidity buffer our illiquidity analysis argues predicts shock absorption, or avoiding borrowers whose income depends on a single volatile sector. But that's a different underwriting objective from maximizing predictive accuracy, and almost no credit policy is written to pursue it. A model optimized for discrimination will not incidentally reduce correlation.

What actually reduces it

  1. Spread originations across time. The most actionable and least practiced. Concentrating volume in a strong quarter creates a vintage concentration that no geographic spread offsets. Steady origination is a risk decision, and it conflicts directly with growth targets — which is why it rarely wins.
  2. Diversify acquisition channels, since channel selection is a hidden concentration.
  3. Vary product structure and repayment mechanism, so mechanical failures don't hit the whole book.
  4. Monitor employer, industry, and local economy concentration, not just state.
  5. Screen for household-level liquidity, which is the borrower characteristic most directly related to surviving a common shock.
  6. Size capital to the correlated scenario. The practical consequence of the table above: a lender holding capital against the independent-case volatility is under-capitalized by an order of magnitude. Stress testing should assume correlated deterioration, not a multiple of the average.
  7. Build the workout capacity in advance, because correlated stress means many borrowers need it simultaneously — and the arithmetic in our forbearance analysis only pays if the operational capacity exists when the wave arrives.
  8. Maintain funding that doesn't withdraw under stress, since funding and credit conditions correlate and a lender losing funding while losses rise faces both at once.

The first is where the tension lives. Every one of these trades growth or margin for stability, and they compete with targets that are usually stated in volume terms. A risk function that cannot influence the origination calendar cannot manage the largest correlation in the book.

The strongest objections

"Consumer credit is empirically well diversified — the data show low correlations." Measured correlations in normal periods are indeed low. That's the problem rather than the rebuttal: correlation in consumer credit is state-dependent, rising sharply in stress. A parameter estimated from calm periods understates the one that applies when it matters, which is the same failure that has recurred across asset classes. A model calibrated on benign data is calibrated on the wrong regime.

"Scale still helps through operational leverage and data." Fully conceded, and it's important. Scale spreads the fixed costs in our selection analysis, generates the defaults needed to build models, and improves negotiating position. The claim here is narrow: scale doesn't reduce loss volatility past a modest point. Those are different benefits and the argument for growth doesn't depend on the diversification claim — which is precisely why the diversification claim should be dropped rather than defended.

"The correlation figures used are asserted rather than estimated." True, and stated deliberately. We've used illustrative values because published estimates vary enormously by product, period, and method, and quoting one would imply more precision than exists. The finding doesn't depend on the value. At any correlation above roughly 0.01, diversification stops working well before the portfolio sizes lenders operate at — and 0.01 is a very low bar. The qualitative result is robust across the entire plausible range, which is what makes it worth acting on despite the parameter uncertainty.

Testable implications

  1. Loss rate volatility should stop falling with portfolio size past a few thousand accounts. Any lender that has grown substantially can check this in their own history, and it's the direct test.
  2. Vintage should explain more loss variance than geography in most consumer books.
  3. Correlation estimated from stress periods should exceed correlation from calm periods by a wide margin — the state-dependence claim.
  4. Channel should predict residual loss after controlling for observable borrower characteristics, evidencing hidden selection.
  5. Books with more concentrated origination timing should show higher loss volatility than books of the same size originated steadily.
  6. Underwriting improvements should lower expected loss without lowering the ratio of loss volatility to expected loss. This is the sharpest test of the section above, and the one we'd most want a lender to run on their own data.

The first and sixth are both computable from any lender's existing history within a day, which makes this unusual among the propositions in this series — no experiment is required, only a query against data already held.

The broader point is about what "diversified" is doing in a sentence. A consumer lending book spread across fifty states and fifty thousand borrowers genuinely has no single-name risk. It also has almost the same loss volatility it had at five thousand accounts, and it will find that out in the same quarter as everyone else.

Frequently asked questions

Why doesn't adding more loans reduce portfolio risk indefinitely?

Diversification only removes independent risk. The component driven by common factors — employment, rates, energy costs — doesn't shrink with scale, so volatility converges to a floor set by correlation.

How correlated are consumer loan defaults?

Low pairwise, consequential in aggregate. Pairs grow with the square of the portfolio while accounts grow linearly, so a correlation invisible between any two borrowers dominates a large book.

What correlations do consumer lenders most often miss?

Vintage, channel, and repayment mechanism — all of which create concentration inside a book that looks well spread by state and employer.

What actually reduces consumer portfolio risk if scale does not?

Spreading originations across time, diversifying channels and structures, monitoring industry and local concentration, and sizing capital to the correlated scenario rather than to the average.

Key takeaways

  • Portfolio loss volatility converges to a floor set entirely by correlation, and it converges within a few thousand accounts.
  • At a correlation of 0.03, growing from 5,000 to 50,000 accounts reduced loss volatility by about 2% — the diversification was already exhausted.
  • Pairs grow quadratically while accounts grow linearly, which is why correlations invisible between two borrowers dominate a large portfolio.
  • Vintage, channel, and repayment mechanism are larger hidden concentrations than the geographic ones lenders monitor.
  • Better underwriting lowers the level of losses without lowering their correlation — those are separate problems and only one is usually being addressed.
  • Spreading originations across time is the most actionable correlation reduction available, and it conflicts directly with growth targets.

This report presents an analytical framework and the authors' interpretation; it is not investment, accounting, or risk management advice. Correlation values used in the worked table are illustrative parameters chosen to demonstrate the mechanism, not estimates for any specific portfolio; published estimates vary substantially by product, period, and methodology.