What Call Monitoring Should Actually Look For | HL Hunt
What Call Monitoring Should Actually Look For
A typical collections quality form asks whether the collector identified themselves, gave the required disclosure, verified the party, and used approved language. Every box can be ticked on a call that went badly — pressure applied to someone who correctly said they couldn't pay, an arrangement agreed that was never going to hold, a disclosed hardship nobody acted on. The checklist measures whether the words were said. It has no view on whether the call should have gone the way it did, and that's where both the recovery and the risk actually live.
What you'll learn
Two different questions
| Compliance check | Quality assessment | |
|---|---|---|
| Asks | Was the required thing said? | Was this handled well? |
| Nature | Largely binary | Judgment |
| Automatable | Largely yes | Largely no |
| Coverage achievable | All calls | A sample |
| Finds | Omissions and prohibited language | Judgment failures |
Conflating them produces the common outcome: a program that measures the automatable thing with human effort and never measures the thing humans are needed for.
The separation is worth making structural. Compliance checking should run across every call automatically; human review should be spent entirely on judgment — and an operation still using reviewers to confirm that disclosures were read is spending its scarcest resource on its most automatable task.
What to detect
Define the target behaviours before designing any form, because a scoring form built without a list will measure whatever is easy to observe.
What a program should be finding:
- Pressure applied after a clear statement of inability to pay. The failure our promise analysis identifies — collectors measured on commitments will obtain them from people who can't deliver, producing a promise and a broken one.
- Arrangements agreed that don't fit the stated circumstances, which per our plan design guide fail and cost the time they ran.
- Disclosed hardship that wasn't acted on — a mentioned job loss, illness, or bereavement that didn't change the call's direction.
- A dispute raised and not recorded as one, which converts a data quality signal into nothing, per our dispute analysis.
- A bankruptcy or attorney mention not escalated, which per our bankruptcy guide should stop the call immediately.
- Options not offered that the account was eligible for.
- Vulnerability indicators not recognized.
- Tone that was technically permissible and clearly wrong.
What these have in common: none is detectable from a script checklist, and each is a decision the collector made. That's the definition of what human review should be looking at.
Sampling toward risk
The change with the largest effect per unit of effort.
Problems are concentrated; random sampling isn't. Reviewing a random 2% of calls spends nearly all of that effort confirming that ordinary calls were ordinary.
Indicators already in your data:
- Unusually long calls, which frequently mean difficulty.
- Calls that ended abruptly.
- Accounts that complained or disputed afterward — the most direct signal available, and per our complaint analysis most operations never listen to the call behind a complaint.
- Arrangements that broke on the first payment, which suggests they were never viable.
- Collectors whose numbers moved sharply in either direction.
- Repeated contact with the same person in a short period.
- Calls on accounts with vulnerability or hardship flags.
- New collectors, in their first weeks.
Keep a random component — a purely targeted sample can't tell you the base rate, and you need to know whether problems exist outside the flagged population. A reasonable split is most of the effort on risk-weighted selection and a minority on random.
The complaint-linked review deserves separate emphasis: pulling the call behind every complaint is cheap, high-yield, and almost never done. The complaint has already identified the call for you.
Scoring decisions
What a quality form should ask, phrased as judgments rather than as presence checks:
- Did the collector establish the person's actual situation? The diagnosis that determines everything, and it requires asking rather than proceeding.
- Was the response appropriate to that situation? Someone temporarily short and someone structurally unable to pay need different calls.
- Were the eligible options offered?
- Was the outcome realistic? An arrangement the person plainly couldn't meet is a failure regardless of it being agreed.
- Was anything disclosed that should have changed the handling?
- Was it recorded correctly? Disputes, hardship, and contact preferences all have to reach the system.
- Would you defend this call?
Question seven is the most useful single item on any form. It captures what checklists can't, it's answerable consistently, and a call that fails it while passing everything else is exactly what the program exists to find.
And question four connects monitoring to recovery rather than only to risk: an arrangement that fails costs the operation the weeks it ran, so scoring realism is a performance measure as much as a conduct one. That framing is what gets a quality program taken seriously by people whose metric is recovery.
What automation can and can't do
Speech analytics changes what's achievable, in one direction more than the other.
What it does well:
- Detecting required language across every call rather than a sample — a genuine step change from 2% coverage to complete coverage.
- Detecting prohibited language.
- Flagging keywords — bankruptcy, attorney, dispute, hardship — for escalation and for review selection.
- Surfacing patterns across collectors and time.
- Building the risk-weighted sample automatically.
What it doesn't do well:
- Judging whether the call should have gone that way, which requires understanding a situation rather than detecting a phrase.
- Assessing whether an arrangement was realistic.
- Distinguishing appropriate firmness from inappropriate pressure, where the same words differ entirely by context.
The useful arrangement: automation for coverage, humans for judgment. That's more than an efficiency point — it changes what human review is capable of finding, because reviewers freed from checklist work can be pointed at the calls automation flagged as unusual.
One caution. An automated system detecting only what it was configured to detect will report improving compliance while judgment failures continue undetected — which is the monotonic metric problem from our measurement analysis. Coverage of the measurable is not coverage.
Routing findings to causes
The step that determines whether monitoring changes anything.
Findings should be categorized by cause and routed to whatever can change it, which is frequently not the collector:
| Pattern | Likely cause | Route to |
|---|---|---|
| One collector, repeated | Individual | Coaching |
| Many collectors, same issue | Training or script | Training design |
| New collectors specifically | Onboarding gap | Onboarding |
| Options not offered | System makes it hard | Systems |
| Concentrated on one plan | Incentives | Compensation design |
| Rose after a change | The change | Whoever made it |
The fourth row is underappreciated. When a collector doesn't offer an option, the usual cause is that offering it takes eleven clicks and a supervisor approval — which is a systems finding, not a coaching one, and coaching it produces frustration and no change.
The general failure: routing every finding back as individual coaching treats systemic causes as personal failings, and the cause stays in place. An operation coaching the same issue across many collectors for months has misdiagnosed it, and the repetition is the evidence.
When the pattern is the plan
The finding a monitoring program should be capable of producing about its own organization.
Behaviour concentrated among collectors on a particular compensation arrangement is telling you about the arrangement. If pressure-after-inability appears disproportionately among collectors paid on promises obtained, the plan is the cause — and no amount of coaching changes what the plan pays for.
How to see it:
- Break findings down by compensation structure, not just by individual and team.
- Compare across teams on different arrangements.
- Look at timing within a period — behaviour clustering near targets is diagnostic.
- Cross-reference with complaint rate per contact, which per our complaint analysis is the counterweight metric.
This requires a monitoring function willing to report upward rather than only downward, which is an organizational condition rather than a technical one. A quality program that can only produce findings about collectors is a program that has been scoped to protect the decisions above them.
Building the program
- Separate compliance checking from quality assessment as distinct workflows.
- Automate compliance detection across all calls where feasible.
- Define target behaviours before designing the quality form.
- Build a risk-weighted sample with a random component retained.
- Review every call behind a complaint or dispute.
- Score judgments, including "would you defend this call".
- Categorize by cause, mandatory before closing a finding.
- Route to the function that can fix it.
- Break findings down by incentive structure.
- Track whether the rate of each finding type falls, which is the only test of whether the program works.
- Apply the same program to agencies, per our agency analysis, since their conduct is attributed to you.
- Calibrate reviewers periodically on the same calls, since a judgment-based form is only useful if reviewers agree.
Item twelve is the one that makes judgment scoring defensible. Without calibration, a quality score measures which reviewer heard the call — and item ten is the only measure that distinguishes a program that reduces problems from one that documents them.
Detect the pattern, not just the phrase
HL Hunt AI Debt Collection flags hardship, dispute, and escalation language across every interaction, builds review queues weighted toward complaint-linked and abnormal contacts, and reports findings by cause and by team — so review effort goes where problems actually are.
Frequently asked questions
Compliance and quality separately. The first is largely binary and automatable; the second asks whether the situation was handled well, which can't be reduced to a checklist.
Problems are concentrated and random sampling isn't, so most effort confirms that ordinary calls were ordinary. Risk indicators are already in data the operation holds.
Most of the compliance checking, very little of the quality assessment. Automation for coverage, humans for judgment — and it should build the review queue rather than replace it.
Categorized by cause and routed to what can change it. Coaching a systemic cause treats it as a personal failing and leaves it in place.
Key takeaways
- A checklist measures whether the words were said, not whether the call should have gone that way — and every box can be ticked on a bad call.
- Automate compliance across all calls and spend human review entirely on judgment.
- Sample toward risk using indicators you already hold, and review the call behind every complaint.
- "Would you defend this call?" is the most useful single item on a quality form.
- An unoffered option is usually a systems finding, not a coaching one — coaching it changes nothing.
- A program that can only produce findings about collectors has been scoped to protect the decisions above them.
Measure what changes outcomes
Get started with HL Hunt AI Debt Collection for coordinated outreach with interaction-level flagging, complaint-linked review queues, and cause-categorized reporting alongside recovery metrics.
This guide is educational and does not constitute legal or compliance advice. Call recording and monitoring are subject to consent and notification requirements that vary by state and by the parties' locations, and required disclosures in collections vary by jurisdiction and account type. Consult qualified counsel about your recording practices and your program design.