How Do You Measure AI Success in Debt Collection?
PTP, kept promise rate, RPC, attribution windows and A/B comparison: the methodology of measurement.
Whether AI is working in collections is not measured by how many calls it made. Measurement looks at the outcome of the conversation and at whether that outcome turned into money. This page sets out which indicators are tracked and how they are read.
There is no promise of numbers here. What is described is methodology: what your institution should look at when it measures with its own data.
Metrics that measure the outcome of the call
Promise to Pay rate (PTP)
Shows what percentage of the customers reached gave a commitment to pay. It is the first indicator of how persuasive the conversation was. On its own it misleads: obtaining a promise is easy, keeping it is the real matter.
Kept Promise rate (Kept PTP)
The percentage of promises given that turned into an actual payment. It has to be read together with PTP. A high PTP with a low kept rate shows that the conversation did not persuade the customer, only that they found the easy way to end the call. A healthy design lifts both together.
Direct payment amount
Collection taken during the call itself. Because it carries no risk of slipping, it is the most valuable outcome.
Metrics that measure contact quality
RPC rate (Right Party Connect)
How many of the numbers dialled actually reached the debtor. A low RPC points to a data quality problem before it points to conversation quality: how current the numbers are, how records are matched, how many duplicates exist.
Reachability rate
The percentage of dialled numbers that were answered. Calling hours, calling frequency and caller ID visibility all affect this rate directly.
Average call duration
Neither good nor bad on its own. A short duration may mean efficiency, or it may mean an unfinished conversation. It is read alongside outcome metrics.
Human capacity equivalent
How much agent workload the automated volume corresponds to. Used for operational planning and investment decisions.
How is a payment attributed to a call?
The attribution window
In collections the payment often does not arrive during the conversation. The customer hangs up and pays two days later from a branch, an ATM, mobile banking or by transfer. To connect that payment to the call, a window is defined.
Common practice is to track payments arriving within 1, 3, 7 and 30 day windows after the conversation separately. The longer the window, the weaker the attribution: a payment arriving on the thirtieth day has a far more debatable relationship to the call than one arriving on the first. For that reason it is more accurate to look at the distribution rather than a single window.
Channel independent measurement
Which channel the payment came through does not matter for attribution. What matters is whether the paying customer was called in that period. For that, the payment record in the collections system and the call record have to match on the same customer identifier.
How do you measure the real effect?
A/B comparison
The attribution window on its own is not enough, because some customers would have paid without being called at all. Seeing the real contribution requires a comparison.
The method is this: customers in a similar risk group are split into two. One group is called by the AI agent, the other is not called or is handled through the existing human team process. The collection rate of the two groups is compared over the same window.
The critical point is that the groups must be alike: they have to be split in a balanced way by days overdue, balance size, product type and prior payment behaviour. Otherwise what you are measuring is not the effect of the AI agent but the difference between the groups.
The infrastructure behind the measurement
For these measurements to be possible, the outcome of every conversation has to be recorded in a structured form: was contact made, who answered, was a commitment obtained, what amount and what date. When records are kept as free text, measurement becomes impossible. We cover the observability side on our analytics and observability page.
TAHSİLDAR writes the outcome of the conversation into your collections system in structured form, which is the foundation needed for the metrics above to be calculated with your institution's own data. For how it works, see the TAHSİLDAR page, and for the scope of the automation, the automated collections system page.