Who will say yes? Ranking cross-sell prospects
Only 9.6% of customers said yes to a loan offer. I built a model that finds most of them, so the team calls the best prospects first.
AUC on held-out customers
of 108 accepting customers found
of flagged customers accepted
of all customers accepted an offer
Contact fewer people, convert more
Only 9.6% of this bank’s customers said yes to a loan offer. Offer it to everyone and you waste most of the budget and annoy people who never wanted it. I wanted a model that ranks customers by how likely they are to say yes, so the team calls the best prospects first.
The data
5,000 customers and 11 attributes: age, experience, income, family size, credit card spending, education, mortgage, and whether they have a securities account, a CD, online banking or a credit card. 480 accepted. No missing values. I split it 75/25, holding back 1,250 customers for testing.
Same trap as churn, only bigger: 1,142 of the 1,250 test customers said no, so “offer it to nobody” is 91.4% accurate. I leaned on the confusion matrix and AUC instead.
Results
| Model | Train accuracy | Test accuracy | Test AUC |
|---|---|---|---|
| Unpruned decision tree | 1.000 | 0.979 | not reported |
| Pruned decision tree (max depth 4) | 0.987 | 0.981 | 0.990 |
Pruning to four levels gave me the same results on new customers with a fraction of the rules. Simpler to explain and less likely to memorize noise, so it was an easy call.
The confusion matrix is the business story. The tree found 99 of the 108 people who said yes (91.7%) and made only 17 offers to people who said no, 1.5% of everyone who declined. About 85% of the people it flagged accepted.
Income and education do most of the work
How I’d use it
Any next-best-offer decision works the same way. For a sales or brokerage team with more prospects than hours, it’s a ranked call list: the same budget, better conversion and fewer unwanted calls. And because it’s a tree, marketing can see exactly why someone is at the top of the list.
Before it goes live
- Teaching data. This is a clean, well-known dataset, and an AUC of 0.99 reflects that. Real customer data will score lower.
- Sensitive attributes. Age and family size are sensitive. A model that shapes credit offers needs its features and outcomes reviewed, including checks for proxies.
- Likely isn’t eligible. A high score doesn’t mean someone qualifies. Approval stays with credit policy.
- Monitoring. Campaign response shifts, so the scores need tracking and periodic retraining.