AI Knowledge · Case study 04 · Ensemble methods

Credit risk scoring: why the most accurate model isn’t the best one

I lined up five models on 1,000 loan applicants. The one with the best accuracy wasn’t the one I’d pick, and here’s why.

0.776

AUC for Random Forest, best of five models

+0.17

AUC lift from ensembles over a single tree

1,000

loan applicants, 16 attributes

10

folds of cross-validation to tune the tree

Two mistakes that don’t cost the same

Approve a loan that defaults and you lose the principal. Decline a good applicant and you lose the interest. Those mistakes have different price tags, so the job isn’t just “build an accurate model.” It’s finding the model and the cutoff that fit those costs, and being able to explain it to a credit committee.

Five models, one test set

I used the German Credit benchmark: 1,000 applicants and 16 attributes covering balances, loan amount and length, credit history, purpose, employment, age and housing. The outcome is default or no default. 80/20 split, with a fixed random seed so the results are repeatable.

Then I lined up five models against the same test set: one decision tree; a tuned tree (GridSearchCV with 10-fold cross-validation over depth, leaf count and split rule); and three ensembles of 100 trees each: Bagging, AdaBoost and Random Forest. I picked AUC as the scorecard before I started.

Test AUC, higher is better

Every ensemble beat a single tree; Random Forest ranked applicants best

Single decision tree0.606
Tuned decision tree0.759
AdaBoost0.763
Bagging0.764
Random Forest0.776
ModelTest accuracyTest AUC
Single decision tree0.6800.606
Tuned decision tree (depth 10, 17 leaves)0.7400.759
AdaBoost0.7650.763
Bagging0.7700.764
Random Forest0.7550.776

What I learned

A single tree is fragile. Mine scored a perfect 1.000 on training data and 0.68 on new applicants, and five-fold cross-validation confirmed that 0.68 wasn’t a fluke. Every ensemble beat it by 0.16 to 0.17 AUC.

The part I’d underline: the most accurate model wasn’t the best one. Bagging had the highest accuracy (0.770). Random Forest had the best AUC (0.776). The approval cutoff gets set by what each mistake costs, so what I care about is how well the model ranks applicants across every possible cutoff. That’s what AUC measures, so Random Forest wins.

What moved the score most: loan amount (0.142), loan length (0.109), applicant age (0.105) and checking-account status (0.067).

How I’d use it

This pattern fits any approve-or-decline call a business makes over and over: lending, credit limits, insurance, tenant screening, supplier risk. I’d let the model handle the clear yeses and clear nos, and send the borderline cases to a person. Decisions get faster, and accountability stays with someone qualified to own it.

Before it goes live

  • Validation. Credit models need documented development, independent validation and ongoing monitoring.
  • Fair lending. Inputs like age need review, and every decline needs a reason someone can explain.
  • Explainability. Ensembles score better but are harder to explain. I’d pair this with explanation methods or keep a simpler model running alongside it.
  • Scale. 1,000 applicants proves the method and the metric discipline. It isn’t a production credit model.
More case studies

Keep reading

Back to AI Knowledge