Predictive Risk Models for Delivery
Built and benchmarked 4 classification models on a heavily imbalanced banking dataset, reaching 99% AUC — and worked out what it would take to point the same approach at early warning in program delivery.
What it is
A supervised classification problem on a real, heavily imbalanced banking dataset: rare positive cases, plenty of noise, and an accuracy score that looks excellent if you predict “no” every single time. I owned it end to end — problem framing, feature engineering, model selection and evaluation — as part of my PG Diploma in Data Science at IIIT Bangalore.
I benchmarked 4 models rather than picking one:
| Model | Why it was in the comparison |
|---|---|
| Logistic Regression | Interpretable baseline — if it wins, use it |
| Decision Tree | Readable rules, easy to explain to a non-technical stakeholder |
| Random Forest | Variance reduction, handles the imbalance better |
| XGBoost | Best expected performance, hardest to explain |
The final model reached a 99% AUC score.
Why AUC and not accuracy
This is the part that transfers to everything else I do. On an imbalanced dataset, accuracy is actively misleading — a model that never predicts the rare class can score in the high nineties and be completely useless.
Choosing the evaluation metric before training is what stops you from celebrating a model that has learned to say no. That decision is not a technical detail; it is the thing that determines whether the work is worth anything, and it belongs to whoever owns the problem.
Where it points
The reason I built this rather than something more novel is that the shape of the problem matches something I deal with constantly: rare, costly events buried in a lot of routine signal.
In program delivery that is a release that is going to slip, an incident that is about to recur, a workstream drifting before anyone raises it. The ingredients are the same — imbalanced classes, expensive false negatives, cheap-but-annoying false positives, and a threshold decision that is really a business trade-off rather than a modelling one.
What I learned
Pick the metric before you train. Every subsequent decision is downstream of it, and changing it afterwards is how teams end up defending a number that does not mean anything.
Benchmark the simple model honestly. Logistic regression is a genuinely useful answer when the gap is small, because interpretability is worth real performance in any setting where somebody has to act on the output.
The threshold is a product decision. Where you set it is a choice about how many false alarms your users will tolerate before they stop trusting the system. That is not a question the model can answer.