All projects

Predictive Risk Models for Delivery

Built and benchmarked 4 classification models on a heavily imbalanced banking dataset, reaching 99% AUC — and worked out what it would take to point the same approach at early warning in program delivery.

)) }

What it is

A supervised classification problem on a real, heavily imbalanced banking dataset: rare positive cases, plenty of noise, and an accuracy score that looks excellent if you predict “no” every single time. I owned it end to end — problem framing, feature engineering, model selection and evaluation — as part of my PG Diploma in Data Science at IIIT Bangalore.

I benchmarked 4 models rather than picking one:

ModelWhy it was in the comparison
Logistic RegressionInterpretable baseline — if it wins, use it
Decision TreeReadable rules, easy to explain to a non-technical stakeholder
Random ForestVariance reduction, handles the imbalance better
XGBoostBest expected performance, hardest to explain

The final model reached a 99% AUC score.

Why AUC and not accuracy

This is the part that transfers to everything else I do. On an imbalanced dataset, accuracy is actively misleading — a model that never predicts the rare class can score in the high nineties and be completely useless.

Choosing the evaluation metric before training is what stops you from celebrating a model that has learned to say no. That decision is not a technical detail; it is the thing that determines whether the work is worth anything, and it belongs to whoever owns the problem.

Where it points

The reason I built this rather than something more novel is that the shape of the problem matches something I deal with constantly: rare, costly events buried in a lot of routine signal.

In program delivery that is a release that is going to slip, an incident that is about to recur, a workstream drifting before anyone raises it. The ingredients are the same — imbalanced classes, expensive false negatives, cheap-but-annoying false positives, and a threshold decision that is really a business trade-off rather than a modelling one.

What I learned

Pick the metric before you train. Every subsequent decision is downstream of it, and changing it afterwards is how teams end up defending a number that does not mean anything.

Benchmark the simple model honestly. Logistic regression is a genuinely useful answer when the gap is small, because interpretability is worth real performance in any setting where somebody has to act on the output.

The threshold is a product decision. Where you set it is a choice about how many false alarms your users will tolerate before they stop trusting the system. That is not a question the model can answer.