Towards Data Science

How to Improve Customer Retention in FinTech

A practical guide to combining pre-churn scoring with uplift modelling for smarter retention. The post How to Improve Customer Retention in FinTech appeared first on Towards Data Science .

Towards Data Science
Sep 8

The Model Validation Playbook for GenAI: Lessons from Banking

The article discusses how model validation standards are evolving for large language model (LLM) based systems, particularly in the banking sector. It examines what aspects of traditional validation break down, what elements remain applicable, and how to effectively test the quality of LLM outputs. The piece offers practical insights into adapting validation playbooks for generative AI applications.

By Ananya Bhattacharyya
arXiv Machine Learning
1d ago

The Hidden Costs of 99% Accuracy: A Trustworthiness Audit of the Telco Customer Churn Benchmark

The paper audits the IBM Telco Customer Churn benchmark, revealing that common practices inflate performance metrics. It shows that pre‑split SMOTE boosts churn‑class F1 by 13.1 points, that isotonic regression is the best calibration method while temperature scaling fails on tree ensembles, and that the cost‑optimal decision threshold is 5–10 times lower than the F1‑optimal one, saving about $77,000 per 1,000 customers. The authors also test generalisation on Iranian Telecom and Bank churn datasets, and propose a four‑component reporting checklist with reproducible code.

By Soumyadeep Roy
Towards Data Science
Aug 27

I Trained Six Models for Fraud Detection, and the Best One Isn't in Production

The article recounts a final‑year project in which the author trained six different models for fraud detection. It highlights the discrepancy between the model that performed best on evaluation metrics and the one that was ultimately chosen for production. The piece reflects on how real‑world constraints can override purely statistical performance.

By Benjamin Nweke
arXiv AI
Aug 28

Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration

The paper presents a framework for integrating explainable AI into customer churn prediction for telecommunications. It benchmarks four classifiers—Logistic Regression, Random Forest, XGBoost, and LightGBM—on the IBM Telco Customer Churn dataset, finding comparable performance with Logistic Regression achieving the highest AUC-ROC and LightGBM the highest accuracy. Explanations are provided via SHAP and LIME at both global and instance levels, and a four‑layer CRM integration architecture is proposed to translate risk scores and attribution vectors into actionable retention strategies, projecting a 3.3–5.3 percentage point reduction in churn and $199K–$319K savings per campaign cycle.

By Sandeep Gaddamwar