arXiv Machine Learning

How Early Is Early Enough? Design-Dependent Observation-Window Sufficiency in Subscription Churn Prediction

arXiv:2607. 00473v1 Announce Type: new Abstract: How many days of early behavior suffice for subscription churn prediction?

arXiv Machine Learning
Aug 20

Seasonal false alarms in customer churn and decline early-warning systems: adjacent-window labels confound seasonality with decline, and a year-over-year correction

The paper investigates how seasonal labeling in customer churn and decline early‑warning systems can misclassify seasonal activity as decline, leading to false alarms. By aligning the baseline to the same calendar months one year earlier, the authors demonstrate a significant improvement in model performance (ROC‑AUC rises from 0.767 to 0.864) and a reduction in the number of flagged accounts. The study quantifies the extent of the problem across public and production datasets and shows that the proposed correction reduces intervention load while maintaining predictive accuracy.

By Md Rezwanul Islam, Wael Mohammed
arXiv Machine Learning
Aug 13

Certifying What Helps Customer-Return Timing: A Screen-and-Confirm Test for Conditioning Signals, and Why Decay Is Nearly Enough

arXiv:2608. 11555v1 Announce Type: new Abstract: Practitioners enrich customer-return models with ever more signals (lifetime value, category, recency/frequency, calendar, geography), and the temporal-point-process (TPP) literature follows suit with covariate- and external-covariate-conditioned intensities.

By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
arXiv AI
Jun 2

ChurnNet: A Optimized Modern AI for Churn Prediction

arXiv:2606. 00169v1 Announce Type: cross Abstract: Increased competition and the growing similarity of products and services offered by retailers have lowered the barriers for customers to switch to competitors.

By Syed Saad Saif, Giulio Maggiore, Paolo Russo, Damiano Distante
arXiv Machine Learning
Jun 9

Benchmark Datasets for Lead-Lag Forecasting on Social Platforms

arXiv:2511. 03877v2 Announce Type: replace Abstract: Social and collaborative platforms emit multivariate time-series traces in which early interactions -- such as views, likes, or downloads -- are followed, sometimes months or years later, by higher impact like citations, sales, or reviews.

By Kimia Kazemian (Department of Computer Science, Cornell University), Zhenzhen Liu (Department of Computer Science, Cornell University), Yangfanyu Yang (Department of Information Science, Cornell University), Katie Luo (Department of Computer Science, Stanford University), Shuhan Gu (Department of Computer Science, Cornell University), Audrey Du (Department of Computer Science, Cornell University), Xinyu Yang (Department of Information Science, Cornell University), Jack Jansons (Department of Computer Science, Cornell University), Kilian Q. Weinberger (Department of Computer Science, Cornell University), John Thickstun (Department of Computer Science, Cornell University), Yian Yin (Department of Information Science, Cornell University), Sarah Dean (Department of Computer Science, Cornell University)
arXiv Machine Learning
Sep 22

A Patient World Model for Early Forecasting of Digital Health Campaign Outcomes: Capabilities and Limits

The paper introduces a compact patient world model that forecasts digital health campaign outcomes by maintaining a latent state per patient and learning exposure‑conditioned dynamics. Evaluated on a large US campaign dataset, the model predicts new‑to‑brand prescription volume with low relative error (2.9% at week‑4 cutoff) compared to much higher errors from baseline classifiers. The study also shows that dense next‑exposure supervision is crucial for accurate forecasts when conversions are rare and highlights limitations in interpreting exposure‑conditioned rollouts causally.

By Yunlong Wang
arXiv AI
Aug 28

Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration

The paper presents a framework for integrating explainable AI into customer churn prediction for telecommunications. It benchmarks four classifiers—Logistic Regression, Random Forest, XGBoost, and LightGBM—on the IBM Telco Customer Churn dataset, finding comparable performance with Logistic Regression achieving the highest AUC-ROC and LightGBM the highest accuracy. Explanations are provided via SHAP and LIME at both global and instance levels, and a four‑layer CRM integration architecture is proposed to translate risk scores and attribution vectors into actionable retention strategies, projecting a 3.3–5.3 percentage point reduction in churn and $199K–$319K savings per campaign cycle.

By Sandeep Gaddamwar
arXiv Machine Learning
Jun 2

Benchmarking Waitlist Mortality Prediction in Heart Transplantation Through Time-to-Event Modeling using New Longitudinal UNOS Dataset

arXiv:2507. 07339v2 Announce Type: replace-cross Abstract: Decisions about managing patients on the heart transplant waitlist are currently made by committees of doctors who consider multiple factors, but the process remains largely ad-hoc.

By Yingtao Luo, Reza Skandari, Carlos Martinez, Arman Kilic, Rema Padman