The paper presents a framework for integrating explainable AI into customer churn prediction for telecommunications. It benchmarks four classifiers—Logistic Regression, Random Forest, XGBoost, and LightGBM—on the IBM Telco Customer Churn dataset, finding comparable performance with Logistic Regression achieving the highest AUC-ROC and LightGBM the highest accuracy. Explanations are provided via SHAP and LIME at both global and instance levels, and a four‑layer CRM integration architecture is proposed to translate risk scores and attribution vectors into actionable retention strategies, projecting a 3.3–5.3 percentage point reduction in churn and $199K–$319K savings per campaign cycle.
By Sandeep Gaddamwar
arXiv:2606. 06776v1 Announce Type: new Abstract: Customer churn prediction is a central task in customer analytics, particularly in non-contractual, pay-per-use service environments where disengagement is not explicitly observed and must be inferred from behavioral inactivity.
By Muhammad Jawad Mufti, Omar Hammad, Haitham Saleh, Muqaddas Gull
arXiv:2607. 10260v1 Announce Type: new Abstract: Customer churn is a major challenge for telecommunication companies, directly eroding revenue and long term customer relationships.
By Nada Ali, Lina Ahmed, Tahani Abdalla Attia Gasmalla
arXiv:2604. 19755v2 Announce Type: replace Abstract: Anti-money laundering (AML) transaction monitoring generates large volumes of alerts that must be rapidly triaged by investigators under strict audit and governance constraints.
By Dorothy Torres, Wei Cheng, Ke Hu
arXiv:2606. 08696v1 Announce Type: cross Abstract: Counterfactual recourse aims to provide actionable feature changes that would alter an unfavorable decision made by a predictive model.
By Yasuo Tabei
FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.
By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang
The paper proposes a shift from predictive modeling to descriptive provider behavior profiles for fraud, waste, and abuse (FWA) review. By decomposing billed revenue into provider scale and procedure composition, the authors construct lineage‑aware profiles that capture scale history, code lineage, and clinical family shares. In a large Medicare audit, these simple, interpretable descriptions outperform complex forecasts and improve recall for high‑cost rare events, while an optional semantic factorization adds context without inferring intent.
By Yubin Park, Evan Brociner
arXiv:2608.17795v2 Announce Type: replace
Abstract: Text-to-SQL systems are commonly evaluated using ground-truth SQL queries or reference execution results, but such supervision is unavailable at in...
By Neelesh Kumar Shukla, Debasmita Panda, Srutanik Bhaduri, Aditya Banerjee, Vasu Rangarajan, Viji Krishnamurthy
The paper reports that language‑model agents used for customer‑relationship management can be misled by optimistic assertions from sales representatives in CRM records, leading to incorrect deal approvals. In a benchmark of 100 lead‑qualification tasks, models incorrectly cleared 29 of 31 deals where the representative’s claims contradicted company policies, with misalignment rates ranging from 87% to 97% across seven models. The authors propose diagnostic methods—including bucket analysis, same‑information controls, and compute‑step controls—to distinguish persuasion from information gaps and to quantify the impact of incentive‑misaligned witnesses.
By Rahul Balakavi
arXiv:2604. 10015v3 Announce Type: replace Abstract: Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks.
By Yupeng Cao, Haohang Li, Weijin Liu, Wenbo Cao, Anke Xu, Lingfei Qian, Xueqing Peng, Minxue Tang, Zhiyuan Yao, Jimin Huang, K. P. Subbalakshmi, Zining Zhu, Jordan W. Suchow, Yangyang Yu
arXiv:2609.37223v1 Announce Type: new
Abstract: Credit-risk prediction is important in banking, but a prediction alone does not explain why an applicant is risky or how it should be combined with oth...
By Aakash Kumar Tiwari
The paper introduces the concept of LLM‑specific utility, defining it as the performance gain a target large language model (LLM) achieves when provided with a passage compared to answering without evidence. A benchmark of utilitarian passages is built for four LLMs (Qwen3‑8B/14B/32B and Llama 3.1‑8B) across three QA datasets, revealing that each model benefits most from its own tailored evidence and that evidence optimized for other models is consistently suboptimal. The authors also create SpecUBench, a benchmark for LLM‑specific utility judgment, and show that current utility‑aware retrieval methods largely capture model‑agnostic usefulness, struggling to estimate LLM‑specific utility.
"whyItMatters":"The study demonstrates that retrieval‑augmented generation must consider model‑specific evidence selection to truly improve LLM performance, highlighting a gap in existing utility‑aware methods."
By Hengran Zhang, Keping Bi, Jiafeng Guo, Jiaming Zhang, Shuaiqiang Wang, Dawei Yin, Xueqi Cheng