arXiv:2608. 08126v1 Announce Type: new Abstract: Credit scoring increasingly relies on models whose decision logic cannot be read off their parameters, in tension with supervisory expectations that adverse decisions be explainable.
By Gregorius Reynaldi Pratama, Kuo-Kun Tseng
arXiv:2505. 24622v3 Announce Type: replace Abstract: Many high-stakes screening tasks require predicting rare outcomes from unstructured text, where errors are costly and decisions must be auditable.
By Ben Griffin, Aaron Ontoyin Yin, Diego Vidaurre, Ugur Koyluoglu, Joseph Ternasky, Fuat Alican, Yigit Ihlamur
The paper investigates the "score granularity gap" in black-box large language model (LLM) classifiers, asking how finely a confidence score can be thresholded for deployment. By comparing seven confidence construction methods across 25 model-dataset pairs, the authors find that single-shot verbalized confidence, when properly converted to a probability, ranks well but offers only a few distinct threshold values, limiting operational flexibility. The study also shows that multi-query aggregation can improve weak models but may harm strong ones, and provides concrete guidance for deployment trade-offs.
By Ao Sun, Tian Sun, Jiaxing Geng
arXiv:2605. 18147v2 Announce Type: replace Abstract: Predictive models play a pivotal role in credit risk management, guiding critical decisions through accurate estimation of default probabilities and losses.
By Bart Baesens, Andreas Goethals, Stefan Lessmann, Simon De Vos, Cristi\'an Bravo, David Martens, Victor Medina-Olivares, Christophe Mues, Maria Oskarsd\'ottir, Seppe vanden Broucke, Tony Van Gestel, Tim Verdonck, Wouter Verbeke
The paper presents a framework for integrating explainable AI into customer churn prediction for telecommunications. It benchmarks four classifiers—Logistic Regression, Random Forest, XGBoost, and LightGBM—on the IBM Telco Customer Churn dataset, finding comparable performance with Logistic Regression achieving the highest AUC-ROC and LightGBM the highest accuracy. Explanations are provided via SHAP and LIME at both global and instance levels, and a four‑layer CRM integration architecture is proposed to translate risk scores and attribution vectors into actionable retention strategies, projecting a 3.3–5.3 percentage point reduction in churn and $199K–$319K savings per campaign cycle.
By Sandeep Gaddamwar
arXiv:2608.20343v1 Announce Type: new
Abstract: This study develops and evaluates a bankruptcy prediction framework that integrates consensus-based feature selection, hybrid resampling, stacking ense...
By Obu-Amoah Ampomah, Edmund Fosu Agyemang, Kofi Acheampong, Louis Agyekum, Enock Adu Bonsu, Eric Nyarko
arXiv:2606. 10347v1 Announce Type: new Abstract: Machine learning is increasingly used in critical domains, where both predictions and their associated confidence levels influence important decisions.
By Vin\'icius Peixoto Chagas, Carlos Henrique Leit\~ao Cavalcante, Thiago Alves Rocha
arXiv:2605.24564v2 Announce Type: replace
Abstract: Backtesting large language models (LLMs) on historical financial data is unreliable when their pre-training data include the evaluated events. An L...
By Weixian Waylon Li, Mengyu Wang, Tiejun Ma
arXiv:2606. 15314v1 Announce Type: cross Abstract: Industrial retrofit planning depends on structured operational data rather than free text: planners must estimate whether a newly registered prototype will require a retrofit, which retrofit package it will need, and how long the work will take.
By Aina Vila Pons, Ioannis Tzachristas, Constantinos Antoniou
arXiv:2606. 12117v1 Announce Type: cross Abstract: Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.
By Selen Erkan, Bastian Boll, Kristian Kersting, Bj\"orn Deiseroth, Letitia Parcalabescu
arXiv:2605. 20716v5 Announce Type: replace Abstract: Random forests construct each tree with a different, randomised representation of the feature space.
By Youngjoon Park
Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false negatives leave harmful activity undetected. The hardest cases require jointly understanding a merchant's textual profile and long behavioral sequence.