Neural networks increasingly combine data across populations, time periods, and operating conditions to improve generalization. This raises a reliability question: whether a model refitted on pooled data preserves an action ordering supported by both sources.
arXiv:2607. 27255v1 Announce Type: cross Abstract: Neural networks increasingly combine data across populations, time periods, and operating conditions to improve generalization.
By Yanli Yan, Yuanzheng Li, Yong Zhao, Hongbo Guo, Shoudong Han
The paper investigates whether neural networks retain a case-based structure in their learned representations, enabling the decomposition of decision margins into contributions from individual training cases. By linking neural networks to Case-Based Decision Theory (CBDT), the authors identify conditions under which this recovered structure can be interpreted within CBDT and demonstrate that the resulting interpretation is consistent with the network’s original decisions. Experiments on a controlled CBDT setting and three real-world decision tasks confirm the viability of this approach.
By Manli Yan, Yaowen Yu, Yong Zhao, Yuebin Lin, Shoudong Han
arXiv:2608.24582v1 Announce Type: cross
Abstract: Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditable fairness constraints. Logistic regres...
By Victor Medina-Olivares, Stefan Lessmann, Jonathan Crook
Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditable fairness constraints. Logistic regression remains attractive because its coefficients ar...
The study investigates whether Large Language Models (LLMs) can translate technical explanations from credit risk models into stakeholder-friendly narratives. Using Freddie Mac loan data, the authors compare standard tabular models (XGBoost + SHAP) with alternative data pipelines (GNN + GNNExplainer and a bimodal mix) and generate explanations with three LLM configurations: a small fine‑tuned Gemma 3 4B, a large fine‑tuned DeepSeek R1 70B, and a zero‑shot Gemini 2.5. Findings show that the quality of explanations is more dependent on the evidence representation than on the LLM, that narratives reliably identify influential factors but are less consistent about the direction of influence, and that credit professionals demand higher evidentiary standards than non‑professionals.
By Sahab Zandi, Noah Kostesku, Christophe Mues, Mar\'ia \'Oskarsd\'ottir, Cristi\'an Bravo
arXiv:2605. 20098v2 Announce Type: replace Abstract: Claim verification is an important problem in high-stakes settings, including health and finance.
By Gabriel Freedman, Adam Dejl, Adam Gould, Mansi, Lihu Chen, Junqi Jiang, Francesca Toni
FinRCA-Bench is a deterministic synthetic benchmark comprising 2,250 accounts‑payable‑to‑bank reconciliation cases that span 14 operational tables and include 1,500 injected failures across 15 causal categories. The benchmark hides root‑cause labels and record‑level evidence contracts from models, enabling independent evaluation of evidence retrieval versus reasoning accuracy. Experiments show that retrieval architecture dramatically influences performance, with retrieval improvements raising macro‑required‑record recall from 0.83% to 77.70% and exact 16‑class accuracy from 2.05% to 72.44%.
By Pratik Ghawate
Legal outcome prediction must disentangle objective case facts from adjudicative context. Merit-based rulings rely on factual evidence while technical disposals may hinge on judicial discretion.
FinRCA-Bench is a synthetic benchmark designed to evaluate evidence retrieval and reasoning in financial AI systems, specifically for accounts‑payable‑to‑bank reconciliation. It contains 2,250 cases across 14 operational tables, with 1,500 injected failures in 15 causal categories and 750 hard‑negative cases, and hides root‑cause labels and evidence contracts to isolate retrieval performance. Experiments show that retrieval architecture dramatically affects accuracy, with structured retrieval methods like Typed Provenance Graph Retrieval vastly improving macro‑recall and exact‑class accuracy compared to dense semantic retrieval or classical ML.
arXiv:2605. 23955v3 Announce Type: replace Abstract: Deploying machine learning in regulated financial environments -- credit risk, fraud detection, and anti-money laundering -- exposes critical vulnerabilities in algorithmic reproducibility.
By Ruizhe Zhou, Xiaoyang Liu, Gaoyuan Du, Yi Zheng, Shouxi Ren, Deepayan Chakrabarti, Dengdu Jiang
arXiv:2605. 18147v2 Announce Type: replace Abstract: Predictive models play a pivotal role in credit risk management, guiding critical decisions through accurate estimation of default probabilities and losses.
By Bart Baesens, Andreas Goethals, Stefan Lessmann, Simon De Vos, Cristi\'an Bravo, David Martens, Victor Medina-Olivares, Christophe Mues, Maria Oskarsd\'ottir, Seppe vanden Broucke, Tony Van Gestel, Tim Verdonck, Wouter Verbeke