arXiv:2606. 18479v1 Announce Type: new Abstract: Reject inference methods are widely used to mitigate survival bias in credit scoring, yet their effectiveness remains poorly understood.
By Bruno Scarone, Ricardo Baeza-Yates
The paper introduces a causal mediation framework to separate direct discrimination from structural inequality in AI-driven credit decisions. Using Pearl’s natural direct and indirect effects, it presents an identification strategy under treatment‑induced confounding and proposes a doubly‑robust estimator with efficiency guarantees. Empirical analysis of 89,465 mortgage applications shows that about 77% of racial denial disparities stem from financial mediators, while the remaining 23% represents a conservative lower bound on direct discrimination.
By Duraimurugan Rajamanickam
arXiv:2607. 19526v1 Announce Type: new Abstract: "Stop Chasing the C-index when Evaluating Survival Analysis Models" (ICML 2026, Spotlight) argued normatively, on synthetic data, that evaluating survival models by discrimination alone, i.
By Rafael da Silva, Danilo Alvares
arXiv:2609.27654v1 Announce Type: cross
Abstract: Verified income is often unavailable in digital loan applications, forcing lenders to rely on reported income and potentially leading to over-lending...
By Sultan Amed, Tanmay Sen, Sayantan Banerjee
arXiv:2504. 21259v2 Announce Type: replace-cross Abstract: Accurate imputation of race and ethnicity (R&E) is essential for fair lending compliance under ECOA, HMDA, and the Community Reinvestment Act, where up to 15% of mortgage applications carry missing race data and regulated institutions bear responsibility for identifying disparities on those records.
By S. Chalavadi, A. Pastor, T. Leitch
The paper presents a locked, multi‑signal audit protocol designed to detect supervision drift in credit‑risk models that use proxy labels. It comprises five layers—transfer performance, an oracle‑gap probe, a calibration diagnostic, feature‑label stability, and a synthetic positive control—each with predefined thresholds and decision rules. Applied to a public LendingClub dataset, the protocol shows stable ranking, small oracle gaps, and identifies a prevalence and probability‑scale mismatch that recalibration largely mitigates, though its root cause remains unclear.
By Mehrdad Shoeibi, Muhammad Shabanpour, Waldemar Karwowski, Niloofar Yousefi
Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditable fairness constraints. Logistic regression remains attractive because its coefficients ar...
arXiv:2608.24582v1 Announce Type: cross
Abstract: Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditable fairness constraints. Logistic regres...
By Victor Medina-Olivares, Stefan Lessmann, Jonathan Crook
The study investigates whether Large Language Models (LLMs) can translate technical explanations from credit risk models into stakeholder-friendly narratives. Using Freddie Mac loan data, the authors compare standard tabular models (XGBoost + SHAP) with alternative data pipelines (GNN + GNNExplainer and a bimodal mix) and generate explanations with three LLM configurations: a small fine‑tuned Gemma 3 4B, a large fine‑tuned DeepSeek R1 70B, and a zero‑shot Gemini 2.5. Findings show that the quality of explanations is more dependent on the evidence representation than on the LLM, that narratives reliably identify influential factors but are less consistent about the direction of influence, and that credit professionals demand higher evidentiary standards than non‑professionals.
By Sahab Zandi, Noah Kostesku, Christophe Mues, Mar\'ia \'Oskarsd\'ottir, Cristi\'an Bravo
arXiv:2609.37223v1 Announce Type: new
Abstract: Credit-risk prediction is important in banking, but a prediction alone does not explain why an applicant is risky or how it should be combined with oth...
By Aakash Kumar Tiwari
arXiv:2608. 19760v1 Announce Type: cross Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance.
By Haiyue Zhang
arXiv:2608. 02664v1 Announce Type: cross Abstract: Deploying ML models in regulated decision-making (credit underwriting, fraud detection, loan approval) requires demonstrating fairness and robustness to auditors without exposing model weights or customer data.
By Mohammad Nasir Uddin, Rahnuma Tabassum Orpita, Eklachur Rahman Bhuiyan, Asaduzzaman Anik