arXiv Machine Learning

Beyond Aggregate Calibration: Decomposing Income-Conditional Recall Disparities in Automated Credit Default Prediction

arXiv:2608. 08202v1 Announce Type: new Abstract: Data-centric curation pipelines frequently rely on model confidence scores to flag and filter noisy or mislabeled training instances.

arXiv Machine Learning
Sep 14

Decomposing Discrimination: Causal Mediation Analysis for AI-Driven Credit Decisions

The paper introduces a causal mediation framework to separate direct discrimination from structural inequality in AI-driven credit decisions. Using Pearl’s natural direct and indirect effects, it presents an identification strategy under treatment‑induced confounding and proposes a doubly‑robust estimator with efficiency guarantees. Empirical analysis of 89,465 mortgage applications shows that about 77% of racial denial disparities stem from financial mediators, while the remaining 23% represents a conservative lower bound on direct discrimination.

By Duraimurugan Rajamanickam
arXiv Machine Learning
Jul 21

STRATA: A Name-and-Geography Race Inference Model for Fair Lending and Housing Equity Applications

arXiv:2504. 21259v2 Announce Type: replace-cross Abstract: Accurate imputation of race and ethnicity (R&E) is essential for fair lending compliance under ECOA, HMDA, and the Community Reinvestment Act, where up to 15% of mortgage applications carry missing race data and regulated institutions bear responsibility for identifying disparities on those records.

By S. Chalavadi, A. Pastor, T. Leitch
arXiv AI
Sep 16

A Decision-Support Audit Protocol for Supervision Drift in Proxy-Labeled Credit-Risk Prediction

The paper presents a locked, multi‑signal audit protocol designed to detect supervision drift in credit‑risk models that use proxy labels. It comprises five layers—transfer performance, an oracle‑gap probe, a calibration diagnostic, feature‑label stability, and a synthetic positive control—each with predefined thresholds and decision rules. Applied to a public LendingClub dataset, the protocol shows stable ranking, small oracle gaps, and identifies a prevalence and probability‑scale mismatch that recalibration largely mitigates, though its root cause remains unclear.

By Mehrdad Shoeibi, Muhammad Shabanpour, Waldemar Karwowski, Niloofar Yousefi
arXiv AI
Aug 19

Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

The study investigates whether Large Language Models (LLMs) can translate technical explanations from credit risk models into stakeholder-friendly narratives. Using Freddie Mac loan data, the authors compare standard tabular models (XGBoost + SHAP) with alternative data pipelines (GNN + GNNExplainer and a bimodal mix) and generate explanations with three LLM configurations: a small fine‑tuned Gemma 3 4B, a large fine‑tuned DeepSeek R1 70B, and a zero‑shot Gemini 2.5. Findings show that the quality of explanations is more dependent on the evidence representation than on the LLM, that narratives reliably identify influential factors but are less consistent about the direction of influence, and that credit professionals demand higher evidentiary standards than non‑professionals.

By Sahab Zandi, Noah Kostesku, Christophe Mues, Mar\'ia \'Oskarsd\'ottir, Cristi\'an Bravo
arXiv AI
Aug 21

Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay

arXiv:2608. 19760v1 Announce Type: cross Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance.

By Haiyue Zhang
arXiv AI
Aug 5

ZK-SR117: A Chunked Zero-Knowledge Attestation Design for Aggregated Fair-Lending Metrics, with a Control Mapping toward Full SR 11-7 Coverage

arXiv:2608. 02664v1 Announce Type: cross Abstract: Deploying ML models in regulated decision-making (credit underwriting, fraud detection, loan approval) requires demonstrating fairness and robustness to auditors without exposing model weights or customer data.

By Mohammad Nasir Uddin, Rahnuma Tabassum Orpita, Eklachur Rahman Bhuiyan, Asaduzzaman Anik