arXiv AI

A Decision-Support Audit Protocol for Supervision Drift in Proxy-Labeled Credit-Risk Prediction

The paper presents a locked, multi‑signal audit protocol designed to detect supervision drift in credit‑risk models that use proxy labels. It comprises five layers—transfer performance, an oracle‑gap probe, a calibration diagnostic, feature‑label stability, and a synthetic positive control—each with predefined thresholds and decision rules. Applied to a public LendingClub dataset, the protocol shows stable ranking, small oracle gaps, and identifies a prevalence and probability‑scale mismatch that recalibration largely mitigates, though its root cause remains unclear.

arXiv AI
Sep 2

Causal Evidentiary Governance for High-Risk Machine Learning Systems

The paper proposes Causal Evidentiary Governance (CEG), a framework that requires regulated institutions to maintain a versioned directed acyclic graph (DAG) separating allowable from disallowed causal pathways in high‑risk machine learning systems. CEG introduces the Causal Harm Rate to quantify prediction variation due to disallowed pathways and pairs each decision with a signed Decision‑Evidence Packet (DEP) that cryptographically links the prediction to the DAG and path‑specific attributions, enabling efficient inclusion proofs via a Merkle tree. Empirical validation on synthetic credit data and the German Credit dataset demonstrates that CEG more clearly isolates causal effects than traditional fairness metrics and that a proof‑of‑concept implementation shows operational feasibility with manageable performance tradeoffs.

By Samah Kareem, Bar{\i}\c{s} \c{C}elikta\c{s}
arXiv AI
Jun 3

Auditable Climate Risk Intelligence from Fragmented ESG Data: Deterministic Orchestration and Imbalance-Aware Learning for Scope 1-3 Validation

arXiv:2606. 02604v1 Announce Type: cross Abstract: ESG and climate risk data remain fragmented across heterogeneous Scope 1, Scope 2, and Scope 3 reporting environments, while conventional validation pipelines lack provenance aware auditability, hidden drift detection, and reproducibility oriented governance.

By Karan Sehgal, Khawar Naveed Bhatti
arXiv Machine Learning
Aug 3

Incorporating data drift to perform survival analysis on credit risk

arXiv:2601. 20533v2 Announce Type: replace-cross Abstract: Survival analysis has become a standard approach for modelling time to default by time-varying covariates in credit risk.

By Jianwei Peng (Humboldt-Universit\"at zu Berlin), Stefan Lessmann (Humboldt-Universit\"at zu Berlin, Bucharest University of Economic Studies)
arXiv AI
Sep 25

When Does Action Credit Need Updating?

The paper investigates when historical action credit for tool‑using agents must be updated after policy changes. It introduces pairwise branch sensitivity to measure how policy updates affect action‑distinguishing branches, and proposes a first‑order anchored credit‑transport estimator along with a Decision‑Sufficient Credit Gate (DSC‑Gate) to decide whether to reuse, transport, or resample credit. Experiments show that DSC‑Gate reduces new tool steps by 39.4% with negligible impact on regret, demonstrating that many policy updates can avoid costly recomputation of action credit.

By Hongye Yang, Boxiao Huang
arXiv AI
Aug 19

Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

The study investigates whether Large Language Models (LLMs) can translate technical explanations from credit risk models into stakeholder-friendly narratives. Using Freddie Mac loan data, the authors compare standard tabular models (XGBoost + SHAP) with alternative data pipelines (GNN + GNNExplainer and a bimodal mix) and generate explanations with three LLM configurations: a small fine‑tuned Gemma 3 4B, a large fine‑tuned DeepSeek R1 70B, and a zero‑shot Gemini 2.5. Findings show that the quality of explanations is more dependent on the evidence representation than on the LLM, that narratives reliably identify influential factors but are less consistent about the direction of influence, and that credit professionals demand higher evidentiary standards than non‑professionals.

By Sahab Zandi, Noah Kostesku, Christophe Mues, Mar\'ia \'Oskarsd\'ottir, Cristi\'an Bravo
arXiv Machine Learning
Sep 24

SR-Fraud: An Outcome-Supervised Reflective LLM Agent Framework for Non-Stationary Payment Fraud Detection

SR‑Fraud is a framework that uses a frozen, stateless LLM agent to score transactions in real time while an offline reflection agent proposes boundary hypotheses based on matured errors. The system then verifies these hypotheses deterministically before updating its knowledge state. On a production payment‑fraud benchmark, SR‑Fraud outperforms both static and periodically retrained CatBoost models and successfully detects an emerging fraud burst.

By Xuwei Tan, Yao Ma, Xueru Zhang