arXiv Machine Learning

Incorporating data drift to perform survival analysis on credit risk

arXiv:2601. 20533v2 Announce Type: replace-cross Abstract: Survival analysis has become a standard approach for modelling time to default by time-varying covariates in credit risk.

arXiv Statistics ML
Aug 28

DTD-VAE: Disentangled Temporal Dependencies VAE for Credit Risk Prediction

The paper introduces DTD‑VAE, a Variational Autoencoder that disentangles temporal dependencies to better predict credit risk. It uses an autoregressive feature inference module to capture temporal patterns among latent variables and an element‑wise gating mechanism in the generative module to assign independent weights to each latent dimension, especially those relevant to credit risk. Experiments on six real‑world datasets show the model outperforms existing methods, improving ROC‑AUC by 3.2%–4.86% and Accuracy Ratio by 6.41%–9.71%.

By Xiaobo Guo, Lu-an Dong, Yanbo Wang, Peng Zhang, Cai Zhi, Youru Li
arXiv AI
Sep 16

A Decision-Support Audit Protocol for Supervision Drift in Proxy-Labeled Credit-Risk Prediction

The paper presents a locked, multi‑signal audit protocol designed to detect supervision drift in credit‑risk models that use proxy labels. It comprises five layers—transfer performance, an oracle‑gap probe, a calibration diagnostic, feature‑label stability, and a synthetic positive control—each with predefined thresholds and decision rules. Applied to a public LendingClub dataset, the protocol shows stable ranking, small oracle gaps, and identifies a prevalence and probability‑scale mismatch that recalibration largely mitigates, though its root cause remains unclear.

By Mehrdad Shoeibi, Muhammad Shabanpour, Waldemar Karwowski, Niloofar Yousefi
arXiv Statistics ML
Aug 25

Random Hazard Forests

Random Hazard Forests (RHF) is a survival tree ensemble that models how a patient's hazard changes over continuous time as new measurements arrive. RHF directly estimates a nonparametric hazard likelihood for predictable covariate processes, using an efficient working model to guide tree construction and then estimating flexible time‑varying hazards at each terminal node. By routing each tree based on the covariate state immediately before each time point, RHF can handle irregular and asynchronous covariate updates, and averaging across trees yields a pathwise hazard estimate that accurately captures changing risk in simulations and an intensive‑care application.

By Hemant Ishwaran, Eileen M. Hsich, Udaya B. Kogalur, Donald K. K. Lee
arXiv Machine Learning
Sep 14

FINESSE: An Agent-Based Simulator and Benchmark Dataset for Multimodal Financial Event Sequences

FINESSE is an agent‑based simulation framework that generates synthetic, structured datasets of multiple interdependent financial event streams, such as transactions, payments, account status changes, and policy interventions. Each stream has its own action space, schema, and variable types, and the streams are coupled through agents’ evolving latent states, allowing temporally rich interactions. The accompanying FINESSE‑Bench dataset supports four tasks—balance forecasting, transaction fraud detection, missed payment prediction, and next event prediction—and baseline results are provided using various time‑series and event‑sequence methods.

By Tyler Farnan, Benjamin Eng, Adam Abate, Xirui Hou, Rizal Fathony, Nam H. Nguyen, Senthil Kumar
arXiv Statistics ML
Sep 16

Statistical Inference for Score Decompositions

arXiv:2603.04275v2 Announce Type: replace-cross Abstract: We introduce inference methods for score decompositions, which partition scoring functions for predictive assessment into three interpretable...

By Timo Dimitriadis, Marius Puke