arXiv Machine Learning

Accurate Ensembles, Fragile Narratives: Multi-Scale Stacking and a Fidelity Audit of LLM-Generated Explanations for Credit Risk

arXiv:2608. 08126v1 Announce Type: new Abstract: Credit scoring increasingly relies on models whose decision logic cannot be read off their parameters, in tension with supervisory expectations that adverse decisions be explainable.

arXiv AI
Aug 21

Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay

arXiv:2608. 19760v1 Announce Type: cross Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance.

By Haiyue Zhang
arXiv Machine Learning
Sep 25

MemGuard-Alpha: Limits of Membership Inference for Detecting and Filtering Memorization-Contaminated Signals in LLM-Based Financial Forecasting

MemGuard-Alpha evaluates whether membership inference attacks (MIA) can detect memorization in large language models (LLMs) used for financial alpha signals. The study combines five MIA methods with a temporal proximity feature and a cross-model disagreement metric, then audits them across seven LLMs, 50 S&P 100 stocks, and 299,600 prompt-model pairs. Findings show that temporal proximity alone perfectly predicts in-sample status, MIA discriminative power largely stems from model scale differences, and filtering based on contamination scores does not improve risk-adjusted performance once transaction costs are considered.

By Anisha Roy, Dip Roy
arXiv AI
Aug 19

Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

The study investigates whether Large Language Models (LLMs) can translate technical explanations from credit risk models into stakeholder-friendly narratives. Using Freddie Mac loan data, the authors compare standard tabular models (XGBoost + SHAP) with alternative data pipelines (GNN + GNNExplainer and a bimodal mix) and generate explanations with three LLM configurations: a small fine‑tuned Gemma 3 4B, a large fine‑tuned DeepSeek R1 70B, and a zero‑shot Gemini 2.5. Findings show that the quality of explanations is more dependent on the evidence representation than on the LLM, that narratives reliably identify influential factors but are less consistent about the direction of influence, and that credit professionals demand higher evidentiary standards than non‑professionals.

By Sahab Zandi, Noah Kostesku, Christophe Mues, Mar\'ia \'Oskarsd\'ottir, Cristi\'an Bravo
arXiv AI
Sep 2

Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure

The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.

By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv AI
Sep 25

Baszta: Data-Centric Fine-Tuning of a Polish Multi-Label Safety Classifier

The paper presents Baszta, a Polish multi‑label content‑safety classifier trained by fine‑tuning the 124M‑parameter allegro/herbert‑base‑cased model on five categories (hate, vulgarity, sexual content, crime, self‑harm) using a Focal + R‑Drop objective. In out‑of‑distribution evaluation on the Gadzi Język benchmark, Baszta achieves a small but statistically significant improvement in micro‑F1 over the Bielik Guard system, though the macro‑F1 advantage disappears when both models are properly tuned. The study also explores calibration techniques, showing that per‑category temperature scaling can recover performance lost by Platt scaling or isotonic regression, and discusses the trade‑offs between robust calibration and adversarial recall.

By Adam G\'orski, Mateusz J\k{a}kalak, Rafa{\l} Jakubowski
arXiv Computation and Language
Aug 31

Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers

The paper investigates whether stacking multiple defenses around large language models (LLMs) truly compounds security. Using the Adversary Access‑Tier Model (AATM) and a cost‑tiering system, the authors analyze a seven‑layer defense stack and find that failure correlations between layers are consistently positive, meaning the residual attack success is higher than the multiplicative prediction. Despite high coverage and low false refusals, the stack’s performance is largely driven by common architectural causes rather than diverse, independent defenses.

By Abrar Alotaibi, Muhammad Shahid Jabbar, Sadam Al-Azani, Moataz Ahmed
arXiv AI
Sep 7

When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models

A 0.6B language model consistently answers YES to 1,200 logical tests, yet its behavior shows no discrimination. Linear probes reveal the correct verdict with high AUC (0.96) and transfer to unseen structures, but a single scalar readout fails due to a saturated decision threshold offset by +4.6 σ. Adjusting this threshold restores behavior accuracy from 50 % to 81 % and improves higher‑scale models, demonstrating that miscalibrated readouts, not hidden knowledge loss, drive performance gaps.

By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli