arXiv:2608.30086v1 Announce Type: cross
Abstract: Credit-default prediction is an important task in financial decision making. Traditional methods use fitted classifiers such as logistic regression a...
By Rishi Datta, Lavanya Prahallad
arXiv:2608. 19760v1 Announce Type: cross Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance.
By Haiyue Zhang
arXiv:2608. 14509v1 Announce Type: new Abstract: Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt.
By Zhelun Wu
MemGuard-Alpha evaluates whether membership inference attacks (MIA) can detect memorization in large language models (LLMs) used for financial alpha signals. The study combines five MIA methods with a temporal proximity feature and a cross-model disagreement metric, then audits them across seven LLMs, 50 S&P 100 stocks, and 299,600 prompt-model pairs. Findings show that temporal proximity alone perfectly predicts in-sample status, MIA discriminative power largely stems from model scale differences, and filtering based on contamination scores does not improve risk-adjusted performance once transaction costs are considered.
By Anisha Roy, Dip Roy
The study investigates whether Large Language Models (LLMs) can translate technical explanations from credit risk models into stakeholder-friendly narratives. Using Freddie Mac loan data, the authors compare standard tabular models (XGBoost + SHAP) with alternative data pipelines (GNN + GNNExplainer and a bimodal mix) and generate explanations with three LLM configurations: a small fine‑tuned Gemma 3 4B, a large fine‑tuned DeepSeek R1 70B, and a zero‑shot Gemini 2.5. Findings show that the quality of explanations is more dependent on the evidence representation than on the LLM, that narratives reliably identify influential factors but are less consistent about the direction of influence, and that credit professionals demand higher evidentiary standards than non‑professionals.
By Sahab Zandi, Noah Kostesku, Christophe Mues, Mar\'ia \'Oskarsd\'ottir, Cristi\'an Bravo
arXiv:2608. 16147v1 Announce Type: new Abstract: Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were properties of the method.
By Diyorbek Musaev
The paper introduces Counterfactual Fragility Certificates (CFC), a model‑agnostic audit protocol that maps each prediction to an evidence‑failure trajectory, summarizing it with metrics such as greedy flip budget, margin‑collapse area, degradation thresholds, and fragility dominance score. CFC is shown to identify brittle high‑confidence predictions on seven tabular benchmarks with an AUROC of 0.915, outperforming existing scalar scores by up to +0.405. The method remains effective across various perturbation and review‑budget scenarios, and can also inform fragility‑aware regularization and temperature correction.
By Filippo Cenacchi, Longbing Cao, Runze Yang
The paper presents Baszta, a Polish multi‑label content‑safety classifier trained by fine‑tuning the 124M‑parameter allegro/herbert‑base‑cased model on five categories (hate, vulgarity, sexual content, crime, self‑harm) using a Focal + R‑Drop objective. In out‑of‑distribution evaluation on the Gadzi Język benchmark, Baszta achieves a small but statistically significant improvement in micro‑F1 over the Bielik Guard system, though the macro‑F1 advantage disappears when both models are properly tuned. The study also explores calibration techniques, showing that per‑category temperature scaling can recover performance lost by Platt scaling or isotonic regression, and discusses the trade‑offs between robust calibration and adversarial recall.
By Adam G\'orski, Mateusz J\k{a}kalak, Rafa{\l} Jakubowski
The paper investigates whether stacking multiple defenses around large language models (LLMs) truly compounds security. Using the Adversary Access‑Tier Model (AATM) and a cost‑tiering system, the authors analyze a seven‑layer defense stack and find that failure correlations between layers are consistently positive, meaning the residual attack success is higher than the multiplicative prediction. Despite high coverage and low false refusals, the stack’s performance is largely driven by common architectural causes rather than diverse, independent defenses.
By Abrar Alotaibi, Muhammad Shahid Jabbar, Sadam Al-Azani, Moataz Ahmed
A 0.6B language model consistently answers YES to 1,200 logical tests, yet its behavior shows no discrimination. Linear probes reveal the correct verdict with high AUC (0.96) and transfer to unseen structures, but a single scalar readout fails due to a saturated decision threshold offset by +4.6 σ. Adjusting this threshold restores behavior accuracy from 50 % to 81 % and improves higher‑scale models, demonstrating that miscalibrated readouts, not hidden knowledge loss, drive performance gaps.
By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
arXiv:2609.39229v1 Announce Type: cross
Abstract: Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary fr...
By Elia Onofri, Roberto Di Pietro
arXiv:2606. 15887v1 Announce Type: cross Abstract: Large language model (LLM) systems are increasingly proposed to assist peer review, yet most evaluations judge the prose of machine-generated review text, not the validity of the numeric score a system assigns.
By Costa Georgantas