arXiv AI By Makar Ulesov, Vladislav Smirnov, Omar Ibrahim, Arsenii Bobovnikov

Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark

Read the original on arXiv AI →

The paper introduces a 96‑item paired benchmark for evaluating large language models (LLMs) on backtest auditing, where each flawed backtest is matched with a clean control that keeps strategy, dates, code style, labels, and reporting scaffold constant while altering a single methodological detail. A deterministic scorer distinguishes flaw recall, clean‑control false positives, evidence localization, and fix relevance. Experiments on 1440 cached audits from four text endpoints show that the DeepSeek auditor achieves perfect closed and clean‑aware code recall, yet open prompts over‑flag 93.8% of clean controls, and clean‑aware specificity is 87.5% even when recall saturates. Introducing a clean‑aware warning eliminates 20.8% false positives to 0% without affecting recall, though the budget anchor still flags many clean controls. Reporting clean‑control rates provides a clearer differentiation among models than reporting recall alone.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.

By Guangzhe Zhang
arXiv AI
Sep 25

TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment

TWIST is a new benchmark suite designed to evaluate the quality of interventions in conversational memory systems, focusing on how well these systems act correctly at belief change points. It includes four tracks—tension detection, draft vetting, belief-consistent answering, and sensitive recall governance—each paired with hard-negative controls to prevent gaming. The benchmark is rigorously validated through double annotation, adjudication, and a separability audit, revealing that current models face trade-offs between contradiction recall and false-positive rates.

By Subrat Panda
arXiv AI
Aug 26

More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight

The paper introduces the twin‑prefix framework to evaluate how the size of the verification unit—i.e., how many actions a pre‑execution LLM monitor reviews in one call—affects its performance. By pairing each gold plan with a twin that differs by a single write and injecting a controlled error, the authors isolate the impact of review length on catch rates and false rejections. Their findings show that longer review windows increase rejection rates but do not improve discrimination, with the highest informedness occurring at one or two actions across all judges and domains.

By Yuchen Han, Cheng Yan, Wuyang Zhang