Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces a 96‑item paired benchmark for evaluating large language models (LLMs) on backtest auditing, where each flawed backtest is matched with a clean control that keeps strategy, dates, code style, labels, and reporting scaffold constant while altering a single methodological detail. A deterministic scorer distinguishes flaw recall, clean‑control false positives, evidence localization, and fix relevance. Experiments on 1440 cached audits from four text endpoints show that the DeepSeek auditor achieves perfect closed and clean‑aware code recall, yet open prompts over‑flag 93.8% of clean controls, and clean‑aware specificity is 87.5% even when recall saturates. Introducing a clean‑aware warning eliminates 20.8% false positives to 0% without affecting recall, though the budget anchor still flags many clean controls. Reporting clean‑control rates provides a clearer differentiation among models than reporting recall alone.
arXiv:2608. 02985v1 Announce Type: new Abstract: The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff.
TWIST is a new benchmark suite designed to evaluate the quality of interventions in conversational memory systems, focusing on how well these systems act correctly at belief change points. It includes four tracks—tension detection, draft vetting, belief-consistent answering, and sensitive recall governance—each paired with hard-negative controls to prevent gaming. The benchmark is rigorously validated through double annotation, adjudication, and a separability audit, revealing that current models face trade-offs between contradiction recall and false-positive rates.
arXiv:2606. 16062v1 Announce Type: new Abstract: We measure the rate at which code RL environments accept incorrect solutions as correct.
arXiv:2608.21606v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of targeted training data from a model while preserving its remaining capabilities, but evaluating whet...
The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.