arXiv AI

Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations

arXiv:2606. 17005v1 Announce Type: new Abstract: Public AI evaluations are often read as terminal leaderboards, yet the underlying evidence is a selective time series shaped by reporting rules, benchmark revisions, and missingness.

arXiv Machine Learning
5d ago

Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model

The paper audits a commercial non‑generative System‑1 model on biosecurity‑relevant benchmarks, evaluating accuracy, calibration, error detection, selective prediction, and sensitivity to answer‑option order. It finds that the model’s accuracy varies strongly by task, is reasonably well calibrated when the vendor’s uncertainty field is interpreted correctly, and that option order can cause significant prediction changes—averaging across rotations improves accuracy. The study also shows that applying averaging only to low‑confidence items recovers most of the gain at a lower cost.

By Kimon Antonios Provatas, Ilias Georgakopoulos-Soares
arXiv AI
Sep 25

How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure

The paper investigates the reliability of ranking tables produced by small-sample evaluations of large language models (LLMs). Using LLM‑inferred prompt structure across eight model variants, the authors find that prompt‑structure recovery is highly unstable, with only the bottom of the ranking consistently reproducible. They demonstrate that standard evaluation practices can misrepresent model performance and propose reporting practices to improve transparency.

By Dipankar Sarkar
arXiv Machine Learning
Sep 25

Time-Series Foundation Models That Understand Data Revisions

The paper introduces VINTAGE-TS, a revision‑aware time‑series foundation model that separates observation time from information‑availability time. It predicts both the next period’s first‑published value and the value available after a fixed delay, maintaining a joint distribution to capture their dependence and uncertainty. The authors provide a detailed evaluation protocol, software tools for validity‑interval reconstruction and delayed‑label filtering, and a synthetic demonstration with a 25‑configuration sensitivity suite to illustrate performance variability and the impact of hindsight contamination.

By Taimoor Ahmad
arXiv AI
Sep 4

Instruction Duplication as an Inference-Time Control Primitive

The paper introduces instruction duplication, a simple inference‑time control that repeats the procedural instruction without retraining or decoding changes. Across seven instruction‑tuned models and 16,800 scheduled generations, duplicating the instruction improves deterministic All‑8 diagnostic‑response success from 90.22% to 93.17% and reduces failures by 30.2%. In downstream Answer Engineering scenarios, duplication further boosts success rates, demonstrating its practical impact on systems that rely on the generated trajectory.

By Victor Lavrenko (PeaceTech VC, Israel)
arXiv Machine Learning
Sep 21

Tracing the Evidence Behind Zero-Shot Time-Series Forecasting: A Source-First Taxonomy and Audit Framework

The paper argues that zero‑shot time‑series forecasting should be treated as an evidence‑access claim rather than merely a no‑parameter‑update condition. It introduces a source‑first taxonomy that distinguishes three evidence sources—frozen LLM prior reuse, parametric time‑series pretraining, and retrieval‑augmented external memory—from the architectures that implement them. The authors further outline four audit questions—task interface, forecast object and scoring, prediction‑time context, and resource budget—to make zero‑shot leaderboards transparent and comparable.

By Delun Kong, Wanyun Ling, Chenxi Liu, Ziyue Li
arXiv AI
Sep 25

When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability Routing

The paper investigates when forecasting agents should employ different behaviors—retrieval, reasoning, deferring to market priors, or using historical analogs—on binary forecasting tasks. It finds that the optimal mechanism depends on the data source, with structured analogs excelling for some processes and market or conservative baselines for others. The authors propose ReliabilityRoute, a rule‑based system that steers agent behavior using reliability features, achieving competitive performance across multiple LLM versions while highlighting that more reasoning is not always better.

By Yufeng Wang
Hugging Face Trending Papers
Jul 13

From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth

Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop. We study the complementary setting in which ground truth is delayed, censored, or private, so deterministic code cannot check correctness at scoring time and must instead issue a code-owned provisional forecast.