arXiv AI

When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit

arXiv:2608. 04896v1 Announce Type: new Abstract: Defensive driving scores are useful only when they preserve distinctions between policies that observe surrounding actors and those that do not.

arXiv AI
Sep 18

Refuse, Decompose, Refresh: A Claim-Safe Protocol for Closed-Loop AI Evaluation

The paper introduces a claim‑safe protocol for evaluating closed‑loop AI systems, consisting of three actions: Refuse, Decompose, and Refresh. It demonstrates the protocol in a simulator with 24 policy components and 1,440 held‑out cases, showing that abstention and stable false admission rates are low while providing detailed statistical diagnostics. The approach emphasizes that evaluation results should be tied to observable support and statistical calibration rather than a single PASS/FAIL label.

By Peiying Zhu, Sidi Chang
arXiv AI
Sep 4

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

The paper reports a preregistered audit of language‑model judges used as measurement instruments, revealing that the assumption that a model’s responses remain stable over time is invalid. Across nearly 53,000 audited requests, repeat rankings and byte‑identical replays fell far below required reliability thresholds, with three identified mechanisms—label‑to‑meaning bias, candidate gaps below the noise floor, and input permutation noise—explaining the discrepancy. The study proposes a three‑level snapshot‑identity framework, eight design rules, and a reporting checklist to prevent such reliability failures in future evaluations.

By Haoyaun Zhu, Jie Zhang
Hugging Face Trending Papers
Sep 3

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

The paper investigates the reliability of language‑model judges used as measurement instruments on shared endpoints. Through two preregistered audits of 52,988 requests, the authors found that repeat rankings and byte‑identical replays fell far short of required thresholds, revealing significant instability. They identify three mechanisms—label‑to‑meaning bias, candidate gaps below the noise floor, and input permutation noise—that explain the gap, and propose a snapshot‑identity ladder, design rules, and a reporting checklist to mitigate such failures.

arXiv Machine Learning
1d ago

The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

The paper introduces the Alignment Flywheel, a governance‑centric hybrid multi‑agent system (MAS) that separates decision generation from safety governance. It defines a Proposer that generates candidate trajectories, a Safety Oracle stack that evaluates safety, and an Enforcement layer that applies risk policies at runtime. A governance MAS oversees monitoring, red‑teaming, verification, and versioned release management, enabling patch‑local fixes to safety failures without retraining the Proposer. The architecture is implementation‑agnostic and is demonstrated in two scenarios: a learned spatial Oracle and a clinical GenAI proxy. The authors provide open‑source code at https://github.com/decide-ugent/Alignment-Flywheel.

By Elias Malomgr\'e, Pieter Simoens
arXiv AI
3d ago

NAQD Env: A benchmark for selective withdrawal in language agents

The paper introduces NAQD‑Env, a synthetic benchmark designed to test language agents’ ability to selectively withdraw and resume actions when new evidence, permissions, or stop instructions arise. It evaluates models against a deterministic reference policy across eleven dependency families, measuring policy agreement, task value, withdrawal, resumption, and event reporting. Experiments on 350 scenarios show low withdrawal recall, no valid resumption, and limited policy alignment, highlighting the need to treat selective withdrawal as a distinct reliability component.

By Mohamed Abouzahra
arXiv AI
Aug 21

Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay

arXiv:2608. 19760v1 Announce Type: cross Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance.

By Haiyue Zhang
arXiv Machine Learning
Aug 19

Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment

The paper argues that cross‑view correspondence, commonly used in agent evaluation and trace‑based learning, functions as a measurement intervention. Removing or altering this correspondence can create artificial sensitivity or invariance, and multiple optimal correspondences can obscure mechanism labels and learning credit. The authors propose a validity theory with two‑sided validation, all‑optima identification, and uncertainty propagation, and demonstrate through experiments that unvalidated correspondences can misattribute credit and erase harmful responses.

By Zhen Zhang, Ahmad Hafez, Amr Alanwar