arXiv AI

When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

arXiv:2608. 07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power.

arXiv Computation and Language
Aug 31

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.

By Qing Ye, Meng-Hsuan Lin
arXiv AI
3d ago

Hard-Gate Candidacy in a Deployed Validator Suite

The paper evaluates hard‑gate candidacy for validators in a deployed generative‑agent system by measuring how well each validator’s firing separates successful from failed builds. Across 13 validators and thousands of builds, only a few checks show statistically significant separation, while many fail to distinguish or never fire. The study highlights that skipped checks are recorded as passes, limiting detectable failure rates and underscoring the need for clearer evaluation records.

By Xin Xu
arXiv AI
3d ago

RAGScope: A Leakage-Controlled, Cost-Aware Evidence-Gating Protocol for RAG Hallucination Triage

RAGScope is a leakage‑controlled, cost‑aware protocol that evaluates evidence‑gating mechanisms for retrieval‑augmented generation (RAG) systems using only the task input, retrieved context, and answer text. The enhanced gate, RAGScope‑E, achieves an AUROC of 0.798 and an average precision of 0.660 on three RAGTruth tasks, outperforming ROUGE‑L by 0.034 in pooled AP and delivering 0.748 precision within a top‑10% review budget. It operates quickly (6.22 ms per example on CPU) and demonstrates that cheap evidence gates can effectively triage RAG outputs, though calibration must be validated and adapted for each target domain.

By Zeming Liu, Qibai Chen, Jingtao Zhang, Hang Lyu
arXiv AI
2d ago

Auditing Routing Entropy as an Uncertainty Signal in Attention-Residual Transformers

The paper audits whether routing entropy in Attention‑Residual transformer variants (Swin‑Tiny and DeiT‑Small) trained on CIFAR‑10/100 can signal prediction uncertainty beyond model confidence. Three tests examine the presence, consistency, and predictive power of routing signals, while a sensitivity audit measures how much injected effect the probes recover. Results show no significant improvement over confidence alone, with only modest recovery of injected signals and no consistent gains across seeds or metrics.

By Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen