AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after using such a tool, what can the user still do on their own?
arXiv:2606. 31273v1 Announce Type: new Abstract: AI-assisted research has entered a stage in which the central question is not only whether systems can generate hypotheses, run experiments, or produce manuscripts, but whether their scientific claims are calibrated to the evidence that supports them.
By Hongmin Li
arXiv:2609.21841v1 Announce Type: new
Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid
arXiv:2609.15624v1 Announce Type: cross
Abstract: Researchers assessing competent generative-AI use at work must choose among self-reports, objective tests, and measures of oversight and reliance. We...
By Daniele Veri'
arXiv:2607. 26159v1 Announce Type: cross Abstract: An AI benchmark result rarely reaches a consequential claim in one step.
By Brett Reynolds
The paper argues that as AI systems increasingly generate code, the bottleneck has shifted to supervising these systems, revealing a vocabulary gap between cybernetic coordination (actions aligning with the world) and epistemic coordination (understanding that can be verified). It critiques current oversight that merely approves outputs, proposing instead that every consequential choice by an agent must include a retrievable condition explaining why it was made, enabling third‑party verification. The authors illustrate this with three delegation episodes, introduce a two‑part reconstruction test, and propose the ORRCF convention to embed such conditions in all recorded decisions.
By J\'er\'emie Lumbroso
The paper proposes a claim‑specific verification audit for modular agents that replaces aggregate task scores with evidence‑based evaluations. Each agent conclusion is recorded with supporting evidence and classified as supported, unsupported, unresolved, or not evaluated, along with the boundary of validity. The audit employs three tools—oracle policies, perfect component replacements, and verifier‑score tests—to trace value changes, locate lost value, and assess verifier effectiveness, demonstrated on a portfolio‑allocation agent in a synthetic market.
By Ali Atiah Alzahrani
The paper introduces the Discovery Certification Protocol (DCP), a framework that transforms claims from AI research agents into executable tests for recovery and feedback. DCP includes multiple gates that validate improvements, provide controlled information, and measure the impact of truthful feedback, while its core requires strict controls and finite‑sample bounds. Controlled audits in SQLite optimization and virtual catalyst control demonstrated zero recoveries across 96 episodes, with rigorous verification by a deterministic, LLM‑free verifier.
By Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng
arXiv:2605. 27914v2 Announce Type: replace-cross Abstract: Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional support, counseling.
By Yuming (Rapheal), Huang, Yao Liu, Pengjie Ding, Lei Wang, Junchen Wan
The paper introduces a diagnostic for reference‑free judge gates in text‑space skill optimization. It formalizes a judge as a latent solver, deriving a closed‑form bound on discriminability (ROC‑AUC) in terms of judge competence and answer‑space size, and shows that discriminability is confounded by item difficulty unless a within‑question estimator is used. A non‑intervening probe demonstrates that discriminability is at chance near the competence floor, rises above it, and that the diagnostic can predict gating errors in closed‑loop experiments.
By Chenle Chen, Yangbo Wei, Chao Yao, Shaoqiang Lu, Junhong Qian, Chen Wu, Lei He
arXiv:2608. 13706v1 Announce Type: cross Abstract: Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagreement, debate verifies an aggregate report rather than individual claims, and such verification occurs only after drafting, leaving inter-agent errors undetected until the final text.
By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain
arXiv:2607. 21268v1 Announce Type: cross Abstract: In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable correctness signal exists.
By Chen Zhu, Xiaolu Wang, Weilong Zhang