arXiv:2606. 05263v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves reasoning and tool use, yet long-horizon language agents still learn unsupported evidence chains, belief drift, and shortcut actions that satisfy terminal checks.
By Renwei Meng
VeriPhy is an auditable physical‑verification system that transforms a text prompt into typed physical obligations and a statically validated execution plan before any video frames are generated. During execution, it gates calls to frozen low‑level experts (segmentation, tracking, counting, depth, OCR, audio‑event detection, etc.) and records provenance‑carrying evidence for each action. The system maps these records to a three‑valued state—supported, contradicted, or unknown—providing traceable verdicts that can be used to refine generation models.
By Wenzhuo Xu, Yuchen Zhu, Chongjian Ge, Xuan Shen, Jing Shi, Jason Kuen, Yongxin Chen, Molei Tao, Christopher McComb, Noelia Grande Guti\'errez, Jiuxiang Gu
VeriPhy is an auditable physical‑verification system that evaluates generated video by compiling prompts into typed physical obligations and a statically validated execution plan before any frames are observed. During execution, it gates calls to frozen low‑level experts (e.g., segmentation, tracking, counting, depth, OCR, audio‑event detection) and returns provenance‑carrying evidence records, which are mapped to a three‑valued state (supported, contradicted, unknown) with full traceability. On a 1,500‑clip corpus of human‑annotated flaw records, VeriPhy accounts for 228 failures out of 304, outperforming a published question‑decomposition evaluator that accounts for 164, while also providing auditable evidence for each verdict.
arXiv:2609.27606v1 Announce Type: new
Abstract: We introduce State-Grounded Conditioning (SGC), a design principle for user-facing LLM agents that must condition on live user state (game state, sessi...
By Qi Liu, Xiaoyang Yuan, Yubin Ruan, Zhuomeng Zhang, Wenjin Wang, Di Wu, Mingye Xu, Xinyi Mou, Xingxi Yin, Ke Feng, Zixun Sun
arXiv:2607.01469v3 Announce Type: replace
Abstract: Agentic Video Question Answering (VideoQA) systems produce answers through adaptive reasoning and tool-use trajectories, yet standard practice eval...
By Rama AlHamidi, Aseel Mohamed, Rasul Khanbayov, Mohamed Rayan Barhdadi, Erchin Serpedin, Hasan Kurban
The paper introduces the twin‑prefix framework to evaluate how the size of the verification unit—i.e., how many actions a pre‑execution LLM monitor reviews in one call—affects its performance. By pairing each gold plan with a twin that differs by a single write and injecting a controlled error, the authors isolate the impact of review length on catch rates and false rejections. Their findings show that longer review windows increase rejection rates but do not improve discrimination, with the highest informedness occurring at one or two actions across all judges and domains.
By Yuchen Han, Cheng Yan, Wuyang Zhang
arXiv:2609.26035v1 Announce Type: new
Abstract: Conversational agents often express answers in a uniformly confident register. We test whether expressed uncertainty, provenance-aware assertion, and e...
By Sebastian Cochinescu
arXiv:2606. 01120v1 Announce Type: new Abstract: In RAG-based fact-checking, LLMs are increasingly used as verifiers to check given claims against retrieved evidence.
By Yuxi Sun, Wenbo Shang, Wei Gao, Xin Huang, Jing Ma
arXiv:2608.31016v1 Announce Type: cross
Abstract: Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the n...
By Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
The paper introduces NAQD‑Env, a synthetic benchmark designed to test language agents’ ability to selectively withdraw and resume actions when new evidence, permissions, or stop instructions arise. It evaluates models against a deterministic reference policy across eleven dependency families, measuring policy agreement, task value, withdrawal, resumption, and event reporting. Experiments on 350 scenarios show low withdrawal recall, no valid resumption, and limited policy alignment, highlighting the need to treat selective withdrawal as a distinct reliability component.
By Mohamed Abouzahra
The paper presents a layered framework for evaluating conversational AI by aligning offline proxy signals with online A/B experiment outcomes. It introduces a three‑step alignment chain—behavioral label to product outcome, classifier to candidate behavior, and offline signal to experiment effect—alongside an audit protocol that compares confidence intervals and rankings. In a real‑world deployment, the composite proxy achieved 81.1% F1 versus 34.3% for the raw classifier, correctly predicting direction on all 113 contrasts and enabling efficient prioritization of candidate models before costly online testing.
By Xuanyi Li, Vaskar Nath, Hossein Amirkhani, Jay Li, Alex Deng
The paper introduces TRUST, a token‑level reward model that predicts the correctness of partial chain‑of‑thought (CoT) traces in vision‑language‑action (VLA) policies, enabling monitoring and selective steering of reasoning. On driving and manipulation VLA benchmarks, TRUST improves reasoning accuracy and reduces collision rates and trajectory errors, though its impact on overall task performance varies across tasks. The study defines two evaluation axes—correctability and actionability—to assess when CoT can serve as a runtime safety interface.
By Sathwik Karnik, Joseph JR. Lee, Aryaman Gupta, Somil Bansal