The paper proposes a claim‑specific verification audit for modular agents that replaces aggregate task scores with evidence‑based evaluations. Each agent conclusion is recorded with supporting evidence and classified as supported, unsupported, unresolved, or not evaluated, along with the boundary of validity. The audit employs three tools—oracle policies, perfect component replacements, and verifier‑score tests—to trace value changes, locate lost value, and assess verifier effectiveness, demonstrated on a portfolio‑allocation agent in a synthetic market.
By Ali Atiah Alzahrani
The paper surveys 160 benchmarks from 2017‑2026 that evaluate predictive embodied intelligence, categorising them into policy suites, embodied agents, world‑model evaluation, and prediction‑to‑action bridges. It finds that most benchmarks are model‑agnostic, rarely compare Vision‑Language‑Action policies to world models, and seldom turn predictions into executed actions. The authors argue that the lack of benchmarks designed to directly test the closed‑loop advantage of world models prevents the field from answering whether such models truly improve robotic performance.
By Gaytri Jena, Kapil Wanaskar, Vinija Jain, Aman Chadha, Vasu Sharma, Amitava Das
arXiv:2610.00601v1 Announce Type: cross
Abstract: Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface fo...
By Sathwik Karnik, Joseph JR. Lee, Aryaman Gupta, Somil Bansal
arXiv:2605. 28114v2 Announce Type: replace Abstract: Language-model agents are moving from single-user assistants into persistent networks that build trust and reputation with one another, and the same models increasingly control physically embodied robots as well as software.
By Messi H. J. Lee
arXiv:2607. 18966v1 Announce Type: new Abstract: Language models trained with reinforcement learning may learn to optimize the grader's judgment rather than the intended objective.
By Axel H{\o}jmark, J\'er\'emy Scheurer, Evgenia Nitishinskaya, Felix Hofst\"atter, Jason Wolfe, Theodore Ehrenborg, Bronson Schoen, Alexander Meinke
arXiv:2609.40306v1 Announce Type: cross
Abstract: Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and phys...
By Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang
arXiv:2606. 08919v1 Announce Type: new Abstract: As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approval gate: risky actions pause and wait for a person.
By Emre Turan
arXiv:2606. 28345v1 Announce Type: cross Abstract: LLM-governed social robots increasingly decide who receives real-world assistance first.
By Carmen Ng, Gjergji Kasneci
arXiv:2607. 28609v2 Announce Type: replace Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world.
By Qiushi Sun, Kanzhi Cheng, Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu, Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong
arXiv:2606. 05395v1 Announce Type: cross Abstract: Reusable robot skills are becoming the basic units through which embodied agents turn open-ended instructions into long-horizon physical behavior.
By Yunhao Yang, Neel P. Bhatt, Kevin Wang, Samuel Tetteh, Zhangyang Wang, Ufuk Topcu
arXiv:2609.01354v1 Announce Type: cross
Abstract: Reinforcement learning with verifiable rewards (RLVR) and standard benchmark evaluation both rely on an automatic verifier that turns a free text ans...
By Esther Xin
arXiv:2606. 26300v1 Announce Type: new Abstract: A classical intuition holds that verifying a solution is easier than producing one.
By Binghai Wang, Chenlong Zhang, Dayiheng Liu, Jiajun Zhang, Jiawei Chen, Mouxiang Chen, Rongyao Fang, Siyuan Zhang, Xuwu Wang, Yuheng Jing, Zeyao Ma, Zeyu Cui