DNative‑Twin is a graph‑native digital twin that records an AI agent’s committed decision as a typed trajectory, linking observed state, decision path, and authority. It re‑executes the decision mechanism under declared conditions, synchronizing and replaying the process in isolation to compare outcomes under controlled changes. Experiments on enterprise decision logs show that adding replay‑contract state and verification evidence improves unresolved‑divergence recall from 0 to 1.0, while end‑to‑end processing time rises from 0.794 to 8.889 seconds across 500–5,000 cases.
By Junjie Pang, Zhenzhen Xie, Haoke Han, Ying He, Jing Wang, Gang Liu
The paper introduces Synthetic Universes, a benchmark that pairs well-known theoretical worlds with closely related twisted variants governed by noncanonical mechanisms. It evaluates scientific agents on each law twice: by testing predictive performance on unseen data and by checking if the law recovers the underlying generating mechanism. Results from a 60‑cell study show a dissociation between predictive adequacy and mechanism recovery, indicating that these are distinct scientific claims requiring separate tests.
By Kargi Chauhan
VeriHarness is a method that enhances verification for large language model agents tackling long‑horizon tasks without needing reference answers at test time. It transforms the base LLM into an agentic verifier by providing a workspace, evidence tools, and reusable verification skills, using disagreement resolution and consensus challenge to evaluate competing claims. Across five benchmarks and two frontier models, VeriHarness outperforms baselines, achieving significant performance gains and demonstrating self‑improvement of verification skills from failure feedback.
By Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu, Nigel Collier, Tomas Pfister, Chen-Yu Lee
The paper introduces the twin‑prefix framework to evaluate how the size of the verification unit—i.e., how many actions a pre‑execution LLM monitor reviews in one call—affects its performance. By pairing each gold plan with a twin that differs by a single write and injecting a controlled error, the authors isolate the impact of review length on catch rates and false rejections. Their findings show that longer review windows increase rejection rates but do not improve discrimination, with the highest informedness occurring at one or two actions across all judges and domains.
By Yuchen Han, Cheng Yan, Wuyang Zhang
arXiv:2603. 14465v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions.
By Shengda Fan, Xuyan Ye, Yupeng Huo, Zhi-Yuan Chen, Yiju Guo, Shenzhi Yang, Wenkai Yang, Shuqi Ye, Jingwen Chen, Haotian Chen, Xin Cong, Yankai Lin
arXiv:2608. 05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions.
By Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He, Peng Zhang, Jiangxia Cao, Yusheng Huang, Guohong Mu, Jian Liang, Ruiming Tang, Shuang Yang, Zhaojie Liu, Wenwu Ou, Kun Gai
arXiv:2609.21423v1 Announce Type: new
Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to...
By Siyuan Liu (Fudan University, Meituan Longcat Team), Fan Yu (Fudan University, Meituan Longcat Team), Dongyu Ru (Meituan Longcat Team), Yizhu Liu (Meituan Longcat Team), Yifan Yang (Meituan Longcat Team), Xuezhi Cao (Meituan Longcat Team), Xunliang Cai (Meituan Longcat Team), Yixin Cao (Fudan University)
The paper introduces A-CEGIS, a lightweight framework that employs counterexamples as feedback to evaluate and improve multi-turn natural-language-to-regex synthesis. In experiments on 30 NL-RX-Turk tasks, counterexample feedback enables agents to solve 90% of tasks within four turns, outperforming zero‑shot generation, generic self‑correction, and error‑only feedback. A full diagnostic run with hardening solves all hidden tasks by the final turn, achieving a mean time‑to‑success of 2.7 turns and robust success of 77% after targeted probing.
By Sidhesh Badrinarayan, Adithya Parthasarathy
The study investigates how language models equipped with tools can still produce unsupported final claims, even when a single tool call could resolve the uncertainty. It defines two metrics—occurrence (how often unsupported claims arise) and conditional repair (how often they are fixed when evidence is provided). Experiments on Qwen3-32B and Gemma 4 show that providing the missing evidence consistently repairs all unsupported claims in the Qwen3-32B setup, while the Gemma 4 model never produced unsupported claims under the tested conditions.
By Justin Bronder
arXiv:2608.25920v2 Announce Type: replace
Abstract: As large language model (LLM)-based multi-agent systems (MASs) are increasingly applied to long-horizon complex tasks, their reliability has emerge...
By Zhongwen Luan, Xiaoyu Zhang, Ming Hu, Yue Yang, Jiongchi Yu, Xiaohong Chen
arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.
By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran
arXiv:2607.01469v3 Announce Type: replace
Abstract: Agentic Video Question Answering (VideoQA) systems produce answers through adaptive reasoning and tool-use trajectories, yet standard practice eval...
By Rama AlHamidi, Aseel Mohamed, Rasul Khanbayov, Mohamed Rayan Barhdadi, Erchin Serpedin, Hasan Kurban