TwinCheck is an inference‑time verification policy for stateful tool agents that only replaces a proposed tool call when a trace‑grounded counterfactual alternative, called a negative twin, satisfies structural checks and is preferred by a pairwise verifier in both candidate orders. The method uses exact replay to isolate intervention effects, and in experiments on 159 multi‑turn BFCL V4 tasks, it increased GPT‑5.6 Sol’s task success from 45.3% to 58.5% without any observed success‑to‑failure regressions.
By Jiaxuan Dai, Tianyi Huang
arXiv:2606. 19616v1 Announce Type: cross Abstract: Autonomous coding agents now open millions of pull requests, yet large-scale studies find their PRs are produced faster but accepted less often - a coordination and trust gap that pull-request-level telemetry cannot explain.
By Dipankar Sarkar
arXiv:2606. 16415v1 Announce Type: new Abstract: Enterprise behavioral simulation requires more than producing a plausible response.
By Ankit Das (Twinning Labs)
arXiv:2609.08015v1 Announce Type: new
Abstract: Long-running AI agents may read state, reason, wait for tools or human approval, and perform an external action much later. The state that justified th...
By Yongjian Lyu, Yang Ren, Ruofei Lai, Wenting Liu
arXiv:2609.24264v1 Announce Type: new
Abstract: Tool-use agent traces identify messages and API calls, but procedural analyses also need explicit units of action and inspectable links to their eviden...
By Songqi Li, Dongqing Li, Zheqiao Cheng
arXiv:2609.13672v1 Announce Type: new
Abstract: AI agents can be interrupted while editing files, calling tools, or carrying out multi-step tasks. Restarting repeats completed work, but continuing fr...
By Zhihui Zhang, Wei Liu
arXiv:2606. 10241v1 Announce Type: new Abstract: Autonomous improvement loops are hard to trust because the improvement process is usually external scaffolding bolted onto the agent: failures go unlogged, diagnoses cannot be replayed, and promote-or-discard decisions land in a side database rather than the agent's own history.
By Yohei Nakajima
Dr. Claw is an open‑source AI scientist workspace that integrates existing command‑line coding agents into a single, auditable, human‑in‑the‑loop workflow. It uses persistent state objects, a reusable skill library, and multi‑executor coordination to link human decisions with AI execution, creating a traceable and recoverable loop for planning, execution, and writing. The authors demonstrate the system with an interactive scenario and a failure‑recovery walkthrough, and show that, when the underlying executor is held constant, Dr. Claw achieves higher research completeness while preserving an auditable process trail.
By Dingjie Song, Hanrong Zhang, Dawei Liu, Yixin Liu, Zongxia Li, Zhengqing Yuan, Siqi Zhang, Henry Peng Zou, Zhiling Yan, Yuxuan Zhang, Yanfang Ye, Philip S. Yu, Lichao Sun
Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve. In conventional transcript-based multi-agent systems, humans typically provide an initial problem, agents deliberate internally, and the system returns a final response.
arXiv:2608. 13046v1 Announce Type: new Abstract: Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve.
By Sanjeev Manivannan
CoreSense is a robot‑system integration architecture that traces episodic evidence and uses a conflict‑aware belief gate to decide whether to proceed, re‑observe, abstain, or escalates. The gate evaluates scope, provenance, time, contradiction, and support before making a recommendation. Evaluation on public robot datasets, simulations, and a live cloud deployment shows that belief gating can eliminate protocol‑defined unsafe proceeds while maintaining auditability.
By Zoe Li
arXiv:2609.21423v1 Announce Type: new
Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to...
By Siyuan Liu (Fudan University, Meituan Longcat Team), Fan Yu (Fudan University, Meituan Longcat Team), Dongyu Ru (Meituan Longcat Team), Yizhu Liu (Meituan Longcat Team), Yifan Yang (Meituan Longcat Team), Xuezhi Cao (Meituan Longcat Team), Xunliang Cai (Meituan Longcat Team), Yixin Cao (Fudan University)