arXiv AI

DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions

DNative‑Twin is a graph‑native digital twin that records an AI agent’s committed decision as a typed trajectory, linking observed state, decision path, and authority. It re‑executes the decision mechanism under declared conditions, synchronizing and replaying the process in isolation to compare outcomes under controlled changes. Experiments on enterprise decision logs show that adding replay‑contract state and verification evidence improves unresolved‑divergence recall from 0 to 1.0, while end‑to‑end processing time rises from 0.794 to 8.889 seconds across 500–5,000 cases.

arXiv AI
Sep 24

TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents

TwinCheck is an inference‑time verification policy for stateful tool agents that only replaces a proposed tool call when a trace‑grounded counterfactual alternative, called a negative twin, satisfies structural checks and is preferred by a pairwise verifier in both candidate orders. The method uses exact replay to isolate intervention effects, and in experiments on 159 multi‑turn BFCL V4 tasks, it increased GPT‑5.6 Sol’s task success from 45.3% to 58.5% without any observed success‑to‑failure regressions.

By Jiaxuan Dai, Tianyi Huang
arXiv AI
Sep 2

Dr. Claw: An AI Scientist Workspace for Vibe Research

Dr. Claw is an open‑source AI scientist workspace that integrates existing command‑line coding agents into a single, auditable, human‑in‑the‑loop workflow. It uses persistent state objects, a reusable skill library, and multi‑executor coordination to link human decisions with AI execution, creating a traceable and recoverable loop for planning, execution, and writing. The authors demonstrate the system with an interactive scenario and a failure‑recovery walkthrough, and show that, when the underlying executor is held constant, Dr. Claw achieves higher research completeness while preserving an auditable process trail.

By Dingjie Song, Hanrong Zhang, Dawei Liu, Yixin Liu, Zongxia Li, Zhengqing Yuan, Siqi Zhang, Henry Peng Zou, Zhiling Yan, Yuxuan Zhang, Yanfang Ye, Philip S. Yu, Lichao Sun
arXiv AI
Sep 18

CoreSense: Traceable Failure Recall and Conflict-Aware Belief Gating for Auditable Robot Decisions

CoreSense is a robot‑system integration architecture that traces episodic evidence and uses a conflict‑aware belief gate to decide whether to proceed, re‑observe, abstain, or escalates. The gate evaluates scope, provenance, time, contradiction, and support before making a recommendation. Evaluation on public robot datasets, simulations, and a live cloud deployment shows that belief gating can eliminate protocol‑defined unsafe proceeds while maintaining auditability.

By Zoe Li
arXiv AI
Sep 21

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

arXiv:2609.21423v1 Announce Type: new Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to...

By Siyuan Liu (Fudan University, Meituan Longcat Team), Fan Yu (Fudan University, Meituan Longcat Team), Dongyu Ru (Meituan Longcat Team), Yizhu Liu (Meituan Longcat Team), Yifan Yang (Meituan Longcat Team), Xuezhi Cao (Meituan Longcat Team), Xunliang Cai (Meituan Longcat Team), Yixin Cao (Fudan University)