TwinCheck is an inference‑time verification policy for stateful tool agents that only replaces a proposed tool call when a trace‑grounded counterfactual alternative, called a negative twin, satisfies structural checks and is preferred by a pairwise verifier in both candidate orders. The method uses exact replay to isolate intervention effects, and in experiments on 159 multi‑turn BFCL V4 tasks, it increased GPT‑5.6 Sol’s task success from 45.3% to 58.5% without any observed success‑to‑failure regressions.
By Jiaxuan Dai, Tianyi Huang
arXiv:2607. 23002v1 Announce Type: cross Abstract: Large language models increasingly write both code and the tests meant to check it; coverage records what ran, not what was verified.
By Jeff Otterson (W. P. Carey School of Business, Arizona State University)
DNative‑Twin is a graph‑native digital twin that records an AI agent’s committed decision as a typed trajectory, linking observed state, decision path, and authority. It re‑executes the decision mechanism under declared conditions, synchronizing and replaying the process in isolation to compare outcomes under controlled changes. Experiments on enterprise decision logs show that adding replay‑contract state and verification evidence improves unresolved‑divergence recall from 0 to 1.0, while end‑to‑end processing time rises from 0.794 to 8.889 seconds across 500–5,000 cases.
By Junjie Pang, Zhenzhen Xie, Haoke Han, Ying He, Jing Wang, Gang Liu
The paper proposes a distributional theory explaining how large language models (LLMs) incorporate external evidence into their decision-making process. It identifies three key predictions: (1) evidence is more persuasive when it aligns with the model’s prior beliefs, (2) models more readily accept errors from their own internal processes than from external sources, and (3) the same evidence can improve weaker models while harming stronger ones. Extensive experiments across ten million trials, twelve LLMs from four families, and eight domains—including quantum mechanics, physics, genetics, and molecular biology—confirm these predictions and reveal that evidence integration occurs late in the network as a structured sequence of steps rather than through a simple trust metric.
By Sebastien Kawada, Manolis Kellis
arXiv:2605. 10246v2 Announce Type: replace Abstract: AI scientist systems are increasingly deployed for autonomous research, yet their academic integrity has never been systematically evaluated.
By Zonglin Yang, Xingtong Liu, Xinyan Xu
The paper investigates how multi‑modal world models can produce inconsistent outputs across different modalities, such as a video showing a ball not rebounding while a text description indicates it should. It defines two types of misalignment—internal (between modalities) and external (against a physical environment)—and introduces a physics‑grounded pipeline to measure these discrepancies. Experiments across multiple settings reveal that while the model’s language output matches the true environment, its video output frequently disagrees, indicating current unified backbones struggle with simultaneous reasoning, consistency, and physical fidelity.
By Geigh Zollicoffer, Minh Vu, Rajiv Ranasinghe, Manish Bhattarai