arXiv AI By Kargi Chauhan

Can AI Scientists Change Their Minds? Prior-Evidence Conflict in Synthetic Universes

Read the original on arXiv AI →

The paper introduces Synthetic Universes, a benchmark that pairs well-known theoretical worlds with closely related twisted variants governed by noncanonical mechanisms. It evaluates scientific agents on each law twice: by testing predictive performance on unseen data and by checking if the law recovers the underlying generating mechanism. Results from a 60‑cell study show a dissociation between predictive adequacy and mechanism recovery, indicating that these are distinct scientific claims requiring separate tests.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 24

TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents

TwinCheck is an inference‑time verification policy for stateful tool agents that only replaces a proposed tool call when a trace‑grounded counterfactual alternative, called a negative twin, satisfies structural checks and is preferred by a pairwise verifier in both candidate orders. The method uses exact replay to isolate intervention effects, and in experiments on 159 multi‑turn BFCL V4 tasks, it increased GPT‑5.6 Sol’s task success from 45.3% to 58.5% without any observed success‑to‑failure regressions.

By Jiaxuan Dai, Tianyi Huang
arXiv AI
Sep 4

DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions

DNative‑Twin is a graph‑native digital twin that records an AI agent’s committed decision as a typed trajectory, linking observed state, decision path, and authority. It re‑executes the decision mechanism under declared conditions, synchronizing and replaying the process in isolation to compare outcomes under controlled changes. Experiments on enterprise decision logs show that adding replay‑contract state and verification evidence improves unresolved‑divergence recall from 0 to 1.0, while end‑to‑end processing time rises from 0.794 to 8.889 seconds across 500–5,000 cases.

By Junjie Pang, Zhenzhen Xie, Haoke Han, Ying He, Jing Wang, Gang Liu
arXiv AI
Sep 7

Evidence Integration in Large Language Models

The paper proposes a distributional theory explaining how large language models (LLMs) incorporate external evidence into their decision-making process. It identifies three key predictions: (1) evidence is more persuasive when it aligns with the model’s prior beliefs, (2) models more readily accept errors from their own internal processes than from external sources, and (3) the same evidence can improve weaker models while harming stronger ones. Extensive experiments across ten million trials, twelve LLMs from four families, and eight domains—including quantum mechanics, physics, genetics, and molecular biology—confirm these predictions and reveal that evidence integration occurs late in the network as a structured sequence of steps rather than through a simple trust metric.

By Sebastien Kawada, Manolis Kellis
arXiv AI
Sep 15

One Model, Two Physical Stories: Auditing Misalignment in Multi-Modal World Modeling

The paper investigates how multi‑modal world models can produce inconsistent outputs across different modalities, such as a video showing a ball not rebounding while a text description indicates it should. It defines two types of misalignment—internal (between modalities) and external (against a physical environment)—and introduces a physics‑grounded pipeline to measure these discrepancies. Experiments across multiple settings reveal that while the model’s language output matches the true environment, its video output frequently disagrees, indicating current unified backbones struggle with simultaneous reasoning, consistency, and physical fidelity.

By Geigh Zollicoffer, Minh Vu, Rajiv Ranasinghe, Manish Bhattarai