arXiv AI

Simulate, Reason, Decide: Scientific Reasoning with LLMs for Simulation-Driven Decision Making

arXiv:2606. 04505v1 Announce Type: new Abstract: Scientific simulators are increasingly being integrated into LLM-driven systems for high-stakes simulation-driven decision-making.

arXiv AI
Jun 16

VeriGraph: Towards Verifiable Data-Analytic Agents

arXiv:2606. 16603v1 Announce Type: cross Abstract: LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a reliance on linear text trajectories makes their reasoning difficult to audit.

By Jiajie Jin, Zhao Yang, Wenle Liao, Yuyang Hu, Guanting Dong, Xiaoxi Li, Yutao Zhu, Zhicheng Dou
arXiv AI
Jun 17

EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning

arXiv:2511. 01650v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly entering specialized, safety-critical engineering workflows governed by strict quantitative standards and immutable physical laws, making rigorous evaluation of their reasoning capabilities imperative.

By Ayesha Gull, Muhammad Usman Safder, Rania Elbadry, Fan Zhang, Veselin Stoyanov, Preslav Nakov, Zhuohan Xie
arXiv AI
Aug 26

LLM Agents Perform Controlled Experiments Using Simulation Models

The paper introduces a multi‑agent framework that lets large language models (LLMs) perform controlled experiments using scientific simulation models, specifically for pharmaceutical process design. Given a user query and baseline configuration, the system builds a structured task, designs and runs comparative simulations, interprets outcomes, and generates evidence‑based recommendations for optimizing process parameters. By integrating high‑fidelity simulations with LLMs, the approach yields more specific, actionable outputs and improves user‑rated correctness and helpfulness compared to language‑only reasoning.

By Yuchen Xia, Michael Weyrich, Nasser Jazdi, Johannes St\"umpfle, Johannes Sigel, Akshay Narla, Gavin K. Reynolds, Anna Jawor-Baczynska, Pol Llopart
arXiv AI
Jun 24

Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs

arXiv:2606. 23938v1 Announce Type: new Abstract: Driving VLA models incorporating Chain-of-Thought (CoT) reasoning are attractive because they leverage pretrained VLM representations and expose intermediate decisions in natural language, yet current rationales often lack the step-by-step decision semantics needed to keep the rationale causally connected to the planned motion.

By Xiangbo Gao, Xiukun Huang, Boyu Lu, Junge Zhang, Mengjie Mao, Jiachen Li, Wei Xiong, Zhengzhong Tu
arXiv AI
Sep 4

Discovering High Level Patterns from Simulation Traces

The paper proposes an unsupervised learning approach that uses program synthesis to translate detailed simulation traces into sparse, high‑level structural patterns, making them easier for large language models to interpret. These pattern detectors can be guided by human‑provided labels such as "rigid collision" or "stretching spring" and produce transparent, explainable functions mapping system states to concise annotations. Experiments on a physics benchmark show that the annotated representations improve natural language reasoning about specific physical systems and enable natural‑language goals to be converted into reward programs for solution search.

By Sean Memery, Kartic Subr
arXiv AI
Aug 11

TempoBench: Reasoning Execution Without Causal Attribution Is Just Simulation

arXiv:2510. 27544v3 Announce Type: replace Abstract: Current training paradigms, optimized for long-horizon reasoning trace execution, have made Large Language Models (LLMs) excel at pattern matching and forward simulation of reasoning, but underperform at counterfactual causal understanding and reasoning.

By Nikolaus Holzer, William Fishell, Baishakhi Ray, Mark Santolucito