World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing the environment dynamics.
arXiv:2608. 13456v1 Announce Type: new Abstract: World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution.
By Avinash Kori, Fabrizio Russo
arXiv:2609.36985v1 Announce Type: cross
Abstract: The central challenge of world modeling is to learn representations that capture how the world evolves. However, existing world models predominantly...
By Ziqi Liu, Songhan Yang, Linfan Zhou, Jiatong Liu, Lijun Peng, Long Wan, Yinqi Bai
arXiv:2609.26037v1 Announce Type: new
Abstract: Explanations are central to causal reasoning, and cognitive science has long established that the human drive to explain is itself a mechanism for lear...
By Nicholas Tagliapietra, Florian Peter Busch, Moritz Willig, Matej Ze\v{c}evi\'c, Lavdim Halilaj, Juergen Luettin, Kristian Kersting
arXiv:2508. 11214v2 Announce Type: replace-cross Abstract: Explanations of cognitive behavior often appeal to computations over representations.
By Atticus Geiger, Jacqueline Harding, Thomas Icard
Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which may omit the dynamics most relevant to the agent's current decision.
The paper introduces a computational model that encodes symbolic knowledge as mental programs combining natural language and source code, and uses LLM-guided Bayesian learning to sequentially infer these programs. It demonstrates that this approach satisfies data‑efficiency, uncertainty handling, and flexibility, reproducing human inductive learning and active inquiry behaviors such as anchoring and garden‑pathing. In contrast, pure LLMs and classic Bayesian models either fail the task, do not match human behavior, or require prohibitive computational resources.
By Wasu Top Piriyakulkij, Sam Acquaviva, Cassidy Langenfeld, Joshua Tenenbaum, Kevin Ellis
Large language models offer a promising foundation for chemical reasoning, bringing together chemical knowledge and multistep problem solving. Chemical intuition can provide an initial sense of plausi...
Latent JEPA is a new framework that trains continuous latent thoughts to anticipate informative aspects of future solutions in chemical reasoning, without verbalizing every intermediate step. It combines autoregressive learning with joint-embedding prediction of one or more future views, using textual and molecular prediction objectives that link latent thoughts to subsequent reasoning and molecular outcomes. Experiments on ChemCoTBench demonstrate improvements in molecular optimization, editing, and reaction metrics, and representation analyses show that future prediction makes latent thoughts more informative about molecular outcomes and better aligned with chemical structure.
By Xinjian Zhao, Yaoyao Xu, Xuemin Chen, Xiaozhuang Song, Tianshu Yu
arXiv:2508. 12448v2 Announce Type: replace-cross Abstract: In-context learning (ICL) lets large language models (LLMs) solve new tasks from prompts alone, across an ever-widening range of domains, yet the mechanisms underlying this ability remain poorly understood.
By Yeongwoo Song, Jaeyong Bae, Dong-Kyum Kim, Hawoong Jeong
arXiv:2606. 30481v1 Announce Type: cross Abstract: Current large language models are extraordinary statistical engines.
By Ziqin Yuan, Jaymari Chua
arXiv:2606. 11445v1 Announce Type: new Abstract: Trust in an AI system is often anchored by explanations of how it works, which one then uses to forecast its behavior on new inputs.
By Mosh Levy, Yoav Goldberg, Asa Cooper Stickland