The paper investigates the safety of Code World Models, where a language model generates an executable world model that a planner uses, and the model is accepted if it reproduces sampled transitions. It defines the pipeline’s danger as the expected risk, showing that the probability of missing a critical event across N independent rollouts is (1‑r)^N, and that an additional acceptance sample adds to the exponent. Experiments on hybrid instruments reveal that mode‑blind models can be exploited, and the authors provide theoretical bounds on localization budgets and demonstrate that acceptance only guarantees sample consistency, covering about two percent of the planner’s queries.
The paper studies how a code‑world model can be perfectly accurate on the portion of the state space that a sampling gate can observe while potentially being arbitrarily wrong elsewhere. By treating the unobservable interior as an annular freeze mode, the authors formalize the notion of reach and show that acceptance with certainty fixes the model only on the reachable query set, leaving the rest as a gauge. Experiments on a minimal ring instrument demonstrate that a single channel width parameter can move the model through regimes of being unfalsifiable and harmless, falsifiable and costly, or instantly falsified, illustrating how topology relative to reach governs danger, repair, and mitigation strategies.
By Javier Aguilar Mart\'in
arXiv:2609.06036v1 Announce Type: new
Abstract: Proposal-based controllers---learned policies, language-model planners, and other black-box \emph{generators}---are increasingly deployed behind runtim...
By Guangxi Wan, Yongbo Xie, Yuqi Liu, Qingwei Dong, Qingxin Li, Hongfei Bai, Peng Zeng
arXiv:2607. 14169v1 Announce Type: new Abstract: Large language models can synthesize a game's rules as executable code - a Code World Model (CWM) - which a classical planner then searches over.
By Javier Aguilar Mart\'in
arXiv:2607. 23386v1 Announce Type: new Abstract: We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional rules of the form "A is required UNLESS B applies, UNLESS C overrides B".
By Paul Simpson, John Kozak, Lisa Doake
arXiv:2607. 07405v1 Announce Type: new Abstract: Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully.
By Vikas Reddy, Sumanth Reddy Challaram, Abhishek Basu